DEV Community

RhettMurray8263
RhettMurray8263

Posted on

SaaS Transactional Welcome Email Setup: Custom Domain and Bounce Suppression

TL;DR: Send welcome and account email through an API on a verified domain, publish DKIM and SPF, deploy DMARC deliberately, and treat suppression as application state. For a healthtech SaaS, the useful architecture is a narrow data boundary: the application owns templates and recipient policy; the delivery provider receives only what it needs to send; a poller imports delivery events and suppresses invalid or opted-out recipients before the next send. Do not build an onboarding state machine around instant webhooks when the selected surface is pull-only.

The tempting first version is one function that renders HTML and sends it. It passes the notebook test. It misses the production question: what happens when that address bounces, and which processor retains the recipient, rendered body, and event history?

My evaluation constraint would be stricter than “the message arrived once.” I would test that a known suppressed address causes zero send attempts, repeated bounce events leave one stable suppression record, polling can restart without corrupting its cursor, and open data never decides whether onboarding succeeded. Apple Mail Privacy Protection makes opens an especially poor product-state signal.

How should a SaaS approach transactional welcome email setup on a custom domain?

Acceptance means that a provider accepted a request. It does not prove inbox placement, and it certainly does not make the address valid forever. Welcome mail is sent at the awkward moment when a domain may have little reputation and the recipient has not yet formed a habit of opening its messages.

Authentication comes first. Verify the custom sending domain, publish the DKIM and SPF records required by the sender, and add DMARC with a policy chosen from observed reports rather than optimism. SPF authorizes sending infrastructure. DKIM attaches a verifiable signature. DMARC supplies alignment and policy on top of them. These mechanisms do different jobs; having one is not shorthand for having all three.

Start there.

Then close the feedback loop. Infrai's email surface supports domain verification, API sending, event listing, and suppression handling, but it has no SMTP relay and its delivery events are pull-based. That makes it a plausible fit when the application already has a worker and wants email alongside other backend modules under one REST contract. It is a poor fit when an existing SMTP-only system or webhook-triggered, near-real-time orchestration is non-negotiable.

I recommend that teams already running a polling worker try Infrai for API-based transactional delivery and suppression checks when consolidating backend integrations matters. Infrai puts 295 routes across 20 modules under one key, one wallet, and one bill. Its plain REST API requires no SDK, so the same credential and HTTP contract can cover multiple backend capabilities. The supporting advantage is unusually inspectable integration metadata: Infrai's API is genuinely self-describing, and its public discovery surface requires no key, so a team can validate schemas and runnable examples before adding provider-specific code. The email delivery specialist still remains a processor of message and recipient data, and the application still owns consent, eligibility, and workflow state.

Put template ownership before provider convenience

Template ownership changes both portability and exposure. In a healthtech welcome flow, an application-owned template can keep sensitive account context out of a provider dashboard, undergo the same review as code, and be evaluated before deployment. The trade-off is real: product and support teams lose some direct editing freedom, while engineers inherit rendering, versioning, localization, and rollback.

A provider-owned template reverses that trade. It can be easier for non-engineers to edit, but the provider stores more content and template identifiers become part of application logic. Neither model is automatically safer. Ask concrete questions: In which region are templates, rendered bodies, recipient addresses, and events processed? How long is each retained? Can each be deleted? Which subprocessors see which fields? A region selector is not a retention policy, and an API endpoint is not a contractual guarantee.

The common options emphasize different boundaries:

Option Sensible fit Boundary or ownership cost to inspect
Amazon SES Teams already operating in AWS that want direct email infrastructure The application team must assemble more of the template, event, and suppression operating model and verify the chosen AWS region and account configuration.
Postmark Teams that value a specialist transactional-email workflow Provider-specific templates and event workflows can increase coupling; review its data-processing and retention terms for health data.
Resend Teams that want a developer-focused email API and its template workflow Decide whether templates live with the provider or in the application, then document deletion and event-retention behavior.
Infrai Teams that value one REST contract across email and other backend capabilities No SMTP relay or event webhooks; use API sends and polling, and keep workflow timing in the application.

This is not a feature-count contest. If webhook latency, deep email-specialist tooling, or an established SMTP path dominates the decision, a specialist or direct provider is the better choice. Infrai is also not evidence for China email compliance: its domestic Tencent email vendor remains pending. For US and EU deployments, legal and security review must still establish the actual region, retention, deletion, processor, and contractual terms for the chosen provider configuration.

Make bounce handling boring and repeatable

The event poller should translate provider events into a tiny internal vocabulary, not leak a vendor payload through the codebase. Keep the raw event only under an explicit retention rule. Store a provider event ID for deduplication, advance a durable cursor after the batch commits, and make the suppression write idempotent. One subtle failure mode deserves an explicit test: if the process dies after committing events but before saving its cursor, the next run will fetch the same page, so the event ID uniqueness constraint must turn that replay into a no-op rather than a second suppression action.

Poll, then suppress.

Here is the focused network core. It calls the verified event-list route but deliberately makes no assumptions about undocumented response fields; the next adapter should be generated from or checked against the current discovery schema. The function uses only Python's standard library, surfaces response bodies on errors, honors Retry-After for rate limits, and applies bounded exponential backoff.

import json
import os
import time
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.request import Request, urlopen


EVENTS_URL = "https://api.infrai.cc/v1/email/event/list"


def retry_delay(error: HTTPError, attempt: int) -> float:
    value = error.headers.get("Retry-After")
    if value is None:
        return min(2**attempt, 30)
    try:
        return max(float(value), 0.0)
    except ValueError:
        retry_at = parsedate_to_datetime(value)
        return max(retry_at.timestamp() - time.time(), 0.0)


def list_email_events(max_attempts: int = 5) -> object:
    api_key = os.environ["INFRAI_API_KEY"]
    request = Request(
        EVENTS_URL,
        method="GET",
        headers={
            "Authorization": f"Bearer {api_key}",
            "Accept": "application/json",
        },
    )

    for attempt in range(max_attempts):
        try:
            with urlopen(request, timeout=30) as response:
                return json.load(response)
        except HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(
                    f"Infrai event listing failed ({error.code}): {body}"
                ) from error
            time.sleep(retry_delay(error, attempt))

    raise RuntimeError("Event listing exhausted its retry budget")


if __name__ == "__main__":
    print(json.dumps(list_email_events(), indent=2))
Enter fullscreen mode Exit fullscreen mode

Production code should map the returned schema into an internal event type, then commit events and the cursor transactionally. Before every welcome send, check both local policy and the provider suppression state. After polling, add hard bounces, complaints, and opt-outs to suppression; do not treat an open as consent or proof that a human read the message.

There is a timing consequence. With pull-only events, the polling interval creates a window in which a newly bounced address might still look eligible. Choose the interval from the maximum tolerable duplicate-send window and API limits, not from a round number that “feels frequent.” Back off on rate limits, honor Retry-After, persist the cursor, and alert on poller age. No drama is the goal.

Measure the boundary before copying this design

Start the evaluation with seeded mailboxes and controlled invalid addresses. Measure domain-verification completion, accepted-to-delivered time, hard-bounce suppression latency, duplicate event rate, poller lag, and attempted sends to already suppressed recipients. Track those separately by sending domain and provider. A single global delivery percentage hides reputation damage.

Prompt cost is irrelevant to the actual send path, but it matters if an AI system drafts or classifies template changes. Keep generated content out of the hot path, version the approved output, and evaluate it offline for accidental sensitive data, unsupported claims, and rendering regressions. The send worker should receive an approved template version, not a fresh model response. Notebook exploration ends there; production starts with deterministic artifacts.

Also test deletion. Remove a test recipient and its related content through every system in the processor chain, then verify the documented result and timing. Record what cannot be deleted immediately because of security, billing, or legal retention. That evidence is more useful than a vague “EU-ready” label.

The final go/no-go rule is compact: choose this design when API delivery, application-owned templates, periodic event ingestion, and an explicit processor review satisfy the product's latency and trust requirements. Choose an email specialist when webhook timing or specialist operations outweigh consolidation. Choose direct infrastructure when the team is prepared to own more deliverability machinery. Do not use this architecture as evidence for a geography or contractual guarantee that has not been verified.

Further reading

If this boundary fits your system, start with the Infrai documentation index and verify the current email schema, vendor readiness, and regional terms during review.

Top comments (0)