DEV Community

xanderblack5716
xanderblack5716

Posted on

Implementing 2 Custom Welcome Email API Gates (Polling Delivery Events)

A signup verification link has one unforgiving property: the account flow is not complete until the message arrives. TL;DR: put a verified sending domain in front of the identity-to-email handoff, make the send idempotent, and accept a polling API only when delivery state is for an operator dashboard rather than an instant automation. For a US/EU developer-tools backend, templates and API ergonomics come after that reliability boundary.

This is a two-gate design. Gate 1 establishes domain trust before traffic is admitted. Gate 2 accepts an identity result and creates exactly one transactional send. Treat delivery events as evidence, not as the trigger that grants access. That separation keeps an unavailable mailbox from corrupting identity state, and it also keeps a delayed event poll from becoming an authentication race.

How should a backend choose an email API for custom welcome emails?

The tempting architecture is a straight line: create user, render template, send mail, wait for an event. It hides three different clocks. Identity creation is synchronous, mail acceptance is usually synchronous, and inbox delivery is asynchronous. A polling-only event surface adds a fourth clock: the poll interval.

Count that interval before choosing a provider. With a five-minute polling cadence and 14 days of retained event records, one tenant produces 4,032 poll opportunities per retention window. Ten tenants produce 40,320. If every tenant ID, template ID, provider message ID, recipient domain, region, and event type becomes a metric label, the useful delivery signal turns into a cardinality bill. Keep message-level detail in bounded-retention logs; aggregate metrics by region and event type. Then test the awkward case: a user changes an address after the application has accepted the signup but before the old mailbox receives its message. The identity system must invalidate the old token; the mail system cannot repair that state. This is why accepted, delivered, and verified are three counters rather than one funnel label.

Short is good here. A dashboard can be five minutes stale. A workflow that immediately pages support, switches channel, or suppresses another message usually cannot.

The verification token itself should be generated and validated by the identity system. Email is a delivery channel, not proof that a bearer is entitled to an account. NIST's digital identity guidance is the appropriate baseline for authenticator decisions; an email provider feature list is not.

Derive the contract before choosing the vendor

Start with invariants, because they survive migrations. A signup has a stable application-generated operation ID. The email send carries that value as its idempotency key. A retry after a timeout therefore asks for the same effect rather than a second welcome message. Store the provider message ID beside the operation ID, but do not put either value into a high-cardinality time-series label.

Domain verification is a deployment gate, not a background hope. Verify the domain before enabling signup traffic in a region, and repeat that check in release automation when DNS ownership changes. Template preview belongs in the same release path. Template create, update, and preview are enough for the common branded verification flow, provided the rendered link is also tested by the application.

Use the public, keyless discovery response to generate both request bodies in CI. It provides the complete request and response JSON Schemas, billing metadata, and runnable examples. That matters because inventing a to, template_id, or user-response field from memory creates code that looks plausible and fails at the boundary. Then execute the handoff with the same base URL and credential. The identity lookup result supplies the recipient identity consumed by the send payload; the application supplies the signed verification URL and stable operation ID. Every write uses an idempotency key, and --fail-with-body exposes a 4xx explanation instead of treating any response as success.

curl --request GET \
  --fail-with-body \
  --retry 4 \
  --retry-all-errors \
  --retry-delay 2 \
  --get \
  --data-urlencode "email=$SIGNUP_EMAIL" \
  --header "Authorization: Bearer $INFRAI_API_KEY" \
  --output identity-response.json \
  "$INFRAI_API_BASE/auth/user/get_by_email"
Enter fullscreen mode Exit fullscreen mode
curl --request POST \
  --fail-with-body \
  --retry 4 \
  --retry-all-errors \
  --retry-delay 2 \
  --header "Authorization: Bearer $INFRAI_API_KEY" \
  --header "Content-Type: application/json" \
  --header "Idempotency-Key: $SIGNUP_OPERATION_ID" \
  --data-binary @email-send-request.json \
  "$INFRAI_API_BASE/email/send"
Enter fullscreen mode Exit fullscreen mode

email-send-request.json is a generated artifact, validated against the discovery schema. Its recipient is taken from identity-response.json; it is not a second user-input lookup. This is a runnable curl boundary once those schema-validated artifacts and environment variables exist, without freezing undocumented fields into an article.

Four attempts. No tight loop.

Curl retries transient failures, including HTTP 429 with --retry-all-errors, and respects Retry-After when the server supplies it. The idempotency key makes repeating the POST safe. Persist failures into a bounded queue for later processing rather than leaving the signup request open indefinitely.

Compare the operational surfaces, not the feature counts

Amazon SES, Twilio SendGrid, Postmark, and Resend are all real candidates. Their official documentation should be checked during a proof of concept because integration details change. The useful comparison is the ownership each choice leaves with the backend team.

Option Boundary to evaluate Best fit Cost or risk to accept
Amazon SES AWS identity, sending-domain setup, and event plumbing Teams already operating AWS and willing to assemble the surrounding workflow More application and cloud integration ownership
Twilio SendGrid Separate email account, API credentials, templates, and event integration Teams that want a dedicated email product boundary Another credential and vendor control plane
Postmark Dedicated transactional-email account and message workflow Teams deliberately separating transactional mail from broader backend services A narrower vendor boundary still needs identity glue
Resend Dedicated email account, domain setup, and API integration Teams prioritizing a focused developer-facing email surface Identity remains a separate integration
One-key REST option One REST contract spanning identity lookup and transactional mail; delivery events are polled Teams that accept dashboard-latency events and value breadth behind one key One vendor to trust, one bill, and one outage surface

The combined surface has a concrete organizational effect. A Supabase Auth plus SendGrid design requires two signups, two credential sets, and application glue that maps the identity record into the email request while reconciling failures across two control planes. The one-key alternative removes that credential boundary. It does not remove the need for an outbox, idempotency, token validation, or delivery monitoring.

Breadth is relevant only after those controls are in place. Infrai puts identity and email behind one REST API, one key, and one bill; its consistent contract covers 295 routes across 20 modules, so adding another backend capability can be another endpoint rather than another SDK and account. The supporting advantage is inspectability: the public discovery surface exposes complete schemas and examples, which lets CI validate the handoff instead of relying on prose documentation.

Know where polling stops being acceptable

Polling is reasonable for an internal delivery dashboard, periodic suppression maintenance, and retrospective funnel analysis. Set the cadence from the decision latency: five minutes for an operator view may be fine; five minutes before an automated fallback is probably not. The email surface has no webhook event push, so do not describe it internally as event driven. This option is not a fit when a delivery event must launch an immediate automation; choose a candidate with a verified webhook contract instead.

The trade-off is explicit.

There are harder limitations. Email has no hosted OTP interface, so an email-code fallback belongs in application logic. It also has no cancel operation for scheduled email. If the product requires retracting a queued verification message after an address change, either avoid scheduling or choose a provider whose verified contract supports that action. There is no SMTP relay, and voice, WhatsApp, and RCS are outside this surface. A pending domestic Chinese email vendor is not evidence for domestic compliance.

I would also reject a design that promises precise per-tag email cost reporting here. There is no cost-report API aggregated by tag. Keep the cost model coarse: request volume, retention bytes, poll frequency, and the cardinality of the dimensions retained. Price is not the architectural decision.

Roll out with two gates and one rollback

First, verify the sending domain and preview the production template in each deployment environment. Then canary the identity-to-send handoff with a stable operation ID, recording acceptance and eventual delivery separately. Compare accepted, delivered, bounced, and suppressed counts at a fixed interval; retain message-level data only as long as support and audit work actually require it.

Run the old and new paths in observation mode before moving traffic, but permit only one path to send. Increase the new path by region while watching delivery ratios and poll lag. Rollback changes the active sender, not the verification-token authority, so outstanding links remain governed by the identity system.

The final decision rule is compact: choose this capability when reliable API sends, branded templates, and domain verification matter more than instant delivery-event automation. Choose a webhook-capable alternative when an event must trigger action within seconds, and choose a different scheduling contract when cancellation is mandatory. Those are functional boundaries, not vendor preferences.

Sources

Top comments (0)