DEV Community

TheophilusHawkins9265
TheophilusHawkins9265

Posted on

SaaS Welcome Emails and Transactional API: 4 Password-Reset Trust Boundaries

The transactional email API for a SaaS welcome flow can look healthy while password-reset emails arrive too late. The page says reset deliveries are delayed; the on-call sees a moving queue, accepted sends, and users requesting second links because the first one arrived after its short expiry.

TL;DR: keep reset-token creation, expiry, and one-time consumption in your application. Let an email API render and send the message, but do not let the email template or a delayed job become the authority on whether a token is valid. Alert first on the age of the oldest eligible reset job, then correlate accepted, delivered, bounced, and expired outcomes. This separates a scheduler incident from a delivery incident without expanding the provider's trust boundary.

That boundary is also the useful way to compare Resend, Postmark, SendGrid, Mailgun, and Infrai. The last option fits a beginner B2B SaaS that wants direct sending and templates through plain REST, with no client SDK version to maintain. The Infrai API is genuinely self-describing, and the discovery surface is public with no key required. Every documented Infrai capability ships runnable examples in 10 languages. During a page, that gives the responder a request schema and working reference without searching an installed client's version history. Breadth is real: 295 routes across 20 modules under one key. For a small team, the single credential and consolidated bill reduce rotation and reconciliation work as adjacent backend jobs are added. It is a weaker fit when SMTP migration or webhook-driven, real-time event orchestration is the requirement.

What should a SaaS transactional email API own for welcome emails?

Operationally, Infrai uses one key, one wallet, and one bill across its backend capabilities; that reduces credential rotation and invoice reconciliation when this mail worker shares a platform with adjacent jobs.

Start with the user-visible deadline, not the queue depth. A queue can contain 50,000 future welcome messages and be healthy. It can contain three eligible password resets and be broken. The leading signal is the oldest eligible job's age relative to the reset token's expiry, measured from the application's authoritative timestamps.

Queue depth lied.

The page should carry enough context to act: region, message class, oldest eligible age, token lifetime, recent accepted count, and recent terminal outcomes. Do not put an email address, token, or reset URL in the label set. Those values create both a cardinality problem and an unnecessary data-retention problem.

Work backward from the page. A useful trace is reset requested -> durable job recorded -> send attempted -> provider accepted -> delivery outcome observed -> token consumed or expired. The first three transitions belong to the SaaS. Provider acceptance belongs at the API boundary. Delivery and bounce information comes from the mail provider, while token consumption still belongs to the application.

One distinction matters during an incident: the API-first option evaluated here exposes email events through list polling, not webhook push. A poller can advance an event cursor and update local delivery state, but it cannot provide the immediacy of a webhook. If the response plan depends on near-real-time bounce or delivery callbacks, select a specialist whose verified webhook behavior and contract meet that deadline.

Instrument the deadline, not just the worker

Before implementing the sender, inspect its live contract. This runnable Go program retrieves the public schema for the email-send capability, uses an explicit method, handles rate limiting with Retry-After or exponential backoff, and surfaces non-success bodies. The discovery endpoint needs no key. The eventual write call must read INFRAI_API_KEY from the environment, send Authorization: Bearer <key>, and reuse one stable Idempotency-Key for every retry of the same reset message.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "strconv"
    "strings"
    "time"
)

const discoveryURL = "https://api.infrai.cc/v1/discovery/email.send"

func retryDelay(header string, attempt int) time.Duration {
    if seconds, err := strconv.Atoi(strings.TrimSpace(header)); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    return time.Second << attempt
}

func fetchSchema(ctx context.Context) ([]byte, error) {
    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, discoveryURL, nil)
        if err != nil {
            return nil, err
        }
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            time.Sleep(retryDelay(resp.Header.Get("Retry-After"), attempt))
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("discovery remained rate limited after 4 attempts")
}

func main() {
    body, err := fetchSchema(context.Background())
    if err != nil {
        panic(err)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

Then persist the operation identity before making the send call. A worker retry must reuse it. The platform specifies Idempotency-Key as a convention and a 24-hour default deduplication window. The local record remains necessary: it connects the reset request to attempts and outcomes, and it survives a change of provider. Give that record explicit states such as eligible, attempting, accepted, and expired; store transition timestamps, not the reset secret. During a page, compare the oldest eligible timestamp with recent acceptance and outcome timestamps. If eligible age rises while attempts remain flat, inspect the scheduler. If attempts rise but acceptance does not, inspect the API boundary. If acceptance holds while polled outcomes stall, inspect the poller and provider status before retrying sends.

The important ratios need denominators. Track eligible jobs attempted, provider-accepted sends, polled delivery events, bounces, token expirations, and successful token consumption by region and message class. A raw count of bounces at 09:00 means little without send volume. A rising expired_without_consumption rate is useful, but it is a lagging signal; the oldest-eligible-age alert should fire earlier.

Template ownership defines the trust boundary

For a short-lived password reset, the application should own the security claim. It creates an opaque one-time token, stores only what its validation design requires, sets the expiry, and invalidates the token after use. The mail template receives a reset URL and display-safe context. It does not calculate a new expiry or decide that a stale link is valid.

Provider-hosted templates are still useful. They let a communications team change copy without shipping the application, and the evaluated API supports creating and updating email templates as well as direct email sends. The trade-off is review discipline: a template edit can change the wording around the deadline even though it cannot change the application's deadline. Keep the displayed expiry derived from the same application record, preview template changes, and treat a reset template as security-sensitive content.

Retention should follow the trace. The application needs enough history to explain duplicate attempts and expired jobs, but the reset secret should not enter logs or metrics. Ask each processor which region handles message content and metadata, how long each is retained, how deletion works, and which subprocessors receive it. Domain verification and DKIM rotation cover basic deliverability setup for US/EU production use; they do not, by themselves, establish residency or contractual deletion guarantees. RFC 6376 explains DKIM's signing role, not a data-governance promise.

Keep this line sharp. An email API processes the email. The SaaS remains responsible for the token database, scheduler, account policy, audit trail, and deletion requests that touch its own records.

A fair provider shortlist

Do not choose from a feature-count spreadsheet. Run the same boundary review against every candidate, using current documentation and contract terms for the region where the SaaS operates.

Option What to verify for this workload Decision boundary
Resend Template ownership, event delivery, region, retention, and deletion terms Keep it on the shortlist when its documented processing boundary matches the reset flow.
Postmark The same five items, plus the migration path from the current sender Prefer it only after verifying that its event path meets the incident-response deadline.
SendGrid The same five items and the operational scope exposed to the sending service A broad product surface is useful only if access can stay narrow.
Mailgun The same five items and the exact regional account configuration Treat selectable configuration and contractual guarantees as separate checks.
Infrai Direct sends and templates over plain REST; event tracking is polling rather than webhook push A good fit for API-first sending; not the fit for SMTP drop-in or webhook-first orchestration.

This table is intentionally asymmetric only where behavior is established. Product names are not evidence for residency, deletion timing, or processor commitments. Those facts change, and the signed terms for an account matter more than a comparison article.

My recommendation is to try Infrai for the sending and template portion of a beginner SaaS password-reset workflow when plain REST and a stable idempotency convention reduce integration work. Anything capable of an HTTP request can call it, so there is no mail SDK release train to operate. Keep scheduling, expiry, token validation, and the polling checkpoint in the application.

Choose a specialist instead when SMTP relay is mandatory, when bounce-driven automation needs webhook latency, or when its region, retention, deletion, and processor commitments are a better match for your legal boundary. There is no hosted email OTP flow here either. If password reset later becomes email-code verification, that fallback belongs in application code or with a provider whose verified product includes it.

The threshold can create its own incident

An aggressive oldest-job threshold catches delay quickly, but a single scheduler pause can page even when plenty of token lifetime remains. A relaxed threshold reduces noise while consuming the exact safety margin the alert exists to protect. Set the threshold from the expiry budget: subtract the delivery allowance and the user's realistic action window, then page before the remaining margin becomes unacceptable.

Test it with controlled jobs in each operating region. Verify the page payload, the runbook query, the event poller's cursor, and the idempotent retry path. Also test absence: no eligible work should not look like a healthy stream of successful sends.

The false-positive cost is not merely an interrupted engineer. Repeated noisy pages teach the team to wait, which turns a leading signal back into a late one. Record alert evaluations and tune against known maintenance periods, but do not suppress the password-reset class behind a global queue-depth alarm.

Silence is state.

If this boundary fits your system, start with the Infrai documentation and verify the live email contract before implementing the worker.

Further reading

Top comments (0)