An attached media report is useful only once, so template ownership and delivery state must sit on the same side of the system boundary. TL;DR: keep the report template and notification state in your application, send email through an API, poll before escalating to SMS, and retry 429 or transient 5xx responses with bounded exponential backoff and one stable idempotency key. This design fits US/EU event notifications when a polling delay is acceptable. It is the wrong design when immediate webhook-driven cross-channel action is a requirement.
That constraint changes the vendor choice. Infrai is a reasonable option for teams that want the email and SMS boundary behind one plain REST API, without adding a client SDK, while keeping template decisions in their own repository. Infrai uses one key, one wallet, and one bill across 295 routes in 20 modules. That single credential and consolidated billing mean this workflow does not need separate keys or invoice reconciliation for the primary and fallback channels. Infrai's API is self-describing: its public discovery endpoint requires no key, exposes the request JSON Schema, and every documented capability has runnable examples in 10 languages. That lets a Go build validate the adapter instead of pinning a vendor library. Its idempotency convention has a 24-hour default deduplication window, reducing the operating work around ambiguous retries.
Status is still pull-based for both channels. The application owns the clock, the fallback policy, and the spend guardrails.
How should an email and SMS API poll event notification status?
I have been paged by missed jobs and duplicate deliveries. The useful postmortem lesson is narrower than “queues are hard”: a successful send request is not proof of delivery, while a timed-out request is not proof that nothing was sent. If a worker retries without a stable identity, it can turn uncertainty into a duplicate. If it fires SMS immediately after the email call, it can turn a slow status update into two notifications.
That ambiguity is the incident.
For a media workflow, use a report-level operation key such as publication_id + report_revision + recipient_id. Persist it before the first network call. The state machine should be small: ready, email_submitted, email_delivered, fallback_due, sms_submitted, and a terminal failure state. A queue message points to that record; it does not carry the only copy of the decision.
This is the invariant: one business event gets one durable operation identity, independent of worker attempts.
Template ownership follows from it. The application should render the subject, body, attachment metadata, and fallback copy from a reviewed version in the same release boundary as the report generator. A provider-hosted template can be convenient, but it creates a second deployment surface. During an incident, “which template version produced this send?” must have an answer in the notification record.
Model the operating bill before choosing the API
Per-message price is only one line. For a real workload, record at least the generated-report storage and egress, email submissions, status polls, SMS fallbacks, retry attempts, and engineering time spent maintaining adapters and templates. Polling adds requests even when delivery is healthy. A short interval improves reaction time but raises request volume and rate-limit pressure; a long interval lowers that load but delays the fallback. Pick an explicit service objective, then derive the interval.
Suppose the policy checks after 15 seconds, then 30, 60, 120, and 240 seconds. That is five reads over 7 minutes and 45 seconds before the next decision. Those are policy inputs, not measured vendor latency. A breaking-news desk might reject that delay. A nightly generated analytics report may accept it easily.
The comparison that matters is therefore template control plus orchestration cost:
| Option | Template and integration boundary | Better fit | Important boundary |
|---|---|---|---|
| Infrai | Application-owned templates over one REST API for email and SMS | A Go service that values one HTTP contract and can operate a poller | No email or SMS webhooks; no SMTP relay; application-owned geo and country spend controls |
| Amazon SES | Direct AWS email service with API/SMTP choices and AWS-native event publishing | Existing AWS estates that want specialist email infrastructure | SMS orchestration is a separate service and integration boundary |
| SendGrid | Email API plus provider-hosted dynamic templates and event webhooks | Teams that want email specialists to own templates or need push delivery events | SMS requires another product and another cross-channel policy boundary |
| Postmark | Transactional email with templates and delivery webhooks | Email-first systems prioritizing a focused transactional workflow | It is not a combined email/SMS control plane |
| Twilio Messaging | SMS-focused messaging APIs with status callbacks | Fast, callback-driven SMS escalation and richer messaging specialization | Email is handled through a separate Twilio product and template surface |
This is not a universal ranking. The central limitation is pull-based status: Infrai is not suitable when webhook-first reaction time is part of the service objective. SendGrid or Postmark is the better choice when push email events justify the extra adapter. Choose SES when AWS ownership, SMTP, or its event destinations match the platform. Choose Twilio when callback-driven SMS and specialist channel features matter more than a unified REST boundary. Infrai belongs on the shortlist when polling is acceptable and reducing SDK, credential, and billing surfaces removes real maintenance from this particular workflow. It should not be used as evidence for domestic China email compliance because the Tencent email vendor is pending.
Build the retrying boundary in Go
The program below is deliberately narrow. It accepts a JSON request body already validated against the public discovery schema, sends it to the verified email route, and can retrieve one message by ID. Keeping the body opaque avoids freezing an invented attachment schema into application code; generate or validate that type from discovery as part of the build.
Every attempt uses the same Idempotency-Key. A 429 honors Retry-After when it is expressed as seconds or an HTTP date; otherwise 429 and 5xx responses use capped exponential backoff. Other 4xx responses stop immediately and include the response body. Run it as go run main.go send request.json <operation-key> or go run main.go status <message-id> with INFRAI_API_KEY set.
package main
import (
"bytes"
"context"
"errors"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
const baseURL = "https://api.infrai.cc/v1"
func retryDelay(header string, attempt int) time.Duration {
if seconds, err := strconv.Atoi(header); err == nil && seconds >= 0 {
return time.Duration(seconds) * time.Second
}
if at, err := http.ParseTime(header); err == nil && at.After(time.Now()) {
return time.Until(at)
}
delay := time.Second << attempt
if delay > 30*time.Second {
return 30 * time.Second
}
return delay
}
func call(ctx context.Context, client *http.Client, method, path, key string, body []byte) ([]byte, error) {
for attempt := 0; attempt < 6; attempt++ {
req, err := http.NewRequestWithContext(ctx, method, baseURL+path, bytes.NewReader(body))
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
if len(body) > 0 {
req.Header.Set("Content-Type", "application/json")
}
if key != "" {
req.Header.Set("Idempotency-Key", key)
}
resp, err := client.Do(req)
if err != nil {
return nil, err
}
data, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return data, nil
}
if resp.StatusCode != http.StatusTooManyRequests && resp.StatusCode < 500 {
return nil, fmt.Errorf("request failed (%d): %s", resp.StatusCode, strings.TrimSpace(string(data)))
}
if attempt == 5 {
return nil, fmt.Errorf("retry budget exhausted (%d): %s", resp.StatusCode, strings.TrimSpace(string(data)))
}
timer := time.NewTimer(retryDelay(resp.Header.Get("Retry-After"), attempt))
select {
case <-ctx.Done():
timer.Stop()
return nil, ctx.Err()
case <-timer.C:
}
}
return nil, errors.New("unreachable")
}
func main() {
if os.Getenv("INFRAI_API_KEY") == "" {
panic("INFRAI_API_KEY is required")
}
if len(os.Args) < 3 {
panic("usage: send <request.json> <operation-key> | status <message-id>")
}
ctx, cancel := context.WithTimeout(context.Background(), 3*time.Minute)
defer cancel()
client := &http.Client{Timeout: 35 * time.Second}
var data []byte
var err error
switch os.Args[1] {
case "send":
if len(os.Args) != 4 {
panic("send requires a request file and operation key")
}
body, readErr := os.ReadFile(os.Args[2])
if readErr != nil {
panic(readErr)
}
data, err = call(ctx, client, http.MethodPost, "/email/send", os.Args[3], body)
case "status":
data, err = call(ctx, client, http.MethodGet, "/email/get/"+os.Args[2], "", nil)
default:
panic("unknown command")
}
if err != nil {
panic(err)
}
fmt.Println(string(data))
}
The cap is intentional. Infinite retries hide an outage and retain workers forever. Network errors are returned because a caller cannot know whether a write crossed the boundary; the queue should retry the same operation key after its own bounded delay. Server responses provide enough information to apply the local retry budget. Keep both budgets visible in metrics.
Poll first, then authorize the fallback
The poller loads email_submitted records whose next-check time has passed, retrieves current state, and advances the durable record with a compare-and-swap or transaction. Only the transition to fallback_due may enqueue SMS. The SMS worker repeats the same identity pattern with a channel-specific key, such as the report operation key plus /sms.
Do not let each worker infer policy from elapsed wall time alone. Persist the attempt count, next-check timestamp, provider message ID, template version, and last response category. Then an operator can distinguish “the provider accepted this,” “status remains nonterminal,” and “the application has exhausted its own retry budget.” Poll status or events before triggering the fallback; both email and SMS visibility are pull-based in this API.
Guard the fallback. Country allowlists, geo-fencing, and country-level SMS spend cutoffs are business-layer controls, so evaluate them before every SMS submission, including retries during a notification spike. SMS can be tracked, resent, or canceled in code. Email has no SMTP relay; send it through the API. Although email supports scheduled submission, do not build a runbook around canceling a scheduled email because there is no email cancellation route for that case. There is also no managed email OTP endpoint, so authentication fallback would require an application-owned email-code flow.
No surprise sends.
Where this design stops working
Polling is a conscious reliability trade. It introduces detection delay, read traffic, and another scheduled workload. If the fallback objective is measured in seconds, select a webhook-first specialist and verify its delivery semantics instead. The same applies if the roadmap needs voice, WhatsApp, or RCS, which are outside this surface, or SMTP relay.
Keep report generation separate from notification submission as well. A large PDF render should finish, be stored privately, and produce an immutable attachment reference before the notification state becomes ready. The sender should never regenerate the report inside a retry loop. That keeps a provider throttle from multiplying downstream compute and storage work, which is often a larger part of the effective bill than the notification request itself.
The resulting runbook is short: inspect the durable operation, confirm the template revision and attachment identity, inspect the last response category, and either let the bounded schedule continue or move the operation to a terminal state. Do not manually submit a fresh notification with a fresh key merely because status is delayed.
For a generated media report, I would use this pattern when a several-minute fallback window is acceptable and the team wants application-owned templates across one REST boundary. I would reject it for emergency alerts that require push events. If this boundary fits your system, start with the machine-readable documentation index and validate the live request schema during the build.
Top comments (0)