TL;DR: For logistics SMS event notifications sent after payment settles, use a four-outcome Python state machine and make resend an authorized transition for reviewed failures, not an automatic reaction to carrier filtering. Register and verify the sender configuration and signature for each destination market before launch. Then poll delivery evidence, separating queued, delivered, failed, and carrier_rejected; a carrier rejection should close the automatic retry path until the sender or routing issue is reviewed.
The least complex dependable shape is payment event -> durable receipt record -> SMS adapter -> polled outcome. Keep the state machine in the application. The transport behind that adapter may change, but payment code, resend policy, and evaluation fixtures should not.
What should authorize SMS event notification resends after carrier failures?
A resend needs two independent approvals: the delivery outcome must be retryable, and the business policy must still allow another attempt. This distinction matters for US and EU traffic because registration and signature configuration belong to the destination market. Repeating a rejected message does not repair an incorrect sender configuration. It is also why troubleshooting a logistics receipt differs from troubleshooting a telehealth login verification code: urgency and expiry differ, but neither use case can retry its way around sender registration.
Treat queued as unfinished work, not failure. Treat delivered as terminal success. A generic failed result can enter review for a controlled resend, while carrier_rejected should stop automation and produce an operational task. Preserve the provider's reason next to the normalized outcome so an operator can tell transport trouble from a registration problem.
One rule keeps the boundary honest: the payment handler creates one receipt-delivery operation after settlement, and every later action refers to that operation. It never creates a fresh logical receipt merely because a poll was late.
Encode the policy before calling a provider
The following program keeps the decision core provider-neutral, then polls the documented Infrai status route at the adapter edge. It is runnable with SMS_API_BASE_URL, INFRAI_API_KEY, and SMS_MESSAGE_ID, and it makes the risky decision visible: a failed receipt may be claimed for one resend, but a queued, delivered, or carrier-rejected receipt may not. The version field models the conditional database update a production worker should perform before it calls a resend operation. The response is printed without guessing its JSON fields; production mapping should follow the live discovery schema.
from dataclasses import dataclass, replace
from enum import Enum
import json
import os
import time
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen
class Outcome(str, Enum):
QUEUED = "queued"
DELIVERED = "delivered"
FAILED = "failed"
CARRIER_REJECTED = "carrier_rejected"
@dataclass(frozen=True)
class ReceiptDelivery:
order_id: str
country: str
outcome: Outcome
resend_count: int
version: int
def authorize_resend(
delivery: ReceiptDelivery,
expected_version: int,
allowed_countries: set[str],
resend_limit: int,
) -> ReceiptDelivery:
if delivery.version != expected_version:
raise RuntimeError("Delivery changed before the resend was claimed")
if delivery.country not in allowed_countries:
raise PermissionError("Destination country is not enabled")
if delivery.outcome is Outcome.CARRIER_REJECTED:
raise PermissionError("Carrier rejection requires registration review")
if delivery.outcome is not Outcome.FAILED:
raise RuntimeError(f"Cannot resend an outcome of {delivery.outcome.value}")
if delivery.resend_count >= resend_limit:
raise RuntimeError("Resend limit reached")
return replace(
delivery,
outcome=Outcome.QUEUED,
resend_count=delivery.resend_count + 1,
version=delivery.version + 1,
)
def fetch_sms_status(message_id: str) -> dict:
api_key = os.environ.get("INFRAI_API_KEY")
if not api_key:
raise RuntimeError("Missing INFRAI_API_KEY")
base_url = os.environ.get("SMS_API_BASE_URL")
if not base_url:
raise RuntimeError("Missing SMS_API_BASE_URL")
url = f"{base_url.rstrip('/')}/v1/sms/status/{message_id}"
for attempt in range(4):
request = Request(
url,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
try:
with urlopen(request, timeout=20) as response:
body = response.read().decode("utf-8")
if not 200 <= response.status < 300:
raise RuntimeError(
f"Status request returned {response.status}: {body}"
)
return json.loads(body)
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(
f"Status request returned {error.code}: {body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = (
int(retry_after)
if retry_after and retry_after.isdigit()
else 2**attempt
)
time.sleep(delay)
except URLError as error:
raise RuntimeError(f"Status request failed: {error.reason}") from error
raise RuntimeError("Status request exhausted its retry budget")
def main() -> None:
receipt = ReceiptDelivery(
order_id="order-1048",
country="US",
outcome=Outcome.FAILED,
resend_count=0,
version=7,
)
claimed = authorize_resend(receipt, 7, {"US", "DE"}, resend_limit=2)
print(claimed)
message_id = os.environ.get("SMS_MESSAGE_ID")
if not message_id:
raise RuntimeError("Missing SMS_MESSAGE_ID")
print(json.dumps(fetch_sms_status(message_id), indent=2))
if __name__ == "__main__":
main()
The limit of 2 in this example is an application policy, not a carrier promise. Make it configurable. More important, commit the transition and its incremented version before invoking the provider's resend capability. If two workers read version 7, only one conditional update should win; the other must reload the record rather than send again.
This is a small state space, which makes it excellent eval material. Run fixtures for all four outcomes, an unapproved country, an exhausted resend allowance, and two claims against the same version. Score the chosen action and the number of sends, not merely the final message text. Seven fixtures can expose a duplicate-delivery path that a notebook demo rarely reveals. I prefer that compact trace suite to dozens of happy-path payload snapshots because it tests the expensive decision: whether another notification leaves the system.
Short code. Strict policy.
The adapter then has only a few responsibilities. It submits the initial message with a stable idempotency key derived from the stored delivery operation, polls status or events with backoff and jitter, and invokes the resend operation only after the state transition is claimed. An HTTP 429 is a transport retry: honor Retry-After when present and use exponential backoff otherwise. It is not permission to create another business-level receipt. Infrai specifies idempotency as a platform convention on 171 of 294 capabilities, with a 24-hour default deduplication window; the application should still keep its own durable operation ID because its business record must outlive a transport window.
Delivery evidence changes the vendor choice
Provider selection should follow the operating model rather than a feature-count spreadsheet. For this workflow, ask how delivery evidence arrives, where sender registration is managed, and whether your team is prepared to own polling or callbacks.
| Product | Delivery evidence | Practical boundary for this workflow |
|---|---|---|
| Twilio Programmable Messaging | Message status plus status callbacks | Fits a callback-oriented stack; the application still validates and deduplicates callback-driven transitions. |
| Vonage SMS API | Delivery receipts | Offers a focused SMS path; normalize its receipt vocabulary before it reaches order state. |
| Amazon SNS SMS | Delivery status logging through Amazon CloudWatch Logs | Fits teams whose operational evidence already lives in AWS and CloudWatch. |
| Infrai | Pull-based SMS status and events, with resend and cancel operations | Fits when one REST contract should remain stable while the vendor behind the capability changes; the application must own polling and country controls. |
These are different commitments. Twilio's callbacks can reduce polling work but add an inbound event surface. Vonage keeps the integration SMS-specific. Amazon SNS ties diagnosis to CloudWatch and its surrounding AWS operations. Infrai keeps the application-facing contract fixed and exposes 295 routes across 20 modules through one key, which is useful when the same backend will later need other capabilities without accumulating SDK-specific boundaries. I would choose among them by ownership: a team already operating authenticated callbacks may value push delivery evidence, while a small Python service with a dependable worker can accept polling in exchange for a stable capability contract.
The trade-off is concrete. Infrai does not push email or SMS webhook events, so a near-real-time workflow must choose and operate its polling interval. It also does not provide geo-fencing or per-country spend cutoffs. Enforce the country allowlist, velocity limits, and spending circuit breaker before an adapter call. Those controls belong beside the resend authorization because both decide whether another message may leave the system.
Do not assume channels are interchangeable. SMS has resend and cancel operations, but cancellation only has meaning while work is pending; it cannot recall a delivered message. The email surface has no hosted OTP interface, and scheduled email does not have an equivalent cancel operation. There is no SMTP relay, voice, WhatsApp, or RCS channel in this capability set, so selecting any of those fallbacks means selecting and governing another integration.
Polling is an operations contract
Polling deserves an explicit service-level policy even though no measured latency or uptime figure is available here. Store the next poll time and attempt count, add jitter, and stop at an application deadline. A fleet that polls every open message on the same fixed second creates its own load spike and still cannot distinguish a slow queue from a rejection without reading the returned outcome. Suppose 1,000 settled payments arrive together: fixed-interval workers create 1,000 aligned checks, while persisted next-poll times plus jitter spread the work and survive a process restart. That number is an illustrative load case, not a measured platform limit, but it forces the queue design to be explicit.
Avoid logging receipt bodies, phone numbers, or verification secrets. Persist identifiers, timestamps, the normalized outcome, the original provider reason, the state-machine version, and the actor that authorized a resend. That record answers the operational question: did the carrier reject the sender, did the provider still report the message queued, or did two workers compete for the same transition?
One red counter is not enough.
An aging queued population asks for delivery-path investigation. A rise in carrier_rejected asks for sender registration, signature, carrier filtering, and market-routing review. A cluster of denied country checks means the fraud or routing policy is doing work. Combining them as “SMS failures” sends responders toward the wrong lever, especially when event notifications and login codes share an on-call dashboard.
There is also a reporting boundary to plan around: no cost report is aggregated by tag, and SMS templates do not have a list operation. Keep the order, market, and policy dimensions needed for internal reporting in your own receipt-delivery record rather than expecting the messaging surface to reconstruct them later.
The release gate for each destination market
Before enabling a country, confirm that the correct sender configuration and signature are registered and verified for that destination. Then run the state-machine fixtures, replay the same settled-payment event, and race two workers against one failed record. The expected result is one logical delivery operation and at most one successful resend claim.
Exercise a queued message that later becomes delivered, a reviewed failure that is resent, a carrier rejection that never enters the automatic resend path, and cancellation while an SMS is still pending. Verify the country allowlist and spending cutoff before any provider call. Finally, make the operations view show queued age separately from carrier rejection volume.
That is the production decision: ship the market only when registration evidence, transition tests, and business-layer controls agree. Vendor switching can remain an adapter change. Resend authority cannot.
Further reading
- Twilio, outbound message status: https://www.twilio.com/docs/messaging/guides/track-outbound-message-status
- Vonage, SMS delivery receipts: https://developer.vonage.com/en/messaging/sms/guides/delivery-receipts
- Amazon SNS, SMS delivery status in CloudWatch: https://docs.aws.amazon.com/sns/latest/dg/sms_stats_cloudwatch.html
- OWASP, Forgot Password Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Forgot_Password_Cheat_Sheet.html
- Yahoo, sender best practices and requirements: https://senders.yahooinc.com/best-practices/
Top comments (0)