DEV Community

TheophilusHawkins9265
TheophilusHawkins9265

Posted on

Reset Email Operations — Auditing Token Expiry, Single Use, and Database Evidence

To build a secure password reset email flow, make every committed state transition provable before a report deadline creates an emergency. The page fires when requests stop reaching a terminal state. On-call sees 31 open recoveries for B2B SaaS administrators who are trying to download generated compliance reports, the oldest already 18 minutes old, but no raw email addresses and no reset links in the alert.

TL;DR: build the flow as an auditable sequence of committed transitions: accept an indistinguishable public request, create a short-lived random secret, store only its SHA-256 digest, hand the email to the transport, and redeem the secret exactly once in the same transaction that changes the password. Correlate those steps with an internal recovery ID. This is the least complex design that gives an investigator useful evidence without turning the evidence store into a cache of bearer credentials.

The first useful alert is therefore not "email failed." It is "recovery records are aging between two known states." Work backward from that page, and both the security design and the instrumentation become much harder to fake.

How should a secure password reset email flow prove single use?

A user-visible redemption error is late. By then, an administrator has requested recovery, waited, opened a message, and followed its link. For a report due before a review meeting, support may learn about the break before operations does.

Model four durable facts: the request was accepted, a recovery record was committed, a message handoff was acknowledged, and a password change was committed. The public request response must not disclose whether the address belongs to an account. Internally, the outcome can be recorded under restricted access so abuse and failures remain distinguishable to authorized responders.

Those facts are deliberately narrower than "the email arrived." An application can prove that its configured transport accepted a message; without separate delivery evidence, it cannot prove inbox placement. The distinction matters during a compliance review because an event trail should support the claim being made, not a more convenient claim.

Claims need receipts.

Use one internal recovery ID across the sequence. Do not use the email address, token, or reset URL as the correlation key. The identifier lets the responder follow a single attempt while the secret stays out of dashboards, traces, and tickets.

The earlier signal is an age-based gap: committed recovery records with no acknowledged handoff after the service's documented operating window. A second signal watches records that were handed off but never became either redeemed or expired. The number 18 minutes above is example alert data, not a universal threshold; each team has to derive its window from its own queue behavior, token lifetime, and response obligation.

Small denominators bite. A ratio based on one request can page at 100% failure, while a high-volume interval can hide a meaningful absolute backlog behind a calm percentage. Pair age with a minimum count, and send low-confidence deviations to investigation rather than the pager.

Put every committed claim under one owner

Start the implementation review by assigning each claim to an owner. Security owns the single-use rule, the identity service owns the commit, messaging owns the handoff fact, and the report service owns authorization after recovery. Single use is decided under concurrency, not by an if statement in application memory. Two requests can read an unused record before either marks it used. The password update and token consumption need one transaction and one winner.

Generate 32 random bytes with the operating system's cryptographic source, encode them for the URL, and email the plain value. Persist only a SHA-256 digest with the account reference, creation time, absolute expiration, and consumption time. Hashing limits what a database read exposes, but it does not improve a predictable token; the random input still carries the security property.

This Go sketch places the audit boundary after commit. Its SQL uses PostgreSQL placeholders and row locking, so a different database needs equivalent conditional-write semantics rather than a mechanical copy.

package recovery

import (
    "context"
    "crypto/rand"
    "crypto/sha256"
    "database/sql"
    "encoding/base64"
    "errors"
    "time"
)

var ErrInvalidReset = errors.New("reset link is invalid or expired")

func NewToken() (string, [32]byte, error) {
    b := make([]byte, 32)
    if _, err := rand.Read(b); err != nil {
        return "", [32]byte{}, err
    }
    plain := base64.RawURLEncoding.EncodeToString(b)
    return plain, sha256.Sum256([]byte(plain)), nil
}

type AuditEvent struct {
    RecoveryID string
    Stage      string
    OccurredAt time.Time
}

func Redeem(
    ctx context.Context,
    db *sql.DB,
    plain string,
    now time.Time,
    setPassword func(context.Context, *sql.Tx, int64) error,
    emit func(context.Context, AuditEvent) error,
) error {
    tx, err := db.BeginTx(ctx, nil)
    if err != nil {
        return err
    }
    defer tx.Rollback()

    digest := sha256.Sum256([]byte(plain))
    var resetID, userID int64
    var recoveryID string
    err = tx.QueryRowContext(ctx, `
        SELECT id, user_id, recovery_id
        FROM password_resets
        WHERE token_hash = $1
          AND used_at IS NULL
          AND expires_at > $2
        FOR UPDATE`, digest[:], now).Scan(&resetID, &userID, &recoveryID)
    if errors.Is(err, sql.ErrNoRows) {
        return ErrInvalidReset
    }
    if err != nil {
        return err
    }

    if err := setPassword(ctx, tx, userID); err != nil {
        return err
    }
    result, err := tx.ExecContext(ctx, `
        UPDATE password_resets
        SET used_at = $1
        WHERE id = $2 AND used_at IS NULL`, now, resetID)
    if err != nil {
        return err
    }
    rows, err := result.RowsAffected()
    if err != nil || rows != 1 {
        return ErrInvalidReset
    }
    if err := tx.Commit(); err != nil {
        return err
    }
    return emit(ctx, AuditEvent{
        RecoveryID: recoveryID,
        Stage:      "credential_changed",
        OccurredAt: now,
    })
}
Enter fullscreen mode Exit fullscreen mode

There is a sharp edge in that compact example: if the process exits after the transaction commits but before emit succeeds, the credential changed without the external event. For compliance evidence, use a transactional outbox row written inside the same transaction, then publish it asynchronously. The database record is the source of truth; log emission alone isn't.

This relational design is not suitable when the password store cannot participate in the same transaction as token consumption. Splitting those writes across independent stores creates an ambiguous partial-success case, and adding retries does not restore atomicity. In that environment, put both decisions behind a service that owns a transactional store or use a database with conditional writes that can enforce one winner. There is another trade-off: row locking is easy to inspect for an ordinary regional B2B SaaS workload, but a globally active service may reject the added coordination latency and choose a globally coordinated conditional write instead. That choice costs operational complexity. The invariant stays the same even when the storage mechanism changes: one request may move the record from live to consumed, and the password change must share its commit boundary.

Never log the plain token.

Return the same generic failure for an unknown, expired, or consumed value. Keep the precise reason in a restricted event category. This preserves a useful investigation trail without giving the caller an account-discovery oracle.

Keep report authorization outside account recovery

The report attachment keeps the design honest. Recovery restores control of an account, but it does not prove that the recovered account may access a particular tenant's report. Re-run tenant membership and report authorization after the password change. Do the same at attachment download time if the email links to a download rather than carrying the document itself.

An evidence map is more useful than a pile of logs:

Review question Evidence that answers it Data to keep out
Did the system accept a recovery attempt? Recovery ID, tenant ID, time, normalized internal outcome Plain email address in general logs
Was a bounded secret created? Digest, creation time, absolute expiration Plain token and full reset URL
Did the message cross the application boundary? Transport handoff ID and acknowledged time Unsupported claim of inbox placement
Did exactly one attempt change the credential? Committed transaction result and consumption time Password or password hash
Was report access authorized afterward? Tenant, report ID, authorization decision, time Report contents in the recovery trail

Access to this trail should be narrower than access to routine service logs. Retention needs a written owner and deletion behavior that is tested. Keeping recovery evidence indefinitely expands exposure; deleting it before the applicable review window makes the control difficult to demonstrate. No universal duration follows from the supplied standards, so legal, security, and support owners must set it for their obligations.

Mail authentication is a separate layer. Google's sender guidelines require SPF or DKIM for all senders to personal Gmail accounts and describe additional SPF, DKIM, and DMARC requirements for bulk senders. They also require TLS for transmitting email. These controls help receiving systems evaluate the sender; they do not make a reset token single-use and they do not replace the application's authorization decision.

NIST SP 800-63B requires verifiers to implement rate limiting for failed authentication attempts. Recovery endpoints deserve the same operational scrutiny around repeated requests and guesses, but an exact limit must follow the application's threat model and traffic. Record the policy version with the decision so a reviewer can tell which control was active at the time.

Audit the control without replaying its secrets

Deployment testing should exercise state transitions, not only the happy-path email. Submit a request for a known account and an unknown address, expire a token, race two redemptions, reject a transport handoff, and roll back the password transaction. The public responses must avoid account disclosure. Internally, one racing redemption commits, the other fails generically, and no evidence claims credential_changed before the database commit.

Then remove the mail worker from service in a test environment. The age-based backlog should rise, the page should link to a runbook, and the runbook should tell the responder how to distinguish record creation failure from handoff delay. If the only instruction is "check the email provider," the signal lacks enough state.

This rehearsal is also the release gate. A change that cannot reproduce the stalled transition and its evidence should not ship merely because a test message appeared in an inbox. Inbox receipt exercises one path; the control exists for the paths where a write, queue, handoff, or concurrent redemption fails.

Escalate by evidence confidence, not raw error count

Thresholds spend attention. Set the age window too tightly and ordinary queue variation repeatedly wakes someone; set it too loosely and the report deadline passes before the page. Revisit the threshold with observed traffic and queue latency, and document why the chosen value leaves time for a useful response. I favor paging only on a sustained, actionable state gap; a weak signal belongs in a dashboard or ticket.

That is the operational finish line: the runtime can demonstrate a bounded secret, one committed redemption, authenticated mail handling, and a separate authorization decision for the generated report. The alert is valuable because it identifies which transition stopped. Its false-positive cost remains part of the design, because a page that responders learn to distrust is another missing control.

Further reading

Top comments (0)