DEV Community

Zero Heartbeat
Zero Heartbeat

Posted on Originally published at delta1labs.com

What your app should do when the license server is down

Every online license check has a failure mode most teams don't design for: the server simply isn't reachable. Maybe your licensing service is having a bad day; maybe the customer's corporate proxy ate the request; maybe their laptop is on a train. If your check treats "I couldn't reach the server" the same as "the server said this license is invalid," then your outage becomes your customer's outage — and a thirty-second network blip locks a paying user out of software they own. The fix is to design explicitly for unreachability, with grace windows and a deliberate fail-open/fail-closed policy.

"No" and "I don't know" are different answers

The single most important distinction in resilient licensing is between two very different results of a check-in:

  • An authoritative negative — you reached the server and it said this license is expired, revoked, or over its seat count. That is a real "no," and it should take effect.
  • An inconclusive result — you could not reach the server at all: DNS failure, timeout, 503, TLS error. That is not a "no." It is "I don't know right now."

A naive check collapses both into "not valid → stop." That's the bug. The overwhelmingly likely cause of an inconclusive result is a transient network problem, not a customer who pirated your software mid-session. Treating silence as a denial punishes the honest majority to inconvenience an attacker who has far easier options anyway (they can just run offline). Resilient licensing keeps the two cases apart and responds to each appropriately.

The cached lease is what you fall back on

Resilience rests on a signed lease the app already holds from its last successful check-in — the same offline-validation primitive described in offline license validation. The lease is a server-signed statement, "this license is valid through time T," that the app can verify locally with an embedded public key, no network required. Its validity window is the backbone of grace: between check-ins the app trusts the lease, and when a check-in can't complete, it keeps trusting the lease until that window — plus an explicit grace period — runs out.

The key design parameter is making the lease window comfortably longer than the check-in interval. If you check in daily but issue a 7-day lease, a customer can be offline (or your server can be down) for most of a week before anything degrades. A check-in, then, is not a gate that must pass on every launch; it is an opportunistic refresh that extends the runway:

async Task<LicenseState> EvaluateAsync(CachedLease lease, DateTime now)
{
    try
    {
        var fresh = await _server.CheckInAsync(lease.Key);   // reach the server
        if (fresh.Status == Status.Revoked || fresh.Status == Status.Expired)
            return LicenseState.Denied(fresh.Status);        // authoritative NO — act on it
        _store.Save(fresh.Lease);                            // refresh: extend the runway
        return LicenseState.Active(fresh.Lease);
    }
    catch (Exception)        // timeout, DNS, 5xx, TLS — INCONCLUSIVE, not a denial
    {
        return EvaluateOffline(lease, now);                  // fall back to the cached lease
    }
}

LicenseState EvaluateOffline(CachedLease lease, DateTime now)
{
    if (now <= lease.ValidThroughUtc)
        return LicenseState.Active(lease);                   // still inside the lease window
    if (now <= lease.ValidThroughUtc + GracePeriod)
        return LicenseState.Grace(lease, until: lease.ValidThroughUtc + GracePeriod);
    return LicenseState.NeedsReconnect();                    // runway exhausted — degrade, don't crash
}
Enter fullscreen mode Exit fullscreen mode

Notice the asymmetry: a Revoked/Expired response ends the session, but any exception falls back to the lease. The network error never denies; only the server denies.

Fail-open, then degrade — rarely hard-lock

With the two cases separated, the policy writes itself for most products:

  1. Inside the lease window: run normally. Nothing to decide.
  2. Past the lease, inside grace: run normally, but start nudging — a quiet "we haven't been able to verify your license; we'll keep trying" banner. Retry in the background with exponential backoff (don't hammer a server that may already be struggling).
  3. Past grace: degrade, don't crash. Drop to read-only, disable paid-only features, or block new work while letting the user save and export what they have. A clear "reconnect to continue" message turns a dead end into a solvable problem.

The guiding principle is that your licensing should fail safe for the customer, not fail aggressive. The cost of wrongly locking out a paying customer during your own outage — support tickets, refunds, churn, reputation — dwarfs the cost of a determined pirate getting a few extra days of grace they could have gotten by staying offline anyway. Entitlement-gated products degrade especially gracefully: you can disable the premium flags while leaving the core app usable.

When fail-closed is the right call

Fail-open-within-grace is the right default, not a universal law. Two situations justify tighter behaviour:

  • An authoritative revocation. If the server is reachable and returns "revoked" — a refund, a chargeback, a compromised key — that is a real no and should take effect promptly, which is why revocation rides the check-in path and isn't subject to the grace window. Grace covers silence, never an explicit denial.
  • High-value or compliance-bound tiers. For a floating/concurrent model where a seat must be released, or a regulated deployment where unverified use is itself a problem, a shorter lease and a shorter grace are appropriate — the business model, not the network, sets the tolerance. Make this a per-tier policy: a consumer app can afford a generous week of grace; a concurrent-seat server tool might allow only hours.

The point is that fail-closed should be a deliberate choice for specific tiers, driven by the value at stake — not the accidental default you get from treating a timeout as a denial.

Tuning the knobs

Three parameters shape the whole experience, and they should be explicit, not emergent:

  • Lease validity window — how long a single successful check-in buys. Longer = more resilient to outages, slower to reflect a revocation. Days for most desktop/B2B software.
  • Grace period — extra time past the lease before degrading. A safety margin specifically for your outages; even a day meaningfully reduces false lockouts.
  • Check-in cadence + backoff — how often you refresh when healthy, and how you retry when failing. Refresh well before the lease expires (so a few missed check-ins are harmless), and back off on failure so a struggling server isn't stampeded.

Set these per tier and you get resilience by construction: brief outages are invisible, longer ones degrade gently with clear messaging, and only a real, authoritative "no" ever stops a customer cold.

The takeaway

A license check is not just a yes/no gate — it has a third state, "couldn't ask," and how you handle that state is what separates licensing that protects your revenue from licensing that occasionally destroys your goodwill. Cache a signed lease, give it a window longer than your check-in interval, add a grace period on top, and make the terminal state a graceful degrade with a path back. Treat the server's "no" as authoritative and the network's silence as "not yet." Do that, and the next time your licensing service has a bad hour, your customers won't even notice.

Top comments (0)