Try a small experiment with your favorite AI-powered app: revoke the API key, switch off the network, and open it again. A surprising number of products fail this test — not the AI features, the product. Spinners that never resolve. A settings page that errors out. A startup sequence blocked on a model ping. Somewhere between the demo and the release, the AI integration stopped being a feature and became load-bearing infrastructure.
WorldScript Studio is an open-source writing studio where AI assists with things like outlines, character work, and prose feedback. It is also, by design, fully usable with no key, no model, and no network. This article is about the mechanisms that keep it that way — a provider seam, policy gates in code, honest failure semantics, and a fallback layer that is allowed to say "no fallback exists." Code references are from the repository at commit 2d9157c0 (2026-09-28), release v1.28.8; simplified excerpts are labeled.
Optionality is not a settings toggle
Plenty of apps have an "AI: off" switch. Fewer have an architecture where off is a real, tested state. The difference shows up the first time a provider has an outage and your error handling turns out to be a toast notification saying "something went wrong" above a dead feature.
I ended up with a rule: the AI layer may fail in every way it wants, as long as the failure is typed, explained, and contained. Four mechanisms enforce it.
1. One seam, not fifty call sites
Every AI capability in the app — text generation, structured JSON, streaming, image generation — goes through one unified provider service. Providers sit behind it as adapters: Gemini, OpenAI, OpenRouter, Anthropic, and any OpenAI-compatible local server (Ollama, LM Studio) selected by base URL. A small factory normalizes them into a single LanguageModel type from the Vercel AI SDK:
// services/ai/providerFactory.ts (excerpt, comments trimmed)
export type WorldScriptLanguageModelConfig =
| { provider: 'gemini'; modelId: string; apiKey: string }
| { provider: 'openai'; modelId: string; apiKey: string;
headers?: Record<string, string> }
| { provider: 'openaiCompatible'; baseURL: string; apiKey: string;
modelId: string; headers?: Record<string, string> };
export function createLanguageModelForWorldScript(
config: WorldScriptLanguageModelConfig,
): LanguageModel { /* … */ }
The point of the seam is not abstraction for its own sake. It is that "the AI is down" has exactly one place to happen. Features never talk to a provider directly, so no feature can accidentally grow its own retry logic, its own key handling, or its own definition of what "offline" means.
Bring-your-own-key lives behind the same seam: keys are stored encrypted (in the browser build, under a random non-extractable AES-256-GCM key in IndexedDB), and no provider SDK ever sees storage.
2. Policy gates that throw, not conventions that hope
Routing is user-visible: four modes — hybrid, cloud, local, eco — with hybrid as the default. The part that matters architecturally is that the rules are enforced as code at the seam, not as conventions spread across components:
// services/ai/aiPolicy.ts (excerpt)
export function assertCloudAiAllowedSync(
provider: AIProvider,
privacy: PrivacySettings | undefined,
): void {
if (LOCAL_INFERENCE_PROVIDERS.has(provider)) return;
const mode = getActiveAiMode();
if (mode === 'local' || mode === 'eco') {
throw new Error(`Cloud provider blocked: AI mode is "${mode}" (local-only).`);
}
if (!privacy) return;
if (privacy.localStorageOnly) {
throw new Error('Cloud provider blocked: local-only mode is active.');
}
// …
}
A gate that throws is testable in a way a gate that "should be checked" is not — the policy test suite exercises the mode matrix directly. The same file carries a second hard gate: model training (LoRA fine-tuning) is restricted to local providers by allowlist, because training data is manuscript data, and manuscript data does not leave the device for that path. Note the scope of that sentence: it describes the training path, not a blanket privacy guarantee for every AI feature — a cloud provider you explicitly call necessarily receives the context you send it.
3. Failure semantics: retry what is retryable, explain what is doomed
Provider errors are not one thing. A rate limit is not a wrong API key, and neither resembles "the laptop is offline." The seam classifies every failure into a small taxonomy:
| Category | Retryable? | Why |
|---|---|---|
| transient, rate limit, network | yes | connection-class; a later attempt can succeed |
| auth, policy, invalid request | no | deterministic — retrying repeats the failure |
| offline | no | doomed until connectivity returns |
| canceled, permanent | no | user intent / unrecoverable |
Each class carries a stable message key, so the UI can say "check your key" or "you are offline" instead of showing a generic error — and so the retry layer fails fast on doomed calls instead of backing off politely on a request that will never succeed. This is the difference between a degraded app and a lying app.
4. A fallback layer that is allowed to say no
When an AI call is terminally unavailable, some features can fall back to local heuristic generators — registered per task (an outline generator, a character-profile generator). The design decision I care about most is what the registry does when nothing is registered:
runHeuristicFallback(task, ctx)
→ generator registered? run it, return its result
→ nothing registered? return null
caller keeps its existing behavior
The fallback layer is "always safe to ship empty," as the code comment puts it. null means: no fallback exists, tell the user the feature needs a configured provider. What the registry never does is invent a plausible-looking answer to cover for a missing model. A fallback that quietly fabricates quality is worse than an honest refusal, because users cannot calibrate trust against output they cannot distinguish from the real thing.
feature call
│
▼
unified provider seam ──► policy gate (mode, privacy) ──throws──► typed error
│ │
▼ ▼
provider adapter (cloud/local) UI hint via message key
│ │
├─ terminal failure ──► heuristic fallback ──null──► honest refusal
▼
streaming response
Simplified: the real chain includes cancellation, request deduplication, and a provider fallback chain for transient cloud failures.
What "offline" honestly means here
Precision matters more than marketing, so: with no network, the manuscript editor, planning tools, and storage work fully — they never touch the seam. Browser-local and Ollama-served models keep working if they were set up beforehand; the first model download obviously needs a connection, and local inference needs hardware that can carry it. Cloud providers are unreachable, and the offline error class says so. Hybrid mode is deliberately cloud-first; local mode is the setting that guarantees the device never emits a request.
AI being optional also means the app's identity does not collapse without it. There is no onboarding step that demands a key, no feature gate on the editor, no telemetry about your text leaving for a server you did not choose.
The checklist I would hand another team
- One seam. All AI traffic through a single layer; providers as adapters behind it. If a feature imports a provider SDK directly, optionality is already gone.
- Gates that throw. Mode and privacy policy enforced in code at the seam, with tests — not documented conventions.
- Typed failure. Classify errors by whether a retry can succeed; map every class to a user-actionable message.
- Refusal-capable fallback. Degrade to simpler local behavior where it genuinely helps; return "no fallback" where it would only pretend.
Then run the unplug test in CI or by hand: no key, no network, cold start. Whatever still works is your product. Whatever breaks that should not have is your real dependency graph.
Source note: WorldScript Studio is open source (github.com/qnbs/WorldScript-Studio). Code references correspond to main at 2d9157c0 (2026-09-28); release anchor v1.28.8. Key files: services/aiProviderService.ts, services/ai/providerFactory.ts, services/ai/aiPolicy.ts, services/ai/aiErrorTaxonomy.ts, services/ai/heuristicFallback/registry.ts. Part of the "Engineering WorldScript Studio" series.
Practical extension: model capability state, not a boolean
“AI available” hides important states. A product can distinguish:
| State | User-facing behavior |
|---|---|
| Not configured | explain the missing setup; core editing remains usable |
| Ready | offer the action and name the provider boundary |
| Temporarily unavailable | preserve the input and offer retry or local work |
| Permission denied | explain the authorization needed without retrying secretly |
| Request in progress | allow cancellation and prevent stale responses from replacing newer work |
| Partial response | label it as incomplete and let the user keep or discard it |
Keep provider imports and network setup behind an optional module so a missing provider cannot break application startup. A fallback is valid only when its behavior is understood; silently switching from a remote model to a local one may change privacy, cost, quality, or data residency. “Offline” should describe a tested user flow, not merely a setting that says local.
Measure optionality with a negative test: start the application with provider packages absent or configuration invalid, then verify that ordinary editing, save, export, and recovery still work. Record what was and was not exercised. For agent-backed implementations, the OpenAI evaluation guide is useful for separating task success from the mere fact that a model returned text.
AI disclosure: AI tools helped with repository research, structure, and editing of this article. I reviewed the technical claims against the referenced source files before publication.
Top comments (3)
The choice to let the fallback registry return null is the detail I keep thinking about. It’s tempting to make every failure produce something, but a plausible heuristic answer can be worse than an honest “this needs a configured provider”—especially in a writing tool where users may not know which outputs came from a model and which didn’t.
I also like that the offline guarantee is scoped carefully: the editor and planning tools work offline, but cloud inference doesn’t magically become private or available. That kind of precision makes the architecture easier to trust. How do you keep the “one seam” rule from eroding as new features get added—do you enforce it with tests or another kind of guardrail?
@carbonlayer Mostly tests plus fail-closed routing today — but I'd separate the execution seam from the dependency seam.
For the main application generation paths, requests go through the shared provider service, where provider selection, policy, fallback behavior, and refusal semantics are tested. The provider factory itself stays inside the AI layer, and unsupported mappings fail closed rather than silently acquiring a plausible integration. The null cases you called out are pinned too: no registered heuristic, a declined fallback, or a failed fallback can all remain an honest
null.What those tests do not prove is that provider-shaped dependencies can never leak outward. There is a concrete example on current main: one plot-board thunk imports
Typefrom@google/genaito build a response schema, while the actual generation request still goes through the shared service. So it is not a direct-provider or policy bypass — but it is exactly the kind of dependency edge a structural guard would make visible.There is no dedicated static AI import-boundary CI check today. The repository already uses that kind of zero-tolerance gate for its desktop/Tauri boundary, so an AI analog is the logical hardening direction.
Tests prove behavior; a structural boundary proves the topology does not quietly erode.
That distinction is useful: the shared service can protect the execution path without proving the dependency boundary is holding. The plot-board example makes that concrete. It isn’t routing around policy, but it still lets a provider-specific dependency reach into application code—and those edges can accumulate quietly while every behavior test stays green.
A structural check seems like the right complement, especially if it catches type-only imports too; those can create real coupling even when they don’t make a provider call. I’d be curious how you’d scope the rule: prohibit provider SDK imports outside the AI layer outright, or allow narrowly defined exceptions for things like schemas?
Some comments have been hidden by the post's author - find out more