Every developer has met a Heisenbug, the defect that vanishes the moment you attach a debugger. Its evil twin gets far less attention: the feature that only works while someone important is looking. A recent breakdown of the true cost of a staged demo traces that pattern through a gravity-powered truck, an edited AI video and a self-driving clip that resurfaced in a deposition, then closes with advice for the people sitting in the audience. This piece is for the people on the other side of the stage, the ones asked to make it look finished by Thursday. Because when engineers stage a demo, we rarely reach for a camera. We reach for a conditional.
The golden path had an expiry date
On January 9, 2007, Steve Jobs walked onto the Macworld stage holding a phone that, by any honest engineering standard, did not work yet. As Fred Vogelstein later reconstructed for The New York Times Magazine, the prototype could handle a snippet of a song or video but not reliably a whole clip, and it behaved if you sent an email before opening the browser, but not necessarily the other way round. Engineers spent hours hunting for a route through the software that wouldn't fall over, then scripted the keynote around it. They called that route the golden path. AT&T hauled a portable cell tower to the venue. And because the radio might crash and reboot somewhere in a 90-minute keynote, the team, with Jobs's blessing, made a call every developer will recognize. As Andy Grignon, who ran the iPhone's radios, put it: "we just hard-coded it to always show five bars."
The engineers watched from about the fifth row, each downing a shot of Scotch when the segment they owned survived. By the finale the flask was empty.
That was staging, plainly. So why does nobody call it fraud? Two reasons, and both matter to anyone who ships software. First, the gap had a deadline the team controlled: the iPhone went on sale on June 29, roughly six months later, and the people who faked the signal bars were the same people responsible for making them true. Second, the fakes concealed instability, not absence. The phone genuinely placed calls on stage; the hard-coded bars hid a radio that couldn't yet be trusted for an hour and a half. A forgivable staged demo is a loan with a repayment date. An unforgivable one is a capability that was never coming.
The demo that ran itself on 11 million cars
Now take the presenter out entirely. In September 2015, the U.S. Environmental Protection Agency revealed that Volkswagen diesels carried software able to recognize an emissions test and change how the engine behaved while the test was running. If you've never looked at how the trick worked, the BBC's plain-language explainer on the Volkswagen scandal is worth five minutes of your time. In engineering terms, it was a golden path with nobody holding the phone: on the test rig the emissions controls did their full job, while in normal driving the cars could emit up to 40 times the permitted level of nitrogen oxides. Volkswagen eventually admitted the software sat in about 11 million vehicles worldwide.
What finally exposed it wasn't a firmware teardown. It was a road trip. In 2013, researchers at West Virginia University took a $70,000 grant from the International Council on Clean Transportation, borrowed a Jetta, a Passat and a BMW X5, strapped portable emissions analyzers to them and drove real roads, including a run from San Diego to Seattle. Two of the Volkswagens came back far over U.S. limits. Nobody had to outsmart the code. They simply walked the demo off its golden path.
The bill matched the scale. By 2020, Volkswagen put its cost in fines and settlements at €31.3 billion. And the conditional had authors. In August 2017, a federal judge in Detroit sentenced James Liang, a VW engineer whom prosecutors called a "pivotal figure" in designing the cheating systems, to 40 months in prison and a $200,000 fine, ten times what the government had asked for. His lawyers argued he had followed orders out of loyalty to his employer. It didn't save him. When a demo has no presenter, the code becomes the presenter, and code keeps a commit history.
The leaderboard is the new keynote
For AI teams, the stage has moved to benchmarks. In April 2025, Meta's Llama 4 Maverick debuted in second place on LM Arena, the leaderboard where human raters vote between anonymous chatbot answers. Developers who downloaded the weights quickly noticed the public model didn't behave like the one on the board. Meta's own launch materials had disclosed that the Arena entry was an experimental chat variant, and as TechCrunch reported when the discrepancy surfaced, the leaderboard version was far chattier and heavier on emojis than anything developers could actually run. LM Arena published more than 2,000 of the head-to-head battles, said Meta hadn't been clear enough that its entry was a customized model, and rewrote its leaderboard rules. When the unmodified release was added to the board, it landed in 32nd place.
The interest kept compounding. In January 2026, Yann LeCun, by then on his way out of Meta, told the Financial Times that the Llama 4 results were "fudged a little bit," with different models used for different benchmarks. By his account, Mark Zuckerberg stopped trusting the people behind the release and sidelined the whole generative AI group. The leaderboard spot lasted days. The explanation took nine months.
Meta didn't hide a switch inside a model; it submitted a different model. But for developers deciding what to build on, the effect rhymed with Volkswagen: the thing that was measured was not the thing that shipped. An evaluation is a demo that runs while you sleep, and Goodhart's law applies in full. Once a score becomes the target, every team under pressure learns to perform for the scorer.
Where heisenfeatures hide in ordinary codebases
Most of us will never touch emissions firmware or a frontier model. The same shape still turns up in mundane code, usually committed with good intentions the night before a sales call:
-
Identity checks inside business logic. A branch like
if (account.isProspect)that routes the demo tenant to a warmer cache, a dedicated replica or precomputed results. The audience gets the fast path; paying customers get the other one. - Fixtures wearing a live badge. A dashboard that animates beautifully because it replays a JSON file, sitting under a label that says "real-time."
-
Time that lies. A
setTimeoutand a spinner standing in for a pipeline that doesn't exist yet, or the reverse: a cached answer returned instantly so an expensive model looks cheap. - Benchmark-shaped behavior. Prompts, retries or model routing that quietly change when an input matches a known eval format or dataset.
- A kinder environment. Demo builds with rate limits disabled, retries cranked up and error toasts swallowed, so the failures every real user sees simply never render.
None of these begins as fraud. Each becomes dangerous when it outlives the meeting it was written for, because six months later nobody remembers which numbers on the screen were real.
Building software that behaves the same with or without an audience
Treat observation invariance as a property you test, the way you test idempotency. Run your demo script against the production build with production configuration, then diff the outputs and latencies against what the demo environment shows. If the demo only works on a different code path, you haven't found a feature. You've found a liability with a nice UI.
When you genuinely need to simulate something, push the fake to the edges of the system. A mock service with "simulated" in its name and a visible banner in the interface is honest staging. A conditional buried in core logic that checks who is watching is a heisenfeature. Fixtures can be audited at a glance; branches hide in plain sight.
Give every fake an owner and an expiry date, the way the iPhone's golden path quietly expired on June 29. A demo flag that reaches production configuration should fail the build, and a demo-only marker that outlives its deadline should fail review. Then lead the audience off the path on purpose: let the prospect type their own input, publish your evaluation harness, and report scores only for the exact artifact you ship, with the same weights, the same config and the same hardware tier.
And remember the Detroit courtroom. The person who writes the conditional owns it, no matter who asked for it. If a commit message would sound bad read aloud by a prosecutor, the code probably shouldn't merge.
The audience that never sees the golden path
One question belongs above every demo you will ever give: what would this system do if nobody were watching? That is the only version your users ever meet. They don't follow the golden path. They open the browser before the email, play the whole clip, and use your product on one bar of real signal at two in the morning. Build for them, and the demo stops being a performance. It becomes a preview.
Top comments (0)