DEV Community

VLAD
VLAD

Posted on

Knight Capital: How One Forgotten Server Lost $440 Million in 45 Minutes

On August 1st, 2012, Knight Capital lost more than $400 million in about 45 minutes. It wasn't a hack or a bad bet. It was old code that nobody remembered was still running on one of eight servers.

Everything below comes from the official record: the SEC order against Knight Capital (Release No. 34-70694, October 2013). It reads like a postmortem, and most of its lessons apply far outside finance.

The machine: parent orders, child orders, and a counter

Knight handled about one in ten trades in U.S. listed stocks in 2011–2012. One of its core systems was SMARS, an automated router. It took a big "parent" order, split it into smaller "child" orders, and sent those to exchanges.

The key part was a counter. SMARS tracked how many shares of the parent order had already been filled, and stopped sending children when the order was complete. Remember that counter.

The ghost in the code

Inside SMARS lived an old feature called Power Peg. Knight stopped using it in 2003, but the code was never deleted. It stayed on the servers, still callable.

In 2005, the counter that told Power Peg when to stop was moved to a different place in the code. Power Peg was never retested after that. Nobody checked whether it still knew when to stop. It didn't.

In 2012, the NYSE launched a new retail program starting August 1st, and Knight wrote new code for it. The new code needed a switch, a flag on each order that meant "use the new feature". Instead of adding a new flag, it reused the old flag that used to activate Power Peg. The plan was to replace the old code everywhere, so the flag would simply mean something new.

The deploy: seven out of eight

SMARS ran on eight servers. Starting July 27th, a technician copied the new code to them by hand over several days. Seven servers got it. The eighth didn't, so it still had Power Peg.

There was no second person reviewing the deployment, and no written procedure requiring one. Other teams at Knight had written procedures; the SMARS team didn't.

97 warnings nobody read

At 8:01 a.m., before the market opened, a Knight system started emailing a group of employees about SMARS orders with the error "Power Peg disabled". By 9:30 it had sent 97 of them. They weren't designed as alerts, and nobody generally read them. They named the exact problem.

9:30: 212 orders become millions

When the market opened, orders arrived carrying the reused flag. Seven servers ran the new code. The eighth ran Power Peg, which kept sending child orders because its counter wasn't where it expected. Another part of Knight's system knew the orders were filled, but that information never reached SMARS.

212 customer orders turned into millions of child orders: about 4 million executions in 154 stocks, over 397 million shares.

People inside saw positions piling up in one account with a $2 million limit. But the risk tool only displayed numbers to humans. It wasn't connected to the order system, it didn't raise automated alerts, and nothing cut the machine off.

The fix that made it worse

The engineers reasoned that the trouble started with the new code, so they removed it from the seven servers where it was working. The orders still carried the reused flag. Now all eight servers woke up Power Peg.

The bill

When it stopped, Knight held about $3.5 billion of stock it never wanted to buy and had sold about $3.15 billion it didn't own. Knight first estimated the loss at about $440 million; the SEC later put it at over $460 million. In 37 stocks, prices moved more than 10% with Knight doing most of the trading.

Five days later, investors put in $400 million to keep the firm alive. Less than a year later it merged with a rival. The SEC fined it $12 million for not having the controls the market access rule requires.

What a developer should take from this

None of the individual decisions looked dangerous on the day they were made. That's the point.

  1. Delete dead code. Code that's switched off is still code that can be switched on.
  2. One flag, one meaning. Give every new feature its own flag, and plan the day you'll remove it. Reusing a flag makes old code reachable in ways nobody is thinking about.
  3. Verify every deploy. Automate it, and make every instance report which version it runs. The eighth server can't hide from a script.
  4. Build the kill switch before you need it. If your code can spend money or send messages at machine speed, stopping it must not depend on first understanding the bug.
  5. An alert nobody reads is not an alert. If a message can say "Power Peg disabled" 97 times without anyone noticing, it's noise, and noise hides the one message that matters.

What's the oldest dead code you've ever found in production?

I make Vlad's Stack — how the tools you use every day actually work, for people who write code: https://www.youtube.com/@VladsStack

Top comments (4)

Collapse
 
dsiacci profile image
Dominique Siacci •

The reused flag is the part that maps onto what we deal with every day. On our platform, which builds apps for other people, every app is a description read by an engine they all share, and the binaries reading those descriptions ship whenever each owner decides to rebuild, some of them years ago. So we live with a permanent eighth server, lots of them in fact, and the rule we ended up with is that a config key can be added but never renamed, retyped or given a new meaning, because an old binary keeps reading it the old way for as long as it's installed. The other half sits on the server: a client says which generation it is, and new capabilities only get served to the apps that can speak them. Your point 2 is the one I'd underline, with one change for systems like ours, which is that you can't always plan the day you remove the flag when you don't control when the last old reader goes away.

Collapse
 
vladut02 profile image
VLAD •

This is a great addition, thanks. You're right that "plan the day you remove it" assumes you control every reader. In your case the eighth server never goes away, so the only safe rule is the one you landed on: keys are add-only, and a meaning never changes once it ships.

It's the same idea as Protobuf's rule for field numbers: you never reuse one, you mark it reserved, because some old client out there will still read it the old way. And the other half you describe, the client saying which generation it is so the server only offers what it can handle, is what would have saved Knight: the eighth server would have announced "I'm still the old version" instead of silently doing something else with the flag.

Curious how you handle the other end: do you ever get to retire a capability, for example once telemetry shows no clients of that generation are left, or does it stay forever?

Collapse
 
dsiacci profile image
Dominique Siacci •

We don't wait for a zero. We don't see the fleet the way Knight saw eight servers: an app we published years ago can still be installed, and the owner may never rebuild, so "no clients of that generation left" isn't a number we get to trust. A key that shipped stays readable. What we retire is the form we write: new code stops using the old encoding, the historical one keeps working next to it, and we don't even wrap it in a warning, because a warning is a request to someone we don't control. The only place we get to say no is before the word exists. If the compatibility burden would outlive the usefulness, we don't add it.

Thread Thread
 
vladut02 profile image
VLAD •

"The only place we get to say no is before the word exists" is the best one-line summary of this whole problem I've read. It's also the exact opposite of Knight's mistake: they treated Power Peg as retired in 2003, but for the code the word still existed and still had a reader. You start from the assumption that the old reader lives forever, which is the only honest assumption when you don't control the fleet.

The point about warnings is a good one too. A deprecation warning only works if someone on the other end can act on it.