I run a small, deliberately boring trading bot: daily bars, spot only, no leverage, about ¥24,000 of my own money on a Japanese exchange (GMO Coin). It wakes up once a day at 06:05, looks at yesterday's candle, and usually does nothing.
On day six it halted itself and reported a 58% drawdown. My balance had not moved by a single yen.
(The free lite backtester for this system, public API only, no key: GMO Coin Trend Lab, source on GitHub.)
Nothing was lost. No order was sent. But the bug is a good one, because every link in the chain looked reasonable on its own.
The timeline
- The PC had been powered off for three days. When it came back, Windows had not yet resynced the clock. It was running about 16 seconds fast.
- The scheduled task fired. The bot signed its first private API call with the local timestamp.
- The exchange rejected it:
ERR-5009 Timestamp for this request is too fast. - My startup code treated "key validation failed" as "no usable key" and fell back to dry-run mode. I had written that fallback on day one so that pasting the wrong kind of key would not crash anything.
- Dry-run mode has its own paper balance, ¥10,000, and it shared the same state file as live mode.
- The bot computed equity as ¥10,000, compared it with the stored high-water mark of ¥23,900, got a 58% drawdown, and did exactly what it is supposed to do at 30%: flatten and halt.
Because it was in dry-run, "flatten" sent nothing. The halt flag, however, was written to the real state file. The bot would have sat there, halted, until a human noticed.
Why each decision looked fine in isolation
"Fall back to dry-run if the key does not validate." Friendly on day one, when the likely failure is a user pasting the FX key into the crypto slot. Dangerous on day six, when the key is fine and the failure is transient.
"One state file." Simple. Until two modes with different balances write the same high-water mark.
"Halt at 30% drawdown." Correct, and it worked. The kill switch is not the bug. The bug is that a number from one world was compared against a number from another.
Fix 1: sign with the server's clock, not mine
The exchange checks that your timestamp is close to its own. So ask it what time it is. Every HTTP response carries a Date header, and that is accurate enough for a tolerance measured in seconds.
import email.utils, time, requests
PUBLIC = "https://api.coin.z.com/public"
class GmoClient:
def __init__(self):
self.s = requests.Session()
self._offset = None # local clock minus server clock, in seconds
def clock_offset(self, refresh=False):
if self._offset is None or refresh:
try:
t0 = time.time()
r = self.s.get(PUBLIC + "/v1/status", timeout=15)
t1 = time.time()
srv = email.utils.parsedate_to_datetime(r.headers["Date"]).timestamp()
self._offset = (t0 + t1) / 2 - srv
except Exception:
self._offset = 0.0
return self._offset
def _timestamp_ms(self):
# the Date header has one-second resolution, so lean half a second early:
# "slightly slow" is tolerated, "too fast" is rejected
return str(int((time.time() - self.clock_offset() - 0.5) * 1000))
Two details matter. Take the midpoint of the request so latency does not bias the estimate. And lean slightly early, because this API rejects timestamps from the future more strictly than ones from the recent past.
After this change the bot does not care whether the OS clock is right.
Fix 2: a failed validation must not touch live state
c = GmoClient()
if not c.dry_run: # keys are present, so we intend to be live
try:
c.assets()
except GmoError as e:
log(f"key validation failed, aborting this run without touching state "
f"(clock offset {c.clock_offset():+.1f}s): {e}")
return # try again tomorrow
if c.dry_run:
STATE = ROOT / "state_dry.json" # paper trading lives in its own file
The rule I took away: if the system intended to be live and cannot prove it is live, it should do nothing, loudly. Not "degrade gracefully" into a different mode that shares storage with the real one.
What I check now when writing a fallback
- What does the fallback write, and where? If it shares storage with the normal path, it is not a fallback, it is a second writer.
- Is the failure I am catching permanent (wrong key) or transient (clock, network, maintenance window)? Transient failures should abort and retry later, not change mode.
- Would I notice? The halt was silent until I went looking. The bot now writes a dashboard every ten minutes with a red banner when it is halted.
- Can the safety mechanism be fed garbage? A kill switch is only as good as the equity number going into it.
The boring numbers, since people ask
The strategy itself is a 20-day breakout with a 50-day filter and a 2×ATR stop, risking 1% per trade. My backtest on the exchange's own daily data from 2018 to 2026 gives about the same return as volatility-targeted buy-and-hold with roughly half the drawdown, and close to zero in ranging years. Walk-forward out-of-sample Sharpe is 0.85. Its first real trade happened on day eight: 0.0076 ETH, with about ¥240 at risk.
Operational holes scare me more than strategy losses. This one cost nothing, which is the best price to learn at.
If you want to poke at the backtest, the lite version is a free download (public API only, no key needed): https://wataflow1.gumroad.com/l/trend-lab-free
Not investment advice.
Top comments (7)
"A number from one world was compared against a number from another" is the postmortem in one line, and you got there yourself, so let me push on the fix rather than the diagnosis. @_firelinks already took the taxonomy half — permanent and transient errors sharing a fallback path — so I will take the direction instead.
Your day-one rule was: if the key does not validate, degrade to dry-run and carry on. That is a fail-open dressed as a safety feature. The system responds to "I cannot authenticate" by continuing in a different mode rather than stopping. In payments that specific move is the one you are not allowed to make, because an auth failure means you no longer know what state you are in, and every mode you could fall into has its own numbers.
Worth noting the clock did not have to be wrong for this to fire. Any transient auth rejection does it: a rate limit, a brief API outage, a rotated key still propagating. The 16 seconds made a good story, and the bug was waiting for a much more ordinary trigger.
Alongside splitting the state file, I would make the mode part of the comparison rather than just part of the filename. If the high-water mark carries the mode that wrote it, then dry-run equity against a live high-water mark is a type error instead of a 58% drawdown. Separate files stop this collision. Tagging the number stops the next one, where the two worlds are something other than live and paper.
One thing I would watch on the Date-header fix: you have moved a correctness control onto an unauthenticated response header with one-second resolution. That is fine against drift, which is your actual failure. It is worth being clear that it is a tolerance check rather than a time source.
That distinction is the useful one: a mode change is an external effect, not a harmless fallback. I would treat authentication uncertainty as a state that blocks any action whose effect depends on identity, and expose dry-run only when the contract explicitly permits it. Otherwise the system changes the meaning of "continue" while preserving the same status surface. Origin scoping and an explicit mode in the evidence record make that boundary inspectable; TTL can come later. Thanks for pushing the point.
The angle I'd add to yours: the kill switch itself did its job perfectly and that's almost the most unsettling part. It saw a 58% drawdown and halted, exactly as designed — the failure was entirely upstream, in what got fed to it. A safety mechanism that fires confidently on bad input is easy to mistake for the system working, because from the outside "it halted" and "it halted correctly" look identical.
Tagging the number with its mode fixes the type error you're describing. I'd also sanity-check the input before it ever reaches the kill switch: a single-day 58% move on a spot-only, no-leverage daily-bar strategy is implausible on its face, independent of which mode produced it. A magnitude check — "does this even happen in real markets under this strategy" — catches the case your fix doesn't: a bug that produces a bad number of the right type, tagged correctly, that's still nonsense. The kill switch trusted its input completely, which is correct until the input is the thing that's wrong.
The 16-second drift is a great catch, and it is worth saying out loud why this class of bug is so expensive: every timestamp in the system agrees with every other timestamp, and they are all wrong together. Nothing looks inconsistent, so nothing looks wrong.
I hit the same family from the other direction in a trading guard. The rule was "reset the daily loss counter when the trading day rolls over". It used the host clock. Every test passed, every log line was coherent - and on a VPS in a different timezone than the broker, the counter reset four hours early. The failure mode was not a spike, it was a quietly larger daily risk allowance than the one I had configured, for weeks.
Three things that caught it / would have caught it faster:
TimeTradeServer()rather thanTimeLocal(), and the difference is not cosmetic.The nastiest variant is the one where the drift is asymmetric: fine on reconnect, wrong on the first tick after a long disconnect. That is where the "is the feed stale?" question needs an explicit timestamp comparison rather than "did we get a tick".
Thanks for writing this up with the actual numbers - a lot of people will recognise their own incident in it.
This is a really important lesson about time synchronization! I've seen similar issues when building APIs that need to handle rate limiting and caching. The devil is definitely in the details with timing. What monitoring tools do you use to catch these issues early?
Clean postmortem. The chain is elegant because each step looked reasonable on its own. The clock drift fix solves the immediate problem, but the architecture underneath is the bigger concern. The bot holds the signing key in the same process that can be fed wrong data by a 16-second clock skew. When the environment can confuse the signer, any glitch becomes a potential catastrophe. Separating proposal from authorization removes that coupling. The agent proposes the action, and the signing happens outside its execution environment, where clock drift, network flakiness, and fallback logic cannot touch the authorizer. Nothing the agent experiences affects whether or how the action gets signed. The author's own rule about doing nothing loudly when live status cannot be verified reaches the same conclusion.
The error taxonomy is the point I’d carry into other systems: a permanent credential or configuration error and a transient clock or network failure should not share a fallback path. Once they do, a safe-looking retry can become a mode change that writes into the wrong state. I’d make the mode, clock offset, validation outcome, and state-file identity part of the run telemetry, then alert on any run that claims live intent without live proof. The quiet halt is exactly the kind of control that needs an observable failure.