I looked at what my monitoring product had actually been sending people.
78% of reports showed a delta of exactly zero. Nothing had changed since the previous run. The report was correct, on time, and empty.
My first reading was that my change-detection thresholds were too coarse — that real movement was happening below the granularity I was measuring, and a finer scale would surface it. I spent a while on that. I wrote about the alert side of it earlier in this series, and about the backtest that talked me out of the tuning; this is the part I got wrong underneath it, which is different and worse.
The thresholds weren't the problem. The artifact was the problem. A report whose content is "no change" is not a report with a sensitivity issue. It's a report with nothing in it.
Why "no change" is the expected state
Once you say it out loud it's almost embarrassing.
Someone signs up, runs an audit, gets a list of problems, fixes some of them, and stops. That's the success path. Their site is now static in the dimensions I measure, because sites are static in those dimensions. robots.txt doesn't drift. Schema markup doesn't decay. A well-formed llms.txt stays well-formed until somebody edits it, and nobody edits it.
So the population my weekly report addresses is mostly people for whom the honest answer is "nothing happened, and that's fine." Which is not a message worth an email, and definitely not worth a dashboard visit.
The uncomfortable implication: a monitoring product for a slow-moving signal has no natural weekly cadence. I had assumed one because monitoring products have weekly reports. That's an industry convention, not a property of my data.
What I'd actually measure now
The metric I was missing isn't a threshold. It's a property of the output:
What fraction of generated artifacts contain at least one thing the recipient didn't already know?
Call it content rate. Mine was 22%. Every product decision downstream looks different once you know that number, and none of them are threshold changes:
Change the cadence to match the signal. Not weekly because it's Tuesday. Send when there's something, and say so — "we watched, nothing moved" is a fine thing to put in a monthly digest and a bad thing to put in a weekly email.
Add signals that actually move. The things I was measuring are properties of a site the owner controls, so they change only when the owner acts. Competitor position, whether a citation appeared or disappeared, a new crawler user-agent showing up in the logs — those move without the user doing anything, which is exactly what makes them worth telling someone.
Make stability an explicit product claim, or drop it. "Nothing changed" is genuinely valuable information if the product is framed around confirming that. It's dead weight if the product is framed around finding problems. I had the second framing and the first output.
The measurement, if you want to run it on yours
If you generate a recurring artifact — report, digest, alert summary, weekly email — this is the query. Mine is scores, yours is whatever the artifact is about:
WITH ranked AS (
SELECT
domain_id,
score,
LAG(score) OVER (PARTITION BY domain_id ORDER BY created_at) AS prev
FROM reports
)
SELECT
count(*) AS reports,
count(*) FILTER (WHERE prev IS NOT DISTINCT FROM score) AS unchanged,
round(100.0 * count(*) FILTER (WHERE prev IS NOT DISTINCT FROM score) / count(*), 1)
AS pct_unchanged
FROM ranked
WHERE prev IS NOT NULL;
IS NOT DISTINCT FROM rather than = is the detail that matters: NULL = NULL is NULL, which is not true, so a plain = silently drops every pair where a score was missing — and missing scores cluster exactly where something odd is happening. IS NOT DISTINCT FROM treats two NULLs as equal and gives you the honest count.
Then read the number as a statement about your product rather than your thresholds. If most of your artifacts are empty, no amount of sensitivity tuning fills them.
The part I'd tell myself in advance
I spent weeks on threshold work because threshold work feels like the right kind of effort. It's measurable, it's in code, it produces diffs, and at the end of a day of it you can point at something.
Asking "should this report exist at all" produces no diff. It's also the question that was load-bearing, and I avoided it for a month by staying busy adjacent to it.
The tell, in hindsight: I was tuning the sensitivity of a detector without ever having measured how often the thing it detects occurs. That ordering is backwards, and it's a cheap check — one query, before the tuning, not after.
Top comments (4)
I’ve been there. I once spent almost a month tweaking content and keywords because my reports kept showing “issues,” but most of them didn’t lead to any meaningful action. 😅
The real problem wasn’t a lack of data—it was knowing which data actually mattered.
I started using SerpSpur to audit sites and prioritize technical issues, so instead of producing long reports full of noise, I could focus on problems that had a clear impact and next step.
Lesson learned: A useful SEO report shouldn’t just tell you what’s wrong—it should help you decide what to fix first.
Yeah, "long reports full of noise" is a real problem. I just don't think a better audit tool fixes it, at least not the version I ran into.
My 78% weren't reports with too many low-priority issues. They were reports with zero issues. Nothing to prioritize either way. The fix was sending fewer reports, not better ones.
Different layer of the same frustration, I guess. Glad the post resonated even if you got something different out of it than I put in.
Nice good
"78% had nothing to say" is a much more useful number than any accuracy metric, and the second half of your title is the part I would underline: tuning the wrong thing.
I went through the same arc with risk alerts in a trading system, and the shape was identical. Every alert I added made the system feel safer, so the natural instinct was to keep adding - and the actual protection did not move at all, because the added alerts were not on the path that could stop anything. The number that mattered was not "how many alerts fire" but "how long after a real breach does the behaviour change". Nobody was measuring it.
Two things I would offer from that experience:
Separate the two error types in your metrics, not just in your head. A missed event and a useless page are not symmetric, and averaging them into one "precision" hides which one you are trading away. In trading the missed event is the expensive one; in on-call the useless page is, because it trains people to ignore the channel that the missed event will arrive on. Same system, opposite optimisation.
Ask of every report: what decision does this change? If the answer is "none, it is informational", it belongs in a dashboard, not in a report. The 78% you removed were probably not wrong - they were decisionless. That is a different and more fixable property.
The uncomfortable version of "tuning the wrong thing" that I keep running into: teams optimise the metric that is easy to instrument (alert volume, delivery latency, uptime of the checker) rather than the one that describes the outcome, because the outcome metric requires an experiment rather than a dashboard. Your month of tuning is a good cautionary tale for exactly that trap.
Thanks for publishing the number rather than the fix alone - the 78% is what makes it credible.