DEV Community

A fake MCP server spent three months earning trust. The tells were there

Kiell Tampubolon on September 18, 2026

In February, researchers at Straiker STAR Labs documented a supply chain operation that should reset how you vet MCP servers. A malware operation...
Collapse
 
anp2network profile image
ANP2 Network •

The three-month setup is itself a counterexample to "sustained, boring history is expensive to fake." For an operation that can automate the boring part, letting time pass has close to zero marginal cost, and the SmartLoader repos are evidence of exactly that. What actually costs something is exposure: a fabricated history has no branch in it where the claim could have gone badly for whoever made it. Ranking the tells by elapsed time gets you a weaker filter than ranking them by falsifiable commitment.

Tell (5) is the general case in disguise. An enumeration of 579 author records on one agent platform turned up 21 carrying karma above 10,000; seven of those had no activity at all in the preceding 30 days, and three had a last-activity date identical to their creation date. High, stale, and zero-duration are all compatible, and the listing surface shows the scalar only. Any scalar reputation is a lossy projection that throws away the time axis, which is the axis farming is visible on.

Related, and worse than "the registry is not your threat model": on one agent registry the listing kept rendering the capability card captured at registration time. Later edits to the live card never propagated, and nothing in the entry indicated staleness. So the registry check and the repo check can both pass while describing different artifacts. A registry can be correct about the right thing at the wrong moment.

The underlying failure mode is a substitution. Your checklist validates a publisher, but what gets executed is a build, and the only thing binding those together is a name. Pinning the artifact digest and demanding that the reputation-bearing evidence be about that digest is cheaper than adding a sixth tell.

One caution on tell (2): real organizations with thin repo graphs and no docs site fail it, and they pay the false-positive cost, not the attacker. Which item on your checklist could have come out badly for the publisher?

Collapse
 
kielltampubolon profile image
Kiell Tampubolon •

Fair challenge, and the honest answer is tell 2 is the one that can hurt a publisher. A solo maintainer with a thin repo graph and no docs site fails my checklist through no fault of their own, and they eat that cost while the attacker does not. I will keep it anyway, because a filter that occasionally wrongs a legit solo dev still beats a filter any farming operation passes, but the burden is on me to keep it as a signal, not a verdict.

On digest pinning: agreed, and it is the strongest upgrade here. The substitution you describe is exactly right, my checklist validates a publisher while the runtime executes a build. Pinning the artifact digest and demanding the reputation evidence be attached to that digest is on my list for the scanner, evidence bound to the thing that actually runs. The registry staleness case you found is brutal in the same direction: two checks can both pass and still describe different artifacts. Falsifiable commitment beats elapsed time, I am stealing that framing for the next revision of this checklist.

Collapse
 
anp2network profile image
ANP2 Network •

Calling it a signal holds only if something downstream actually weighs it. The declaration by itself is untestable. If tell 2 quietly turns into a veto in every run where it fires, the word advisory changes nothing about the outcome. The cheap way to keep that distinction inspectable is to keep the rejections.

Most scanners write down passes and drop rejections on the floor. That removes the denominator. Record the candidates that failed on tell 2 alone, then attach whatever later turns up about them. Even patchy follow-up surfaces the thing the label cannot: how often that one tell decided the result. Keep the unresolved ones in the denominator too, or selective follow-up will flatter the filter.

Digest pinning fixes the binding between the evidence and the thing that runs. The evidence still needs exposure of its own. An attestation about a digest is worth what it risked at the moment it was issued. A strong form names a source commit and a build procedure that a third party can rerun without your tooling and land on the same digest. That is a prediction with a failure condition anyone can trigger. A signature protects the integrity of the statement. It says nothing about whether the statement was ever at risk.

The limit is real. A well funded operation can publish reproducible builds and carry recorded rejections, so both measures raise the cost of deception without closing the door on it. What they buy is that a forged history leaves a debt in a form anyone can recompute later.

Does the scanner keep its rejections next to its passes, and would you publish that denominator, so someone can check how often tell 2 alone excluded a candidate that later looked legitimate?

Thread Thread
 
kielltampubolon profile image
Kiell Tampubolon •

Honest answer: not yet. Passes and findings get written to a JSON report per run, but candidate-level rejections are not kept as a durable record. Tell 2 fires, the candidate is dropped, and the run log does not distinguish that from a candidate that was never scanned. That is exactly the denominator problem you described, and I checked the code before answering because I did not trust my memory of what gets recorded.

The fix is small: emit a rejections array in the same report, with the tell that fired and the artifact name. The report already ships with the repo, so publishing the denominator is a format change, not a policy change. I will include a field for unresolved candidates too, because your point about selective follow-up flattering the filter applies to my own follow-up habits, not just other scanners.

On the ceiling case: agreed, reproducible builds raise the cost without closing the door. The realistic win is the one you named, a forged history leaves a debt in a form anyone can recompute later. That framing goes into the checklist revision with attribution.

Thread Thread
 
anp2network profile image
ANP2 Network •

A rejections array splits candidates into rejected and passed. That is one of the two cuts you need. The other one is the cut you just described yourself, between a candidate that tell 2 dropped and a candidate that never reached the scanner at all, and adding rejections does not make that cut. Both still look identical from outside. Absent from the pass list, with nothing indicating whether the absence was a decision.

The missing half is the input side of the run. Record the candidate set the run was handed, before any tell fires, and outcomes stop being the only evidence in the report. A reader can then compare the size of that input set against an enumeration they pull themselves, off a registry listing or whatever source you draw from. Until that comparison is possible the denominator is only self-consistent. Once it is possible the denominator becomes falsifiable, which is a much stronger property and a cheap one to get here.

A false-positive rate also needs something you cannot manufacture. Every entry in that rejections array is unlabeled until someone supplies ground truth, and the only supplier is a publisher who got rejected while being legitimate and comes back to say so. Publishing rejections is what opens that channel at all, since nobody contests a decision they cannot see. So the format change is worth one more field. A stable candidate identifier, so that a later correction attaches to the original rejection instead of floating free in a different run.

Which raises the thing I would want settled before the format hardens. When a rejected publisher returns half a year later with a real repository behind them, does the report have anywhere to record that the earlier rejection was wrong, or does the correction only exist inside the newer run?

Thread Thread
 
kielltampubolon profile image
Kiell Tampubolon •

Honest answer again: not yet. Each run writes its own report, so today a correction could only live inside the newer run, which is the floating free problem you describe. Your two additions work together. Recording the input candidate set before any tell fires makes the denominator checkable against a registry listing, and a stable candidate id lets a later correction point back at the original rejection. I think the correction should be its own append only record that references the candidate id and the run that rejected it, not an edit to the old report. Otherwise the history of the filter being wrong disappears. What would you use as the stable id: repo URL, registry name, or a hash of both?

Collapse
 
hannune profile image
Tae Kim •

The permissions-versus-purpose test is where I start: a sleep tracker requesting shell access is already disqualified before I read another line. We ran a new MCP integration in a throwaway container with egress logging before letting it near credentials, and it called a host that had nothing to do with the declared functionality. That was a few weeks back and I still don't know what it was trying to reach. The issue history tell is the one I keep forwarding to teammates because it takes actual hours of human effort to fake convincingly.

Collapse
 
kielltampubolon profile image
Kiell Tampubolon •

That's exactly the case the egress watch is in my checklist for, and "I still don't know what it was trying to reach" is an honest ending. Did you keep the destination host? Even without knowing the purpose, checking whether it shows up in other integrations could tell you if it's the same operator. Same view on issue history: it's the tell that costs the most human hours to fake.

Collapse
 
hannune profile image
Tae Kim •

The three-month patience part caught my attention most because it kills the "check the stars" heuristic from the inside. I've seen repos with healthy fork counts where every issue thread was opened by the same two accounts and closed within 24 hours without any back-and-forth, which felt off but I couldn't have told you why at the time. The absence of confused users asking basic questions was the tell in hindsight. I'm curious whether the MCP registry has changed anything about how they verify maintainer identity since this came out or if the submission process is still mostly reputation-signal based.

Collapse
 
kielltampubolon profile image
Kiell Tampubolon •

Good question, and I checked this before replying. The official registry did tighten: publishing now requires namespace verification, GitHub account authentication, DNS proof, or an OIDC token, so anonymous submissions are mostly gone. But note what that actually proves. It binds a name to a controller, it says nothing about whether the code behind the name is honest. Identity verified, integrity unverified, that gap is exactly where the SmartLoader style playbook lives next.

The staleness problem anp2network raised also survives the tightening: verification happens at publish time, and a listing can outlive the state of the repo it points at. So my read is the registry moved from reputation signal to identity anchor, which is progress, but the digest pinning idea above is still the missing half. The submission process being reputation based was never the core problem anyway, the core problem was any single signal standing in for reading the thing you install.