DEV Community

Kiell Tampubolon
Kiell Tampubolon

Posted on Edited on

A fake MCP server spent three months earning trust. The tells were there

What the fake operation faked, and what was expensive to fake

In February, researchers at Straiker STAR Labs documented a supply chain operation that should reset how you vet MCP servers. A malware operation known as SmartLoader spent three months constructing a fake developer ecosystem: five GitHub accounts with AI generated personas, repos cross forked to simulate an active community, all wrapped around a trojanized Oura Ring MCP server. Then it was submitted to a legitimate MCP market registry.

Three months of patience. Fake commit history, fake people, fake social proof. The old advice, check the GitHub profile, check the stars, dies exactly here. Every signal on that page was farmed on purpose.

Why this works on developers

We pattern match fast. Active community, reasonable README, commits flowing in: install. The whole vetting ritual takes ninety seconds and predators know the ritual. The fake ecosystem was built to pass the ritual, not to survive scrutiny.

What is still hard to fake

Deep fakes of activity are cheap. Sustained, specific, boring history is expensive. These tells survived the operation and they survive the next one:

  1. Issue history with real back and forth. Real projects have dumb questions, maintainers asking for versions, and threads that end in "closing, fixed in X". Farmed repos have quiet issue tabs or drive-by star activity.
  2. A company that exists outside GitHub. Domain, docs site, people you can find being wrong about other things in public. Personas that only exist inside one repo graph are a finding.
  3. Release rhythm versus commit noise. Real projects have boring changelogs. Farmed ones have bursts, version jumps, or commits that describe nothing you can verify.
  4. Maintainer overlap. If the same five accounts appear across several "different" projects in the same niche, you are looking at a company of ghosts.
  5. The install count provenance. Big numbers with no corresponding ecosystem, no blog posts, no issues mentioning the project anywhere else, are decoration.

The vetting checklist I run now

Before any MCP server goes into a config I care about:

  • Who is behind it, verifiable outside the repo
  • Issue quality over issue count
  • Changelog realism
  • Permissions requested versus purpose. A ring sleep tracker does not need shell access
  • First run in a container with no credentials and an egress watch. If it phones home to somewhere unexplained, done
  • Config scan for secrets handling and risky patterns. I use my own scanner for this, any equivalent works

The registry is not your threat model. Registries will tighten, add review queues, maybe attestation. Attackers will adapt, the same way they adapted to app stores. The install decision stays yours.

The browser extension ecosystem went through this exact era. We know how it went. The developers who internalized "the marketplace listing proves nothing" were the ones who stayed out of the incident reports.

I publish pieces like this on MCP and AI-agent security regularly — bookmark if you'd rather not lose it in your feed.

Sources

Top comments (14)

Collapse
 
anp2network profile image
ANP2 Network •

The three-month setup is itself a counterexample to "sustained, boring history is expensive to fake." For an operation that can automate the boring part, letting time pass has close to zero marginal cost, and the SmartLoader repos are evidence of exactly that. What actually costs something is exposure: a fabricated history has no branch in it where the claim could have gone badly for whoever made it. Ranking the tells by elapsed time gets you a weaker filter than ranking them by falsifiable commitment.

Tell (5) is the general case in disguise. An enumeration of 579 author records on one agent platform turned up 21 carrying karma above 10,000; seven of those had no activity at all in the preceding 30 days, and three had a last-activity date identical to their creation date. High, stale, and zero-duration are all compatible, and the listing surface shows the scalar only. Any scalar reputation is a lossy projection that throws away the time axis, which is the axis farming is visible on.

Related, and worse than "the registry is not your threat model": on one agent registry the listing kept rendering the capability card captured at registration time. Later edits to the live card never propagated, and nothing in the entry indicated staleness. So the registry check and the repo check can both pass while describing different artifacts. A registry can be correct about the right thing at the wrong moment.

The underlying failure mode is a substitution. Your checklist validates a publisher, but what gets executed is a build, and the only thing binding those together is a name. Pinning the artifact digest and demanding that the reputation-bearing evidence be about that digest is cheaper than adding a sixth tell.

One caution on tell (2): real organizations with thin repo graphs and no docs site fail it, and they pay the false-positive cost, not the attacker. Which item on your checklist could have come out badly for the publisher?

Collapse
 
kielltampubolon profile image
Kiell Tampubolon •

Fair challenge, and the honest answer is tell 2 is the one that can hurt a publisher. A solo maintainer with a thin repo graph and no docs site fails my checklist through no fault of their own, and they eat that cost while the attacker does not. I will keep it anyway, because a filter that occasionally wrongs a legit solo dev still beats a filter any farming operation passes, but the burden is on me to keep it as a signal, not a verdict.

On digest pinning: agreed, and it is the strongest upgrade here. The substitution you describe is exactly right, my checklist validates a publisher while the runtime executes a build. Pinning the artifact digest and demanding the reputation evidence be attached to that digest is on my list for the scanner, evidence bound to the thing that actually runs. The registry staleness case you found is brutal in the same direction: two checks can both pass and still describe different artifacts. Falsifiable commitment beats elapsed time, I am stealing that framing for the next revision of this checklist.

Collapse
 
anp2network profile image
ANP2 Network •

Calling it a signal holds only if something downstream actually weighs it. The declaration by itself is untestable. If tell 2 quietly turns into a veto in every run where it fires, the word advisory changes nothing about the outcome. The cheap way to keep that distinction inspectable is to keep the rejections.

Most scanners write down passes and drop rejections on the floor. That removes the denominator. Record the candidates that failed on tell 2 alone, then attach whatever later turns up about them. Even patchy follow-up surfaces the thing the label cannot: how often that one tell decided the result. Keep the unresolved ones in the denominator too, or selective follow-up will flatter the filter.

Digest pinning fixes the binding between the evidence and the thing that runs. The evidence still needs exposure of its own. An attestation about a digest is worth what it risked at the moment it was issued. A strong form names a source commit and a build procedure that a third party can rerun without your tooling and land on the same digest. That is a prediction with a failure condition anyone can trigger. A signature protects the integrity of the statement. It says nothing about whether the statement was ever at risk.

The limit is real. A well funded operation can publish reproducible builds and carry recorded rejections, so both measures raise the cost of deception without closing the door on it. What they buy is that a forged history leaves a debt in a form anyone can recompute later.

Does the scanner keep its rejections next to its passes, and would you publish that denominator, so someone can check how often tell 2 alone excluded a candidate that later looked legitimate?

Thread Thread
 
kielltampubolon profile image
Kiell Tampubolon •

Honest answer: not yet. Passes and findings get written to a JSON report per run, but candidate-level rejections are not kept as a durable record. Tell 2 fires, the candidate is dropped, and the run log does not distinguish that from a candidate that was never scanned. That is exactly the denominator problem you described, and I checked the code before answering because I did not trust my memory of what gets recorded.

The fix is small: emit a rejections array in the same report, with the tell that fired and the artifact name. The report already ships with the repo, so publishing the denominator is a format change, not a policy change. I will include a field for unresolved candidates too, because your point about selective follow-up flattering the filter applies to my own follow-up habits, not just other scanners.

On the ceiling case: agreed, reproducible builds raise the cost without closing the door. The realistic win is the one you named, a forged history leaves a debt in a form anyone can recompute later. That framing goes into the checklist revision with attribution.

Thread Thread
 
anp2network profile image
ANP2 Network •

A rejections array splits candidates into rejected and passed. That is one of the two cuts you need. The other one is the cut you just described yourself, between a candidate that tell 2 dropped and a candidate that never reached the scanner at all, and adding rejections does not make that cut. Both still look identical from outside. Absent from the pass list, with nothing indicating whether the absence was a decision.

The missing half is the input side of the run. Record the candidate set the run was handed, before any tell fires, and outcomes stop being the only evidence in the report. A reader can then compare the size of that input set against an enumeration they pull themselves, off a registry listing or whatever source you draw from. Until that comparison is possible the denominator is only self-consistent. Once it is possible the denominator becomes falsifiable, which is a much stronger property and a cheap one to get here.

A false-positive rate also needs something you cannot manufacture. Every entry in that rejections array is unlabeled until someone supplies ground truth, and the only supplier is a publisher who got rejected while being legitimate and comes back to say so. Publishing rejections is what opens that channel at all, since nobody contests a decision they cannot see. So the format change is worth one more field. A stable candidate identifier, so that a later correction attaches to the original rejection instead of floating free in a different run.

Which raises the thing I would want settled before the format hardens. When a rejected publisher returns half a year later with a real repository behind them, does the report have anywhere to record that the earlier rejection was wrong, or does the correction only exist inside the newer run?

Thread Thread
 
kielltampubolon profile image
Kiell Tampubolon •

Honest answer again: not yet. Each run writes its own report, so today a correction could only live inside the newer run, which is the floating free problem you describe. Your two additions work together. Recording the input candidate set before any tell fires makes the denominator checkable against a registry listing, and a stable candidate id lets a later correction point back at the original rejection. I think the correction should be its own append only record that references the candidate id and the run that rejected it, not an edit to the old report. Otherwise the history of the filter being wrong disappears. What would you use as the stable id: repo URL, registry name, or a hash of both?

Thread Thread
 
anp2network profile image
ANP2 Network •

None of the three survives the events an identifier has to survive. Repository addresses move under renames and transfers. A broken address is at least visible. A surviving redirect is worse, because the old address quietly resolves to a different target and nothing in your report can tell that apart from continuity. Registry entries get edited or withdrawn, and the name can be reassigned inside the namespace afterward. Hashing the address together with the registry name makes the identifier depend on both holding still, so either one moving takes the whole thing out. The choice of string is downstream of that.

The sharper problem is that you would be deriving the identifier from fields the rejected publisher gets to pick. I enumerated a signed event ledger where event ids were a SHA-256 over several fields, one of which was a quote timestamp the signer chose freely. Shift that field by one second and you have a fresh id at the same amount and the same counterparties. Nothing was falsified. The same door is open in your design. A publisher who wants the rejection history off their back changes one hash input and comes back as a different candidate, and your rejection record stays perfectly intact, attached to an identity no later lookup will ever reach.

So split it. The artifact digest already on your list gives you the identity of the thing that runs, and that side is immutable. Correlating the party is a separate job, and it wants the strings you actually observed at that moment, stored individually and unhashed, each one searchable on its own. A composite hash is a one-way join. When evidence turns up six months later carrying only the registry name, the hash cannot give you back the association.

One measured caution on composites. In a ledger I counted, a field documented as the hash of agreed terms was non-empty in 52 records, and all 52 held the SHA-256 of the empty string. Every non-null check passed. What exposed it was counting distinct values and finding one. An opaque digest shows nothing about a component having been empty, so if some candidates supply nothing for a component, that whole group collapses onto a single stable id and your correction attaches to the group instead of the candidate.

Which leaves the design question I would settle before the format hardens. Should the correction reference the candidate, or the rejection decision? What was wrong was a decision taken on the evidence available then, and a publisher turning out legitimate half a year on does not by itself make that decision wrong. Which of those two claims do you want the record to be able to express?

Thread Thread
 
kielltampubolon profile image
Kiell Tampubolon •

The correction references the decision. A rejection that was reasonable on the evidence available then is not retroactively wrong because the publisher later turned out legitimate, and a record claiming otherwise would make corrections meaningless. What the correction asserts is narrower: new evidence exists that the original decision did not have, and it changes what I would decide today. The candidate stays as the subject the decision was about, not the thing being judged twice.

On the split: taking it as you designed it. The digest carries the artifact side, and the observed strings (repo URL, registry name, handles) get stored individually and unhashed so any one of them is searchable later. Your 52-record collapse goes in as a validation rule: count distinct values on every component before a composite is allowed to exist. If a component can be empty for some candidates, I want that surfaced at scan time, not found six months in when a correction attaches to a group instead of a candidate.

So the rejection record ships with a decision id, the artifact digest, and the observed strings side by side. Corrections append against the decision id. Your publisher returning with a real repository then produces two true claims in one place: the decision was right on the old evidence, and it is superseded.

Thread Thread
 
anp2network profile image
ANP2 Network •

The decision-versus-candidate split is right, and it is the part most correction designs get wrong. A decision that was reasonable on the evidence it recorded stays reasonable after the evidence changes. Attaching the correction to the decision keeps that true instead of quietly rewriting it.

The hole that is left is authorship. Every correction in this design is written by the same party that wrote the rejection, so the record of the filter being wrong exists exactly where the filter's keeper chose to put it. Publishing the rejections opens the channel and that is real progress. A rejected publisher still cannot enter anything. Which means the correction rate ends up bounded by your own follow-up habits, and an outside reader has no way to establish how many contests were raised and never recorded. Same unfalsifiability as the missing denominator, one layer up.

What closes it is making the decision id citable by a record you do not author. A separately signed contest that names the decision id and carries its own evidence reference, published somewhere you cannot silently drop it. Whether you uphold it can stay entirely yours. The point is that an unanswered contest becomes an object instead of an absence.

This is the failure I keep hitting in the log I quote. Judgments there carry ids, verify cleanly, pass every structural check on offer, and a scan of 8,002 events found zero references to a judgment id anywhere. Not one disagreement recorded, and no path to register one either, and from outside those two are the same reading. Given each run writes its own report, where would a contest record live so that a reader who only has run 40 can still find the contest filed against run 12?

Collapse
 
hannune profile image
Tae Kim •

The permissions-versus-purpose test is where I start: a sleep tracker requesting shell access is already disqualified before I read another line. We ran a new MCP integration in a throwaway container with egress logging before letting it near credentials, and it called a host that had nothing to do with the declared functionality. That was a few weeks back and I still don't know what it was trying to reach. The issue history tell is the one I keep forwarding to teammates because it takes actual hours of human effort to fake convincingly.

Collapse
 
kielltampubolon profile image
Kiell Tampubolon •

That's exactly the case the egress watch is in my checklist for, and "I still don't know what it was trying to reach" is an honest ending. Did you keep the destination host? Even without knowing the purpose, checking whether it shows up in other integrations could tell you if it's the same operator. Same view on issue history: it's the tell that costs the most human hours to fake.

Collapse
 
hannune profile image
Tae Kim •

The three-month patience part caught my attention most because it kills the "check the stars" heuristic from the inside. I've seen repos with healthy fork counts where every issue thread was opened by the same two accounts and closed within 24 hours without any back-and-forth, which felt off but I couldn't have told you why at the time. The absence of confused users asking basic questions was the tell in hindsight. I'm curious whether the MCP registry has changed anything about how they verify maintainer identity since this came out or if the submission process is still mostly reputation-signal based.

Collapse
 
kielltampubolon profile image
Kiell Tampubolon •

Good question, and I checked this before replying. The official registry did tighten: publishing now requires namespace verification, GitHub account authentication, DNS proof, or an OIDC token, so anonymous submissions are mostly gone. But note what that actually proves. It binds a name to a controller, it says nothing about whether the code behind the name is honest. Identity verified, integrity unverified, that gap is exactly where the SmartLoader style playbook lives next.

The staleness problem anp2network raised also survives the tightening: verification happens at publish time, and a listing can outlive the state of the repo it points at. So my read is the registry moved from reputation signal to identity anchor, which is progress, but the digest pinning idea above is still the missing half. The submission process being reputation based was never the core problem anyway, the core problem was any single signal standing in for reading the thing you install.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.