DEV Community

Cover image for MCP Costs 72% of Your Context: The Fix Pi Shipped
Kiell Tampubolon
Kiell Tampubolon

Posted on

MCP Costs 72% of Your Context: The Fix Pi Shipped

The agent that spent a year refusing to support MCP just shipped MCP support, and the reason why matters to anyone running tool-heavy agents. On September 29, Pi (the minimal coding agent whose homepage used to say "Pi does not support MCP") released v0.99.0 with native Model Context Protocol support, announced in a post titled "You Said No MCP!". By October 1 it was sitting at #1 on Hacker News with hundreds of comments arguing about what it means.

I was researching the release for my own MCP servers, mostly because I have been on both sides of this fight. My agent context has been eaten alive by tool schemas before, and I wanted to know if Pi's approach is something I can actually steal. Short answer: yes, and you do not need Pi to do it. Here is the cost breakdown, the pattern behind their fix, and a minimal version you can wire into your own harness tonight.

How bad is the context tax, really?

The most quoted number in the whole saga came from Perplexity. Their CTO said at a conference in March 2026 that three MCP servers consumed 143,000 of their 200,000-token context window before the agent did any real work. That is a 72% tax. They dropped MCP internally after that.

You can reproduce the shape of the problem with any MCP client. Every session, the client loads full schema definitions for every connected tool. Individual tool definitions reportedly cost 550 to 1,400 tokens each. Pi's cited example was a Playwright MCP server dumping 21 tool definitions and 13,700 tokens per session.

Let me be honest about my numbers here: I have not lab-benchmarked Pi or Perplexity. What I can show is what the naive loading pattern costs in principle, using a single tool definition you would recognize:

# What the model sees for ONE tool, before it ever calls it
TOOL_SCHEMA = {
    "name": "create_issue",
    "description": "Create a GitHub issue. Use this when the user "
                   "reports a bug, wants to track work, or asks to "
                   "file something against the repository. Accepts "
                   "title (required), body (optional markdown), "
                   "labels (optional list of strings that must match "
                   "existing repo labels), assignees (optional), and "
                   "milestone (optional numeric id). Returns the "
                   "created issue object including number, html_url, "
                   "and state. If label validation fails, retry "
                   "without labels and mention them in the body.",
    "input_schema": {
        "type": "object",
        "properties": {
            "title": {"type": "string"},
            "body": {"type": "string"},
            "labels": {"type": "array", "items": {"type": "string"}},
            "assignees": {"type": "array", "items": {"type": "string"}},
            "milestone": {"type": "number"}
        },
        "required": ["title"]
    }
}
Enter fullscreen mode Exit fullscreen mode

Count it: that is roughly 200 tokens of description plus 80 of schema, for one tool. Multiply by 20 or 40 tools across a few servers and the tax is real. The model pays it every single session whether or not it ever calls the tool.

What did Pi actually change?

Pi did not bolt on the naive implementation. It shipped Codemode, and the design difference is worth studying:

# Naive MCP: schemas pre-loaded into context at session start
context_budget_spent = sum(len(t["description"]) for t in all_tools)

# Codemode pattern: tools are documentation, not context
# The model writes JavaScript that runs in a sandboxed
# interpreter with MCP tools available as functions.
agent_plan = """
// runs in QuickJS inside the harness, not in my context
const issue = await mcp.github.create_issue({
  title: "Flaky retry in upload client",
  labels: await mcp.github.list_labels("api")
});
return { number: issue.number, url: issue.html_url };
"""
# Only the return value comes back into context
Enter fullscreen mode Exit fullscreen mode

Two properties make this work. First, deferred loading: tools are visible to the model as discoverable documentation, not as context-filling schema blocks. Second, the sandbox lives in the harness, not inside each MCP server. That second part matters more than it sounds: earlier code-execution designs shipped one sandbox per server, and the sandboxes could not call into each other. One shared sandbox means the agent can compose tools from different vendors in a single script, and only the final result occupies context.

Pi's team gave three reasons for reversing course. The July 2026 MCP spec revision removed the initialize handshake and session layer, the parts that made production deployments painful. Their QuickJS sandbox already existed, so MCP support was cheap to add. And their stated philosophy: "We believe the best way to positively influence something is to embrace it."

Can I steal the pattern without adopting Pi?

Yes, and it is less code than you would expect. Here is a minimal deferred tool registry I wrote while researching this post. I have not benchmarked it in a real agent loop yet, so treat the output as illustrative:

# deferred_registry.py: minimal sketch of the Codemode idea
# Tools are exposed as docs; schemas are fetched on demand.

class DeferredToolRegistry:
    def __init__(self, servers):
        self.servers = servers          # name -> MCP client
        self._schemas = {}              # lazy cache

    def index(self):
        """One line per tool. This is ALL the model sees upfront."""
        lines = []
        for name, client in self.servers.items():
            for tool in client.list_tools():
                lines.append(f"{name}.{tool.name}: "
                             f"{tool.description.splitlines()[0]}")
        return "\n".join(lines)

    def schema(self, qualified_name):
        """Model pulls the full schema only when it needs it."""
        if qualified_name not in self._schemas:
            server, tool = qualified_name.split(".")
            self._schemas[qualified_name] = \
                self.servers[server].get_tool_schema(tool)
        return self._schemas[qualified_name]
Enter fullscreen mode Exit fullscreen mode

The upfront cost drops from "every schema for every tool" to one summary line per tool. In my sketch that is roughly 15 to 30 tokens per tool instead of 550 to 1,400. That is the whole trick: move the schema dump from session start to first use.

The full Codemode version goes further and puts tool calls inside sandboxed code so multi-step chains return one result instead of bouncing through context. If you build on the registry sketch, the incremental step is an interpreter with the registry bound into its namespace.

Does the criticism still apply, or did it win?

This is where the HN thread split, and I think both camps are partly right.

The capitulation reading: every prominent MCP skeptic now ships MCP, the protocol has 17,000+ public servers and reports hundreds of millions of monthly SDK downloads, and OpenAI, Google, Microsoft, AWS, Cloudflare, and GitHub are all involved in its governance. When even the team whose homepage taunted the protocol folds, the debate was settled by ecosystem gravity, not by which side argued better.

The evolution reading: the MCP Pi adopted is not the MCP that was criticized. The stateless spec revision, deferred tool loading, and harness-level sandboxes fix the exact complaints the skeptics raised. The critics did not lose; their feedback became the roadmap.

One caveat I want to be careful about: a QuickJS sandbox is a token-efficiency mechanism, not a security boundary. Do not read "sandboxed code execution" as protection against prompt injection or malicious tool output. That problem is unchanged, and the MCP Python SDK OAuth credential theft disclosure from two days ago is a reminder the trust surface is still rough.

What should I actually do this week?

The practical takeaways from the whole thread, ordered by effort:

  1. Stop pre-loading full tool schemas. Whatever protocol you use, expose a one-line index and fetch schemas on demand.
  2. Return structured data from your tools, not prose. Tools that dump text cannot be chained by a script.
  3. If you build harnesses, watch where frontier models are trained. Code-first tool calling is increasingly in the training data, so models will perform better at it regardless of which design is theoretically superior.

Is code-first orchestration the new default?

Here is my honest position, and I expect pushback: I think pre-loading tool schemas was always a transitional artifact, like pasting a README into a prompt. Deferred discovery beats upfront dumping for the same reason lazy loading beats eager loading everywhere else in software.

But there is a real counterargument. Dumping schemas gives the model a complete picture of its capabilities in one shot, which can produce better planning on tasks that need unusual tool combinations. If your agent routinely needs to discover cross-tool workflows on its own, the tax might be buying you something. I genuinely do not know where that line sits yet, and I would love to hear from anyone who has measured it.

If you have run both patterns against a real workload, tell me what happened. My own next step is wiring the registry sketch into a harness and running the same 20-tool suite both ways, and I will write up whatever I find, including if I am wrong.

Further reading

I have been circling this topic for a while. Related pieces of mine:

Sources: Pi 0.99 release notes and the "You Said No MCP!" announcement, byteiota's Codemode explainer (Oct 1), the Hacker News thread, and Perplexity CTO Denis Yarats' comments at Ask 2026.

Top comments (0)