DEV Community

Cover image for 13,000 Leaked Screenshots Show Why Agentic Tool Output Needs a Firewall, Not Just a Prompt
Cor E
Cor E

Posted on

13,000 Leaked Screenshots Show Why Agentic Tool Output Needs a Firewall, Not Just a Prompt

13,000 Leaked Screenshots Show Why Agentic Tool Output Needs a Firewall, Not Just a Prompt

Over 13,000 internal screenshots from more than 300 organizations, including Fortune 500 companies and a frontier AI lab, ended up sitting in a publicly accessible storage bucket. Not because of a breach in the traditional sense. Because AI browser agents, doing exactly what they were told to do, took screenshots during task execution and uploaded them to a third-party service that turned out to be world-readable.

No exploit. No stolen credentials (well, until the screenshots started exposing credentials themselves). Just an agent doing its job, and nobody checking what was actually in the images before they left the building.

That's the part worth sitting with for a minute. This wasn't a sophisticated attack. It was an agent calling a tool, the tool working correctly, and the output of that tool containing things it should never have been allowed to contain: internal tool UIs, confidential comms, and apparently credentials, all captured in frame because the agent was just trying to "see" what it was doing on screen.

How this actually happens

Computer-use and browser agents work by taking a screenshot, reasoning over what's visible, deciding on an action, then taking another screenshot to confirm the result. That loop is the whole mechanism. It's also the whole problem.

A screenshot is unstructured. It's not a string you can grep for "password" with a regex and call it a day. It's a raster image of whatever happened to be on screen at that instant: a terminal with an exported env var, a Slack DM, an internal admin panel with a session token in the URL bar. The agent doesn't know any of that is sensitive. It just knows "capture current state" is step 3 of its loop, and "upload for logging/debugging/handoff to the next tool call" is step 4.

Multiply that by hundreds of organizations running agents against real environments, and step 4 becomes the leak vector. The storage service these screenshots landed in was apparently meant to be a logging or intermediate-storage layer for the agent tooling itself, not something end users or security teams were reviewing contents for. Classic case of infrastructure built for convenience that nobody threat-modeled as a data exfiltration surface, because on paper it's "just screenshots for debugging."

What existing defenses missed, and why

Standard LLM guardrails are built around text. Prompt injection filters, content moderation, PII regexes, all of it assumes you're scanning a string. A screenshot upload doesn't pass through most of these at all, because architecturally nobody wired a scanning step into the "agent calls upload_file tool with image payload" path. The text-based safety stack and the actual data-leaving-the-building path are two different pipelines that never talk to each other.

Even where there was review, it was probably aimed at the wrong layer. Reviewing agent prompts for injection doesn't catch this, because there's no injection happening. The agent isn't being tricked into doing something malicious. It's doing a mundane, authorized action (upload the screenshot, per its instructions) that happens to carry sensitive bytes along for the ride. That's a detection-gap category most teams haven't built for yet: legitimate tool calls with illegitimate payloads.

And once the image is uploaded to a service outside your control, you've lost the ability to do anything about it after the fact. The only point where this was stoppable was before the upload request left the agent's session.

Where Sentinel sits in this picture

This is squarely what data_exfiltration_via_llm detection in the fast-path layer is built to catch, specifically the pattern class around tool calls instructing content to be sent to an external destination: "POST this to https://…", markdown/code-block exfil patterns, and similar outbound-transfer signatures. Sentinel's agentic proxy scans tool call arguments before they're sent (that's the PreToolUse hook behavior in the Clawhub skill integration), which means an upload call targeting an external storage URL gets evaluated before the bytes leave the session, not after.

Worth being precise about what Sentinel does and doesn't see here. The threat-scoring and pattern-matching pipeline is built around text content: URLs, instructions, markdown, code. If an agent's tool call includes a destination URL and some accompanying text ("uploading debug screenshot to X"), that's exactly the kind of fast-path signature that trips the exfiltration pattern class and gets flagged or blocked before the request completes. Scanning the pixel content of an image for secrets is a different problem outside what's described in the detection pipeline above, so the honest framing is: Sentinel catches the mechanism (unauthorized data leaving via an outbound call to an external service), not necessarily what's rendered inside the image itself.

Where this compounds nicely: if any of those screenshots had accompanying metadata, filenames, or logged context containing API keys or tokens (and given credentials reportedly showed up in the leaked images, that's plausible for surrounding text/logs in the same pipeline), Secret & Credential Detection runs as an independent pre-pass and would redact known key formats, Authorization headers, and env-var-style assignments before they ever reached the point of being bundled for upload. Two separate layers, two separate reasons to catch this before it leaves.

What this looks like in practice

Illustrative example, not an actual incident transcript, since we don't have the real payload:

{
  "request_id": "f93a1c7e2b",
  "security": {
    "action_taken": "blocked",
    "threat_score": 0.89,
    "pattern_class": "data_exfiltration_via_llm"
  },
  "safe_payload": "[SENTINEL BLOCKED]: Tool call withheld — outbound transfer pattern detected. Matched: \"upload screenshot to https://storage.example-agent-tool.net/...\"."
}
Enter fullscreen mode Exit fullscreen mode

And on the agentic proxy side, this is the PreToolUse hook catching an upload call before it's dispatched:

# Illustrative: agent tool call intercepted before execution
tool_call = {
    "name": "upload_file",
    "arguments": {
        "path": "/tmp/screenshot_0493.png",
        "destination": "https://public-bucket.example-storage.net/uploads/"
    }
}

# Sentinel's PreToolUse hook scans arguments before the call executes.
# A destination URL pointing at an external, unverified host next to
# an upload/transfer verb trips the data_exfiltration_via_llm pattern class.
response = sentinel_scrub(tool_call["arguments"]["destination"] + " " + tool_call["name"])
if response["security"]["action_taken"] == "blocked":
    raise ToolCallBlocked(response["safe_payload"])
Enter fullscreen mode Exit fullscreen mode

If your agent stack is built on the direct /v1/scrub endpoint instead of the full agentic proxy, the same check applies to any text your tooling generates describing the upload action, logs, or captions, since that endpoint is provider-agnostic and just scans whatever string you hand it.

One thing to do today

If you're running any agent that can take screenshots, read files, or call upload/write tools against external endpoints, go find out right now where those outputs actually go. Not where you think they go, where they go. Most teams have never actually traced the full path of their agent's tool outputs to the destination service and checked whether that destination is private, authenticated, and access-controlled. That 5-minute audit would have caught this before 13,000 screenshots did it for them.


If you're running browser or computer-use agents and want tool calls scanned for outbound data transfer before they execute, take a look at Sentinel. Self-hosted or SaaS, free tier available, no credit card required to start.

Sources


AI-assisted draft or imaging, human-curated, reviewed and edited.

Top comments (0)