A CVE published yesterday afternoon says the Python sandbox in CrewAI, a widely used agent framework, blocked nine module names and still lost. The record's wording is blunt for a CVE: the blocklist "operates at the wrong level of abstraction." The escape it describes never uses an import statement at all.
I was researching agent sandbox failures this week anyway. Between the GreyNoise agent swarm report, the SGLang pickle RCE last weekend, and Bengio's agent-misbehavior paper trending on Hacker News, every feed I follow is agent security right now. When my trending scan surfaced CVE-2026-37008 this morning, I searched for writeups and found none. Zero HN threads, zero Dev.to posts, no news pickup. So I pulled the NVD record, the MITRE CNA data, and the fix commit from GitHub, and read all three. Here is what the record says, what the vulnerable code actually did, why the fix deletes the feature instead of extending the list, and three checks worth running on any sandbox your agent framework ships.
This is the entire security model the CVE is about, verbatim from the pre-fix source:
# From crewai-tools code_interpreter_tool.py, pre-fix source (SandboxPython)
BLOCKED_MODULES = {"os", "sys", "subprocess", "shutil",
"importlib", "inspect", "tempfile",
"sysconfig", "builtins"}
def restricted_import(name, *args, **kwargs):
if name in SandboxPython.BLOCKED_MODULES:
raise ImportError(f"Importing '{name}' is not allowed.")
return __import__(name, *args, **kwargs)
Nine names. One string comparison. Keep that in mind while we walk the record.
What does CVE-2026-37008 actually say?
Start with the record itself, because it is unusually readable. Published 2026-09-13 at 21:17 UTC by MITRE as the CNA. CVSS 3.1 base 8.1 HIGH with the vector AV:L/AC:H/PR:N/UI:N/S:C/C:H/I:H/A:L. No CVSS 4.0 score exists yet, and NVD has ingested the record without analyzing it, so the only score on any official source is the CNA's own. The weakness is CWE-424, "Improper Protection of Alternate Path," which in plain words means: you guarded one path and an alternate path stayed open.
The description, verbatim:
CrewAI before fb2323b offers a Python blocklist approach that operates at the wrong level of abstraction, a different vulnerability than CVE-2026-2275. Import-time blocking of module names does not address the availability of Python's complete object graph. For example, calling ctypes.CDLL(None) loads the C library without relying in any import statements. In other words, a within-process sandbox cannot merely account for the import system and instead must account for the complete runtime of the Python interpreter.
Read the middle sentence twice. The claim is not "one dangerous module was missing from the list." The claim is that a list of module names cannot work in principle, because the interpreter's object graph is reachable without ever touching the import system.
Why the affected range is a commit hash
The affected range is not a version number. It is defined against git: every revision before commit fb2323b3deb3ec62b3965526857e77a2264e4cd0 is affected, everything after is not. If you pip install CrewAI, there is no version string to compare against, and the GitHub advisory (GHSA-2q68-3cp7-72v9) is unreviewed with "Unknown" in both the affected and patched fields. This is the awkward part of commit-boundary CVEs, and it matters more than usual here because the fix landed six months ago.
Why did a nine-name blocklist fail?
The sandbox lived in crewai-tools, in the CodeInterpreterTool module, in a class called SandboxPython. When the framework needed to run model-written Python and Docker was unavailable, it fell back to running that code with plain exec() in the same process, with the builtins dict filtered to remove ten unsafe names like exec, eval, and open, and with restricted_import swapped in for __import__. Notice what is absent from the nine: ctypes, socket, pathlib. The vendor's own statement to CERT/CC about the sibling disclosure concedes the ctypes gap directly.
But the deeper problem is what the record says: even a complete list of names guards the wrong thing.
Two ways past a name filter, no import required
The record's own example is the first one. ctypes.CDLL(None) asks the platform loader to open the running process itself, which hands back a handle whose symbol table includes libc. No Python import statement executes, so restricted_import never fires. From that handle, write-ups of the earlier March disclosure describe reaching native calls directly, up to and including system(). I am deliberately stopping at the handle rather than printing a chain; the record gives exactly that much and it is enough to understand the class.
The second one is documented in the fix commit itself. Commit fb2323b added a test named test_sandbox_escape_vulnerability_demonstration, marked xfail, which recovers the original __import__ by walking Python's object graph: take the empty tuple's class, list its subclasses, find one whose module still holds the real __builtins__, and pull the original import function out of it. After that step, the blocklist is guarding a door the attacker is no longer using. Both escapes defeat the same filter by never passing through it.
How did code end up in this sandbox at all?
The path matters as much as the code, because the dangerous fallback was never the default. CERT/CC's record for the March disclosures describes the trigger: an operator sets allow_code_execution=True or manually attaches the CodeInterpreterTool, and then Docker is unreachable. The tool checks Docker first. If it is there, model code runs in a container. If it is not, the tool silently falls back to SandboxPython in the host process.
# Shape of run_code_safety() before and after the fix
# (reconstructed from commit fb2323b's diff and message, not verbatim source)
# BEFORE: fail-open. Docker missing means a silent in-process fallback.
if self._check_docker_available():
return self.run_code_in_docker(code, libraries_used)
return self.run_code_in_restricted_sandbox(code)
# AFTER: fail-closed. Same check, loud failure.
if not self._check_docker_available():
raise RuntimeError("Docker is required for safe code execution "
"but is not available.")
There is also a nastier sibling. CVE-2026-2287, one of the March disclosures, covers Docker silently stopping mid-session: the same fallback fires even though nobody configured anything unsafe. So the recipe for reaching SandboxPython includes a runtime failure you do not control. I wrote about the same fail-open shape when the two CVSS 9.8 Cua and AutoAgent sandbox bugs landed on the same day, and the pattern keeps repeating: the isolation is real until the moment it is quietly skipped.
Is anyone actually affected today?
Here is the part the bare record does not tell you, and the part most summaries will get wrong. The fix is old. Commit fb2323b merged on 2026-03-15, and the repo changelog shows v1.11.0rc1, released the same day, was the first tagged release containing it. I verified by ancestry that current tags, including the latest 1.15.21, all contain the fix. Better: v1.14.0 in April removed CodeInterpreterTool entirely and deprecated the code execution parameters in favor of external sandboxes. If you pip install crewai today, the vulnerable code was never in your tree.
The realistic audience is narrower: projects pinned below 1.11.0rc1, forks and vendored copies of old crewai-tools, and anyone who read an old tutorial explaining that "the sandbox keeps you safe."
Why publish a CVE for code fixed in March? The record was reserved on 2026-04-06 and published 2026-09-13. The record does not say why, and I will not guess. What I can say is that it arrived with no coverage at all: I checked Hacker News via Algolia, the Dev.to API, and news search this morning and found nothing but mirrors. That is why this is a code-level walkthrough rather than a fourth rehash of a press release.
Could your sandbox have the same class of bug?
I wrote before that a happy-path demo proves almost nothing about tenant isolation, and the same rule applies to sandboxes. Three checks, derived from this case. Honest caveat: these are code-derived, not a lab-tested methodology.
First, find the name filters. Any string-membership check on module names, any import hook, any allowlist, is a speed bump, not a boundary.
$ grep -rn "BLOCKED_MODULES\|allowed_modules\|is not allowed" ./your-agent-src
$ grep -rni "allowlist" ./your-agent-src | grep -i "import\|module"
Second, find every fallback path. Ask what happens when the primary isolation is unavailable. A fallback that runs restricted code in-process converts a missing daemon into an execution path. The Cursor allowlist bypass I covered earlier is the same species: trusting a name instead of the resolved thing.
Third, ask what the code can reach without importing. If executed code can touch ctypes, cffi, gc.get_objects(), object.__subclasses__(), or sys.modules, an in-process filter cannot contain it. The record's last sentence is the design rule: a within-process sandbox must account for the complete runtime, which in practice means the boundary belongs outside the interpreter. Container, VM, or a hosted code runner, and the container must fail closed.
Should a string filter ever ship as a sandbox at all?
Here is where I expect disagreement, and I want it. My position: the fix is right, and it should be the norm. A sandbox that cannot guarantee isolation should raise, loudly, rather than hope nobody notices. A filtered fallback does not just protect less than it claims, it advertises a boundary that does not exist, which is worse than no boundary in a code review.
The counterargument is real, though. Many users do not have Docker installed, and a maintainer can reasonably say a restricted fallback with warnings is better than raw exec in the same process, better at least against the laziest scripts. So where is the line between "better than nothing" and "false advertising"? I lean toward: the moment the feature is described as a sandbox, it has to fail closed. But I can see the other side, and a defensible compromise is an honest name like run_code_unsandboxed_with_fewer_builtins.
The second argument is the score. Official: 8.1, local attack vector, high attack complexity, scope changed, confidentiality and integrity high. The reporter's own page lists 9.0, Critical. The gap comes from the local vector and the complexity adjustment, which encode "needs opt-in code execution plus a Docker failure." Is that the right lens when the thing being scored was, per the record itself, never a sandbox? I honestly do not know, and I would like to hear how you would score it.
What I would do this week
- If you pin CrewAI below 1.11.0rc1 or vendor an old
crewai-tools: upgrade. Current is 1.15.21 and the tool is gone. - If you ship any restricted-execution fallback of your own: delete it or make it raise. Silence is the bug.
- Re-read your sandbox assuming the import system is not the boundary. In Python, it never was.
Primary sources: CVE-2026-37008 on NVD · MITRE CNA record · GHSA-2q68-3cp7-72v9 · Fix commit fb2323b · CERT/CC VU#221883
Top comments (0)