DEV Community

ZAKARIA KHCHICHE
ZAKARIA KHCHICHE

Posted on

Three silent bugs in Mistral AI's Python libraries

A tool schema that looks right. A validator that says "valid". A structured output that parses. In each case the code runs without an error, and the result is still wrong.

I audited the two Python libraries at the heart of Mistral AI's stack and found three bugs of that kind. All three are reproduced, fixed with regression tests and submitted upstream.

Context: the two libraries

  • mistral-common validates chat requests and turns them into tokens. vLLM and Hugging Face Transformers rely on it to serve Mistral models.
  • client-python is the official mistralai SDK. Its extra package turns Python functions into tools and Pydantic models into structured outputs.

Agents built on Mistral go through both on every call. A silent bug there spreads to every agent built on top.

Bug 1: optional tool parameters become mandatory

create_tool_call turns a Python function into a tool the model can call. With the idiomatic Pydantic style, a default value is lost:

def get_weather(
    city: Annotated[str, Field(description="City name")],
    unit: Annotated[str, Field(description="celsius or fahrenheit")] = "celsius",
) -> str: ...

create_tool_call(get_weather).function.parameters["required"]
# ['city', 'unit']   expected: ['city']
Enter fullscreen mode Exit fullscreen mode

The default lives in the function signature, not in the Field taken from Annotated, so it never reaches the schema. Tools are sent with strict=True: the model is forced to fill unit every time, and the developer's default is never used.

Reported in issue #632, with the fix and two tests ready on my fork.

Bug 2: a dict field that can never hold a key

chat.parse lets you ask for an answer that matches a Pydantic model. Before sending it, the SDK makes the schema strict by setting additionalProperties: false on every object.

The problem: for a dict[str, int] field, Pydantic already uses additionalProperties to describe the values. The SDK overwrites it.

stock: dict[str, int] Schema
Pydantic {"additionalProperties": {"type": "integer"}}
Sent to the API {"additionalProperties": false}

With the second schema, {"apples": 3} is invalid: the field can only be empty. The fix only adds the key when it is missing, which is what the OpenAI SDK already does.

Reported in issue #633, with the fix and a test. I could only check the client side: I did not test the API end to end.

Bug 3: a newline slips past the validator

mistral-common checks that function names match ^[a-zA-Z0-9_-]{1,64}$ and that tool call ids have 9 alphanumeric characters. The check uses re.match, and in Python $ also matches right before a final \n.

So "get_weather\n" and "abcDEF123\n" (10 characters) are accepted, and the invalid name ends up in the tokenized prompt. The fix is one word per call site: fullmatch instead of match.

Fixed in pull request #354 (issue #353).

What this means for teams shipping agents

  1. Check what you send, not just what you write. Print the schema your SDK actually sends to the model at least once. Bugs 1 and 2 are invisible in the Python code.
  2. Strict mode raises the stakes. A wrong schema in strict mode is not ignored: the model must follow it.
  3. re.match + $ is not a full match. Use re.fullmatch for validation, in your code too.
  4. Contribute back. Each fix is a few lines. Sending it upstream protects everyone on the same stack.

Building agents on Mistral? I'd like to hear which surprises you've hit.

Let's talk, or get your own agents audited: LinkedIn · Malt · Medium · Website

Want your team to build agents like these? I run a hands-on Copilot Studio training (in French, with Spar-x, Qualiopi-certified, eligible for OPCO funding in France): https://zakariakhchiche.github.io/formation-copilot-studio/ and a free AI Act article 4 kit: https://zakariakhchiche.github.io/kit-ai-act/

Top comments (2)

Collapse
 
grunzai profile image
schultzbehrnt9-jpg •

runs without an error and the result is still wrong is the whole category that eats my time. nice that u sent fixes upstream w tests instead of just writing it up. how did u find the optional param one, a real agent misbehaving or reading the code? (i build grunz, a coding agent on open models, and silent format stuff is most of what i chase)

Collapse
 
zakaria_khchiche_490919ed profile image
ZAKARIA KHCHICHE •

Reading the code. I audit the paths that turn Python into what the model actually receives (tool schemas, response formats, validation), then confirm each suspicion with a minimal script before writing a fix. For create_tool_call, the code took the Field out of Annotated and never looked at the signature default; a 6-line repro showed unit in required. Same for the dict bug: dump the schema actually sent and diff it against Pydantic's. For a coding agent, one cheap habit catches this whole category: snapshot the final tool and response schemas in your tests, so a library upgrade that changes them shows up as a diff.