DEV Community

Helgard
Helgard

Posted on

An AI hardware advisor that walks every path — and gets its math checked by the rules in Sanity

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

An agent that answers "what computer do I need to run this AI model?" — including "don't buy, rent" and
"you already have enough". It reads a Sanity dataset of criteria, not a catalog: laws (formulas for weights
memory, KV cache, speed ceiling, own-vs-rent breakeven), machine-checkable rules, 11 solution paths
(keep what you have · upgrade · used · one GPU · multi-GPU · unified memory · MoE experts in system RAM ·
CPU-only small model · smaller model · cloud API · rented GPU), run reports and reference hardware.
New models appear every week, so anything not in the base is looked up on the web by the base's own checklist.

Demo

Live — bring your own agent (no login, no keys): https://hw-advisor.helgardorlm.tech · MCP endpoint https://hw-advisor.helgardorlm.tech/mcp
Replay of all 27 measured runs: https://helgard-orlm.github.io/sanity-ai-hardware-advisor/

Code

https://github.com/helgard-orlm/sanity-ai-hardware-advisor (the public proxy with check_answer is in public_mcp/)

Sanity Project Details

Try it with your own agent (no keys, no login)

The demo is a public MCP endpoint, not a hosted chatbot: you bring the agent, the base brings the knowledge and the checking.
Add https://hw-advisor.helgardorlm.tech/mcp to ChatGPT (developer mode), Claude, Claude Code, Codex, Cursor or VS Code.
There are one-click buttons for Cursor and VS Code on the page. Or paste one line into any agent that can run commands:
"Read https://hw-advisor.helgardorlm.tech/setup.md and follow it."

Behind it is a ~230-line stdlib proxy. It passes the four Sanity Context MCP tools through unchanged and keeps the token
server-side. It puts today's date and the advisor's work order in front of initial_context. It adds one tool, check_answer,
which is the same validator the app uses: it recomputes every calculation with the law.formula stored in Sanity.
In the first outside test (Codex, its own subscription), check_answer rejected the first draft with 4 errors. The second draft passed:
5 calculations recomputed, 11/11 solution paths walked. A test from ChatGPT (developer mode) also went through to a passing check. Agents that can only browse get the same tools as plain HTTPS links
(/api/context, /api/query?q=…, /api/check), listed in /llms.txt.

Known limit: the check covers numbers, completeness, sources and identity. It does not cover whether a verdict makes sense. In one test an agent said yes to "MoE experts in RAM" for a dense model, and the check passed.

How I Used Sanity

Sanity Context MCP tools the agent calls: initial_context, groq_query, schema_explorer, array_field_reader (plus our check_answer). What it does with the content it retrieves — the structure is used, not just searched:

  1. Paths are data. The agent must give a yes/no/maybe with a number for every solutionPath document. Before this, the same model answered a 321B-MoE question with "4×H200" and never mentioned the one-GPU + large-RAM path.
  2. Rules are the search checklist. For any model/card/CPU that is not in the base, the app requires the fields that the base's rule documents test for that type (e.g. cpu.instructionSets, aiModel.variants.fileGb).
  3. Laws are recomputed by code. The agent writes every calculation with its inputs into a hidden check block; the app evaluates the law.formula from Sanity and sends the answer back if the number is wrong.
  4. The person's situation is stored, not remembered: country, budget, existing PC survive a change of mind.

The same code checks the base itself: every law variable and rule path must be a real schema field
(coverage check), and laws get property tests. That caught the filler model writing an own-vs-rent formula that
added electricity instead of subtracting it and billed rent for 24 h/day — 12.3 vs 110 months on the same inputs.

A real trace (GLM-5.3-Flash, not in the base) — open it in the replay

  1. GROQ over the base → no aiModel document → 8 web searches by the base's own checklist: the model card and config.json (layers, KV heads, head dim), file sizes, and the GPU fields the base's rules test (slots, bandwidth, compute capability).
  2. Laws from Sanity: weights 320 × 4 / 8 = 160 GB (NVFP4), KV for 8k context 22.5 GiB (flagged as rough for hybrid attention).
  3. Validator round 1 sent the answer back: the speed ceiling was written as 19.91 tok/s — code evaluated the law.formula from Sanity with the agent's own inputs and got 199.1 (a 10× slip); plus three entities missing fields that the base's rules check (e.g. variants.fileGb, slots).
  4. Round 2: no errors. All 11 solution paths get a verdict — multi-GPU = yes, MoE-in-RAM / unified memory = maybe, cloud = no (the person ruled it out) — and the build is 3 × RTX PRO 6000 (288 GB) with the numbers above.

Other catches by the same recomputation in the measured run: KV cache 0.949 → 6.75 GiB (A4), monthly electricity cost 0.675 → 2.025 (A5).

Measured, honestly

Blind grader (a different model, sees only the dialogue and a frozen checklist), 9 scenarios × 3 runs,
78 checklist items, same base state for all arms:

version score median time
v2: base + instructions 47/78 62 s
v3: + solution paths, validator 53/78 104 s
v3.1 via Sanity Context MCP: + rules-as-checklist, recomputed laws, dated prices 59/78 181 s

What this does not show: that structured retrieval beats plain text for reading. The same 200 documents
pasted as text notes scored the same as v3 (53/78) — at this size the model reads either. Where structure
mattered in our runs is checking: fewer first answers rejected by the validator with structured queries
(8/27 vs 14/27), and the gains of v3.1 are exactly in the scenarios where code recomputes Sanity's laws
(KV for 8 users 7/9 → 9/9, own vs rent 5/12 → 7/12, "don't buy" 4/6 → 6/6, v2 → v3.1). The price: answers take ~3× longer.
An earlier full run of v3.1 scored 62/78 with two bugs in my own checker (nested fields, European number format) that sent
correct answers back for rework; after fixing them the clean run scored 59/78 — with N=3 per scenario, 59 vs 62 is noise.

How it was built

Architect method: a risk catalog (16 classes of agent failure seen in earlier sessions) → acceptance scenarios
written before changes → build → blind grading. The base was filled by a model (GPT Luna) from an empty schema and
two prompts; I never wrote data by hand. An independent judge model watched the session and corrected the method
several times (frozen checklist, snapshot baseline, text control).

Known limits

  • Answers are slow (minutes) because the validator can send them back up to twice.
  • Ratings N=3 per scenario; differences of 1–2 points are noise.
  • "unknown after search" is allowed and can be abused for vague entities.
  • In the measured run 12 of 27 answers came out in Russian to English questions (an account language setting leaked in); the app now checks the answer's language in code and sends it back (3/3 English on a re-test).

Top comments (0)