DEV Community

Sarvar Nadaf
Sarvar Nadaf Subscriber

Posted on AI-assisted

Which AWS limit is actually current? An agent that proves it, 32 vs 5 vs 16

Sanity Challenge Path One Submission

This is a submission for the Sanity Challenge, Path One: Ship an Agent That Queries Real Content

What I Built

Which AWS quota is actually current? You copied a default vCPU limit off the AWS docs, shipped it, and it broke in production. I have done this. The number was stale, and nothing warned me.

Here is what makes it nasty. For one AWS fact, three official pages can each give you a different number. An old User Guide says one thing. The Service Quotas console says another. A pricing page says a third. All look official. None of them tells you which is live today.

So I built an agent that answers "which AWS value is current?" and then proves it. It reads typed facts from a Sanity Context Knowledge Base, picks the winner with a deterministic rule the model is not allowed to override, and shows both the current value and the one it replaced, each with its source. Then it does the part most content agents skip: it calls the live AWS API (read-only) to check whether even the reconciled record is still true. On the EC2 vCPU quota it turns up three different numbers, docs 32, record 5, live 16, and puts all three on the screen.

Stack: Amazon Nova Pro on Bedrock, Strands Agents (Python), a Sanity Context Knowledge Base read over the hosted Context MCP, and a read-only live AWS cross-check.

Why a keyword search returns the wrong AWS value

The challenge sets a hard bar: "the strongest submissions show an agent that only works because the content was structured. If a keyword search would have gotten you the same answer, aim higher."

Fair. So I built a fact where keyword search gets it wrong on purpose. Real one too: the max IOPS per volume for an EBS general purpose SSD. The current number is 80,000 (gp3, from the EBS console). The old number is 16,000, the legacy gp2 ceiling, still sitting in an older SSD guide. That old guide repeats the exact phrase you would type into a search box: "maximum IOPS per volume for a general purpose SSD EBS volume." It is stale and wordy, so keyword scoring loves it.

Watch it pick the wrong answer. Same content, real TF-IDF, the question a person would actually ask:

Keyword/TF-IDF baseline for: "maximum IOPS per volume general purpose SSD"

  0.8626  EBS/limit: Maximum IOPS per volume for a general purpose SSD ... = 16000   <-- WRONG (old)
  0.2026  EBS/limit: Maximum provisioned IOPS per gp3 volume            = 80000   <-- right, ranked lower
  0.0218  S3/price:  S3 Standard storage, first 50 TB / month           = 0.023
  0.0185  EC2/quota: Running On-Demand Standard ...                      = 5
  ...
Enter fullscreen mode Exit fullscreen mode

The stale 16,000 wins by a mile. Keyword search hands you the wrong number and sounds sure of it.

Keyword search ranks 16000 first; the agent reconciles to 80000

Demo

Run the credential-free core yourself in under a minute (public dataset over anonymous GROQ, local reconcile, read-only live AWS check, no model or token needed):

# 1. clone + install
git clone https://github.com/simplynadaf/aws-source-of-truth-agent
cd aws-source-of-truth-agent
pip install -r requirements.txt

# 2. the credential-free path (public dataset + local reconcile + live AWS check)
python -m agent.reconcile_offline --service EBS --type limit --region us-east-1
# -> Verdict: 80000 IOPS (serviceQuotasConsole wins over officialDocs)

# 3. the keyword control that returns the WRONG answer:
python -m agent.baseline "maximum IOPS per volume general purpose SSD"
Enter fullscreen mode Exit fullscreen mode

For the full LLM agent through the Context MCP, copy .env.example to .env and add a Bedrock region plus a Sanity Context Viewer token (the README walks through it). There is a JSON mode too, for piping into something else:

python -m agent.ask --json "What is the maximum IOPS per volume for a gp3 EBS volume?"
Enter fullscreen mode Exit fullscreen mode

The real run, Nova Pro, us-east-1, read-only throughout:

Fact Reconciled (from the record) Superseded Live AWS Result
EC2 On-Demand Standard vCPU quota 5 (console) 32 (old user guide) 16 DRIFT
EBS gp3 max IOPS per volume 80,000 (console) 16,000 (old SSD guide) unavailable trusted (no such live quota)
S3 Standard $/GB-mo 0.023 (pricing page) 0.021 (stale blog) 0.023 AGREE
RDS PostgreSQL oldest major 13 (release notes) 11 (old tutorial) 11 DRIFT
Lambda concurrent executions 1000 (dev guide) n/a unavailable trusted
Graviton4 (R8g) availability Available (instance types) n/a Available AGREE

Look at the EC2 row. Docs say 32. The reconciled record says 5. The live account says 16. Three numbers for one quota, and the agent shows all three with their sources instead of picking one and hoping. The drift is not a failure. It is the honest answer: this is the current record, and here is where reality has already moved past it.

The agent's cited EC2 answer with the live-drift line and tool trail

Code

🌊 AWS Source of Truth

When your AWS docs, pricing page, and Service Quotas console disagree, an agent that knows which one is telling the truth, then checks the live API to see if even that record has drifted.

Sanity Challenge Sanity Context AWS Nova Pro Strands

Live Demo Read the Article

Stars Forks Issues


Sanity project id: 0q5ohtvv · ⭐ If a stale AWS number has ever bitten you in production, give this a star.

The Problem β€’ Why Search Fails β€’ How Structure Fixes It β€’ The Twist β€’ Getting Started β€’ FAQ

πŸ“– Table of Contents

πŸ€” The Problem

You copied a limit straight out of the AWS docs…

Full source, plus the credential-free path judges can run with no token.

How I Used Sanity

What I pointed Sanity Context at: my own Sanity content. Every fact is a typed awsFact document in the production dataset, not a wall of prose:

{
  "_type": "awsFact",
  "service": "EBS", "factType": "limit", "region": "us-east-1",
  "key": "Maximum provisioned IOPS per gp3 volume",
  "currentValue": "80000", "unit": "IOPS",
  "effectiveDate": "2026-01-15",
  "source": { "name": "EBS gp3 volume limits (Service Quotas / EBS console)",
              "kind": "serviceQuotasConsole", "url": "..." }
}
// ...plus a separate record for the old 16,000 (kind: officialDocs, 2020).
Enter fullscreen mode Exit fullscreen mode

source.kind and effectiveDate are real fields, so the rule that picks the winner is dull and readable:

highest source precedence wins (console and pricing page beat changelog, which beats official docs, which beats a blog), and the newest effectiveDate breaks ties.

Prose cannot do this. The schema is the whole trick.

The Knowledge Base: a Sanity Context Knowledge Base indexes these facts. The build reads them ahead of time and writes short cited entries. Where two records fight, the entry keeps both numbers and both sources next to each other.

Which Context tools the agent used: the agent pulls facts over the hosted Context MCP with knowledge_base_search (ranked lookup) then knowledge_base_read (full entry). That is the proof the answer came from Sanity and not from the model's memory. (It can also fall back to anonymous GROQ over the public dataset for the no-token path.)

The Knowledge Base in the Sanity Context dashboard

The ebs/quotas_and_limits entry: current 80,000 vs superseded 16,000, with sources

What the agent does with what it retrieves: Nova Pro runs the tools and writes the sentence, but it does not get to decide the number. A plain function reconciles the value, and a guard checks the model's answer against it. If they disagree, the guard throws out the model's prose and ships the deterministic answer instead. On one EBS run Nova muddled its own wording, the guard caught it, and swapped in the correct answer with no help from me. The model cannot invent the number even if it tries.

The tool trail proves the path through Sanity on every question:

- [Sanity Context MCP] knowledge_base_search('EC2 quota us-east-1 ...') -> top='ec2/quotas_and_limits'; knowledge_base_read(['ec2/quotas_and_limits'])
- fetch_candidate_facts(service='EC2', fact_type='quota', region='us-east-1') -> 2 rows
- reconcile_facts(n=2) -> current=5
- verify_live(EC2/quota) -> drift
Enter fullscreen mode Exit fullscreen mode

Then: what if even the reconciled record is stale? (live AWS drift)

Reconciling the sources gives you the best answer the documents can offer. But documents rot. So the agent does one more thing a pure content agent will not: it calls the live AWS API, read-only, and asks whether the reconciled value is still true right now. Service Quotas, the Price List API, EC2, RDS. The results are in the Demo table above. The live layer is additive and clearly labelled; it is not part of Sanity and it is account-specific.

What didn't work (the honest part)

  • I planned to screenshot a resolved conflict in the dashboard. The build never gave me one. The Knowledge Base read the typed records, reconciled the 32-vs-5 fight into a clean entry with both numbers cited, and left the Issues queue empty. That is the KB doing its job, but it means there is no "I resolved an Issue" screenshot to show, so I am not claiming one.
  • Nova dropped a tool argument on me. fetch_candidate_facts returned a row, then Nova passed an empty string to reconcile_facts, which saw zero facts. The guard fail-closed to "Not verified" instead of guessing. I fixed it by caching the last real fetch server-side, so the facts always come from Sanity, never from whatever the model relayed.
  • A wrong factType guess used to lose a real fact. Nova guessed instanceType when the field value was regionalAvailability, the typed filter matched nothing, and the fact vanished. Now a zero-row typed fetch retries on service and region alone.
  • The live check is not part of Sanity, and it is account-specific. A judge on a different AWS account will see different live numbers. That layer is additive and labelled as such. Two facts (EBS max IOPS, Lambda) have no matching Service Quotas entry, so the agent says "unavailable" rather than fake a check. The seed is labelled demo data at the top.

Reusing the pattern beyond AWS

Drop the AWS parts and the shape is generic: type your sources, index them into a Knowledge Base, reconcile by precedence and date, guard the answer so the model cannot fake the winner, and check a live system of record when one exists. It fits API version support, pricing, compliance clauses, legal terms. Anywhere the real question is "which version is current?"

The thing I took away: the win here was not a smarter model. It was structured content, plus the nerve to say "the record says 5, but reality says 16."

Sanity Project Details

If your docs and your console have ever disagreed with each other, tell me which number bit you.

Top comments (7)

Collapse
 
salman_khan_c31307505285e profile image
Salmankhan •

Much better way to use Ai and limitations were sorted out in better way.

Picked as gem
Collapse
 
sarvar_04 profile image
Sarvar Nadaf •

Exactly πŸ’―

Collapse
 
micheypico profile image
Micheal Heypico •

This matches what we see operating a model-routing layer (32 models, one key at heypico.ai): the deterministic scaffolding around the LLM is what makes multi-model setups viable. When a provider throttles mid-task, the state machine decides retry vs failover vs error β€” the LLM can't make that call reliably. Debugging a 'flaky agent' is usually debugging a missing state machine around a fine model.

Collapse
 
sinarezaei profile image
Sina Rezaei •

This is a really solid approach to a problem that looks simple until you actually have to trust the answer.

I especially like that the model is not allowed to decide which number is correct. The flow is much more interesting: structured facts β†’ source precedence β†’ deterministic reconciliation β†’ model response β†’ guard β†’ live AWS check.

The 32 vs 5 vs 16 example makes the point really well. Instead of hiding the conflict and returning one confident number, the system exposes the disagreement and then checks whether the reconciled value has drifted from the live environment.

That pattern is useful far beyond AWS. I can see the same architecture working for API versions, pricing, compliance rules, dependency versions, or even internal company policies where the question is basically: β€œWhich value is actually current?”

For me, the strongest part isn't the AI agent itself. It's the architecture around the agent that limits what the AI is allowed to claim. That's a much more practical way to build reliable AI systems.

Collapse
 
rulestack profile image
Rulestack •

Ours was Claude Code's limit on MCP tool output. The docs page gives 25,000 tokens. When we measured it in September on v2.1.273, the token count only ran once a result was already long in characters (somewhere between 45,000 and 52,000), so 24,000 characters of CJK text went in uncounted and grew the next request by 49,964 tokens, about twice the documented figure. Unlike your 16,000 row, the page wasn't out of date; it just doesn't say the limit is only checked once a result passes a certain length in characters.

Collapse
 
mrsaynothing profile image
Mr Say Nothing •

The rule the model is not allowed to override is the whole design β€” everything else is decoration. An agent that can argue with its own sanity layer will eventually win that argument. Showing the replaced value next to the live one is the other half: a number without its predecessor is unfalsifiable. One edge worth poking: quota changes propagate unevenly across regions, so the live API can lag the console for a while. Does the agent timestamp its proof and flag API-vs-console disagreement instead of picking a winner?

Collapse
 
arhancanli profile image
Arhan Canli •

The TF-IDF control that returns 16,000 is a great way to meet the "would keyword search get it?" bar: it shows the failure instead of asserting it.

One thing worth separating in the EC2 row, because it changes what DRIFT means. The live Service Quotas value for an account is the applied quota, which includes any increase that account was granted, so 16 vs 5 can be "this account asked for more" rather than "the record is stale". Service Quotas exposes both sides: GetServiceQuota gives the applied value and GetAWSDefaultServiceQuota gives the default. Comparing the record to the default tells you whether the docs are out of date; comparing the default to the applied value tells you about this account. Mixing them into one DRIFT flag will raise false alarms on any account that has had an increase.

The RDS row has a similar twist: describe-db-engine-versions on a live account can still list majors that are past standard support but in extended support, so a live 11 doesn't necessarily contradict a record saying 13 is the oldest standard-support major. A field for which support tier each number refers to would make that row read as AGREE-on-different-questions rather than drift.