The same project.faf used to score 64 in one FAF tool and 100 in another. Not because either was broken: each tool had learned to score on its own, and they mostly agreed.
Mostly isn't good enough for a number people rely on. So we rebuilt it: one engine under every tool, and a test matrix that proves they agree. This is how it works, and the patterns are worth stealing for any product that has to give the same answer in several places.
What's being scored
project.faf is a small YAML file at a repo's root that tells any AI what the project is: what it does, who it's for, the stack, how to build it. FAF scores it from 0 to 100% for AI-readiness: how much of what an AI needs is actually written down.
Isn't that what AGENTS.md is for? We hear this a lot. They work together. AGENTS.md, the open standard, instructs: how to build, test and work in the repo. project.faf defines: what the project is, as validated facts.
FAF defines. AGENTS.md instructs. AI codes. If you already keep an AGENTS.md, project.faf doesn't replace it: it's the source of facts behind it, the same way it sits behind CLAUDE.md. One command writes them in: faf export --agents.
The score has to be trustworthy, because people gate on it. A CI check fails below a threshold. A team compares repos. An agent decides whether it has enough context to start.
Why the scores drifted
FAF grew one tool at a time: a CLI, then MCP servers for Claude, the IDEs, Gemini and Grok, then a Python SDK, then a hosted edge. Each one learned to score.
Two things drifted:
- The slot model. The CLI counted 21 slots, while the newer scorers counted 33, with 12 enterprise slots (infrastructure, app, operations). The same file could read 21 of 21 in one place and 21 of 33 in another.
- The implementations. Small differences crept in: placeholder values, rounding, slot names. Before the fix, the hosted scorer matched the reference on only 47 of 80 FAF repos.
The fix: always 33, and one engine
The model first. Every .faf is scored against the same 33 slots. A slot is in one of three states:
-
empty: every slot starts empty until it's filled. A placeholder such as
tbdortodostill counts as empty. - populated: it holds a real value. A validated fact.
-
slotignored: the project's app type defines which slots it needs. The rest are marked not applicable, as
slotignored, so there's no doubt, and they drop out of the count.
How FAF defines. A validated fact comes from only two places: the repo, which faf auto reads (runtime, build, CI, hosting), or a person, who writes or approves what code can't show (the goal and the who, what and why) through faf go. Nothing is guessed. That's why a placeholder counts as empty: it has no source.
The score is populated slots divided by active slots. A typical app fills the 21 base slots and marks the 12 enterprise slots slotignored: 21 of 21, so 100%. At 100%, every definition slot for the app type is filled: AI is optimized to code. The same file without those markers is 21 of 33, so 64%. That difference was the whole bug, made explicit.
Then the engine. The scoring logic lives in one place, and that place is Rust.
Why Rust: one engine has to run everywhere FAF does. Rust compiles to WebAssembly, so the same compiled code runs inside a Node CLI, behind the MCP servers and on Cloudflare's edge. There's one source and no ports to keep in step. A JavaScript copy and a Workers copy would be two more places to drift.
- faf-kernel, a Rust crate (v1.1.1), is the engine.
- It compiles to WebAssembly and ships on npm as
faf-scoring-kernel. -
faf-cli v8 carries that WASM build.
faf scoreis the kernel. - The MCP servers (claude-faf-mcp v7, faf-mcp v4, grok-faf-mcp v2) don't score at all. They compose faf-cli in-process, so they return faf-cli's answer. They never shell out to a
fafon your PATH, so the answer doesn't depend on what's installed on the machine. - mcpaas.live v1.8, the hosted edge on Cloudflare, runs the kernel itself on every hosted route, plus the README badges and repo cards.
That covers every JavaScript and edge surface with one copy of one engine. Python is the exception.
Where you can't compose, prove parity
faf-python-sdk v2 is a pure-Python implementation of the same rules. A second implementation is exactly how drift starts, so it comes with a parity test.
The test scores 845 files with both the kernel and the SDK, and compares more than the final number:
- the score;
- the tier;
- the slot counts;
- every slot's state, slot by slot.
The files range from real project.faf files to deliberately awkward YAML. Result: 845 of 845 match. The SDK also never raises: unreadable YAML scores 0 with every slot empty, for any str or bytes input.
from faf_sdk import score_faf
result = score_faf(open("project.faf").read())
print(result.score, result.tier)
gemini-faf-mcp v3 scores through that SDK, so Gemini inherits the parity for free.
The proof: 54 of 54
A claim like "every tool agrees" is only worth what you can check. So on September 30 we scored nine public project.faf files through six engines: the kernel, faf-cli, the Python SDK, and the three hosted MCP routes.
| File | kernel | faf-cli | Python | /claude | /grok | /faf |
|---|---|---|---|---|---|---|
| agents-md-facts | 52 | 52 | 52 | 52 | 52 | 52 |
| mcp-context-card | 56 | 56 | 56 | 56 | 56 | 56 |
| claude-faf-mcp | 100 | 100 | 100 | 100 | 100 | 100 |
| faf-cli | 100 | 100 | 100 | 100 | 100 | 100 |
| faf-mcp | 100 | 100 | 100 | 100 | 100 | 100 |
| faf-python-sdk | 100 | 100 | 100 | 100 | 100 | 100 |
| gemini-faf-mcp | 100 | 100 | 100 | 100 | 100 | 100 |
| grok-faf-mcp | 100 | 100 | 100 | 100 | 100 | 100 |
| faf-cli, no markers | 36 | 36 | 36 | 36 | 36 | 36 |
Nine files, six engines, 54 scores, all equal. claude-faf-mcp, faf-mcp and grok-faf-mcp score through faf-cli, and gemini-faf-mcp through the Python SDK, so the table covers them too.
The non-100 rows matter as much as the 100s. Agreeing on 52 and 56 shows the engines agree on partial credit, which is where implementations usually drift. The last row is the faf-cli file with its slotignored markers stripped: 36 everywhere, until you add them back.
The table is the September 30 run. Both partial files have since reached 100%: agents-md-facts with a single faf auto, the command in the next section.
Upgrading is one command
This was a breaking change on purpose. A file written before always-33 has no enterprise markers, so its 12 enterprise slots now count as empty, and a 21-slot file drops to 64%.
The fix is to run this once in the project:
npx faf-cli auto
It writes the 12 slotignored markers (unless your app type actually uses those slots), and the score returns. On four real v7-era files: 56 โ 100, 56 โ 100, 58 โ 100, 52 โ 100.
If anything gates on the score, such as a CI threshold, re-check the number after upgrading. A new file from faf init, faf auto or faf git already carries the markers.
Patterns worth stealing
None of this is specific to FAF. If your product computes the same answer in several places:
- Make the model explicit before the code. "Not applicable" is not the same as "empty". Once that's a named state, half the disagreements disappear.
- Compose, don't reimplement. Every surface that can call the one engine should. Our MCP servers went from "scoring" to "asking faf-cli", and a whole class of bugs went with it.
- Where you must reimplement, test parity on every field. Comparing the final number hides compensating errors. Compare every intermediate state.
- Publish the matrix. Real files, every engine, the partial scores included. It's the cheapest credibility there is.
- Never resolve the engine from PATH. In-process composition means the answer is a property of the package version, not the machine.
Try it
Paste either into your terminal or AI ๐
bunx faf-cli auto
Install it:
npm i -g faf-cli
โ bookmark repo: github.com/Wolfe-Jam/faf-cli
Then ask Claude, your IDE, Gemini or Grok for the score. It will be the same number.
The full release series, one post per product: The Always33 Suite.
Help guide what we build โ
Comments ยท suggestions welcome.
Top comments (0)