DEV Community

Cover image for Implementation That Favors the AI, Review That Favors the Human---Writing a COW Filesystem from Scratch, Part 2
faliye
faliye

Posted on AI-assisted

Implementation That Favors the AI, Review That Favors the Human---Writing a COW Filesystem from Scratch, Part 2

In the last post, I decided to hand the keyboard over to the AI. But how to use the keyboard, what the key layout is, which input method to use: all of that is still unclear. To get a filesystem built sooner rather than later, I still need to make a little more effort.

Part 1: Decisions and the Blueprint

The implementer is an AI Agent. An Agent can write code toward a goal, but we lack a blueprint. A filesystem blueprint takes far more than five decisions. What goes into the self-contained unit (how many rooms)? How is the data format defined (how big is each room)? Are stripes variable (can the partitions be moved)? How is the journal handled (do we install a gate)? What is the first goal (fix the foundation first, or the roof)? All of these need decisions. Otherwise you may end up with the roof finished and the foundation not yet dug. We all know that Agents are superb executors when the goal is clear, but they are a bit weaker at planning complex tasks, and without a blueprint a task easily goes off track.

But filesystem design is extremely complex and involves a mountain of prerequisite knowledge. Even if I started searching through references day and night without sleep, I might have a first design drawing five years from now. Fortunately, the Agent bro sitting next to me is exceptionally learned, so I can ask him to make the design decisions and draw the blueprint. His weakness is that he is prone to going off the deep end: overconfident, charging down one path to the bitter end. So I have to guard against an awkward situation: he may not be trying to deceive me (he is simply hallucinating), yet I have no way to tell whether the information he gives me is true or false. What to do?
As the saying goes, three monks have no water to drink. Well then, let's get three Agent bros to work together anyway. Here is the design:

The conversation where the task starts is called the Main Agent. It only asks questions according to the plan. For example, according to the plan, today we should decide how many rooms to design.

This question is handed to one strong Agent (Strong A), who is responsible for collecting information and finding an answer that looks correct. Once he brings back this seemingly correct answer, the Main Agent notifies two more Agents, B the supporter and C the attacker, to start work.

Supporter B looks for non-overlapping supporting evidence for the answer, that is, other arguments that also support it. If he finds any, he reports them; otherwise he reports that he could not find any.

Opponent C looks for arguments that show the answer does not hold at all. If he finds any, he reports them; otherwise he reports that he found none. But the opponent must find strong evidence; weak evidence is treated as failing to find any.

After collecting these reports, the Main Agent begins independent verification. The Main Agent does not judge who has the better argument; it judges and checks whether the results are valid. If they are valid, the Main Agent summarizes these points and has Strong A repeat the steps above and reason through a second round. Note that in the second round the Main Agent's materials have grown, now including both supporting and opposing views.

At the same time, four points need attention:

  1. An Agent's conclusions are random in nature. So before Strong A starts work, he must set a target for himself. If the conclusion drifts away from that target, the conclusion is invalid. For example, suppose we want to build three rooms and the conclusion is that flooring is better. That conclusion cannot be accepted. The file header comment of an Agent's experiment program contains the following: the criterion (written in stone before the run, not to be changed after it), the failure clause (written in stone before the run), the reverse-acceptance clause, and the questions it cannot answer.
  2. Agents have a natural tendency to agree with the user, and this must be eliminated completely. Even if a proposal comes from the user, the decision must still go through the three-party process; if it does not, it is not recognized as approved. The user can demand that it pass, but the AI will, at lightspeed, leave the evidence in a warning and then submit.
  3. Agents from the same source tend to think alike. Although we bring in two AIs to participate, to avoid this tendency I also brought in a local model to join the thinking and the attacking. A local model is usually useless, but every now and then it has a flash of inspiration.
  4. The Main Agent's data sources are not trustworthy. Each Agent must measure the data independently. This prevents the Main Agent from hallucinating some data into the context and thereby biasing the other Agents' judgment.

Of course, considering that AIs often have no stopping condition, if three rounds of three-party confrontation produce no result, reasoning stops and the question is summarized and handed to me to decide.

With this process, my job becomes much simpler. I generally no longer need to re-check whether the materials are real. My attention is not spent on reading materials and checking for hallucinations; I spend it only on four things a machine cannot judge:

  • Is this test measuring the right thing?
  • Is this invariant itself correct?
  • Does the basis of this decision hold up?
  • Do I accept this trade-off? This is the second half of our first heading: review that favors the human. Make sure the evidence is sufficient and the reasoning is valid, and let the human make the call. And the cost? What's your superpower, Tokenman?

Part 2: Implementation and Code

Before discussing this topic, two observations:

  1. If, in this era, AI context has reached the million-token scale, the context window can even hold all the code and documentation of a medium-sized project at once.
  2. The best practices of the past had human programmers as their subject.

Based on the second point, let's re-examine past best practices: function names should be short, for loop nesting should be shallow, functions should be reusable and abstracted and encapsulated, and so on. These practices actually address the problem of limited human brain bandwidth: short function names, shallow for nesting, and reuse, abstraction and encapsulation are all meant to reduce the load on the human brain. Today, for those of us programming toward Markdown (prompts), does the length of a variable name matter to us? Does shallow for nesting matter to us? Do reuse, abstraction and encapsulation matter to us?

But note that these matter a great deal to the Agent.

When an Agent modifies a function, it greps, and function names and variable names are always the first coordinates. If a name alone can tell you what something is and what it does, that lowers the Agent's hallucination rate, because comments may be forgotten, but the name anchor gets used again and again.

Ten levels of for loops with no if inside are just a traversal to an Agent. An exhaustive match with 12 arms is counted for you by the compiler; 12 independent layers of if/else are 2^12 paths, which you cannot finish testing.

Abstraction and encapsulation are the biggest source of bugs in code. Their main benefit is making code easier for humans to maintain, but for an AI, modifying one place costs about the same as modifying ten.

Today the main force writing and reading code is the model. Sorting out logic, maintaining consistency and exhaustive coverage are exactly what machines are good at. A floor set by the limits of the human brain is far too low for a machine.

So our default attitude toward "best practices" is skepticism, not compliance. This does not mean they are all wrong; it means the reason behind each one has to be re-examined.

Four steps to audit an old rule:

  1. What problem was it originally meant to solve? If you can't say, drop it immediately.
  2. Does that problem still exist today? If it was solving "humans can't remember," "humans can't read it all," or "reviewers are too busy," it most likely no longer exists.
  3. Does it have a second reason, one that has nothing to do with humans? If so, keep it, and rewrite the reason to be that one.
  4. Can the rewritten reason be turned into a gate? If it can't, downgrade it to a suggestion.

There is only one criterion: does this rule make the code easier to verify mechanically, or does it only make the code easier for human eyes to skim? Keep the former; the latter can be dropped.
Here is a brief summary of some past best practices.

Popular practice Disposition How we write it
Names should be short, and the closer to the declaration the shorter Abolished No length limit
Loop counter called i, temporary variable called tmp Abolished Name it for what it is: stripe_index
Functions under 20 lines Kept, reason rewritten Cut by "one thing that can be verified on its own," not by line count
Nesting no deeper than three levels Relaxed No limit on depth; what is limited is the number of paths
Single exit Abolished Early returns freely allowed
Replace conditional branches with polymorphism Abolished for closed sets enum plus exhaustive match
DRY, avoid repetition Relaxed Repetition must be generated, never copied by hand
default: fallback is safer Abolished for closed sets Leave missed cases for the compiler to catch
One assertion per test Abolished One scenario per test; multiple assertions allowed
No more than 80 columns per line Left to tools rustfmt decides; don't shorten names for line width
Avoid premature optimization Downgraded to a suggestion —

Part 3: Practice in Rust

1. No abbreviations, except those registered in the abbreviation table

Without reading comments, without looking at call sites, without looking at the implementation, from the name alone you should be able to say what it is, what it does, and what it does not do.

No abbreviations: write cnt, idx, buf, tmp out in full. Domain abbreviations (lba, crc) may be used only once registered, and the registry is their sole authoritative definition.

No single letters: write i as stripe_index, the generic parameter T as Key, the lifetime 'a as 'journal, and Err(e) as Err(error).

Write preconditions into the name: write_node says nothing; append_verified_node_to_journal says three things: it appends rather than overwrites, the input has been verified, and it lands in the journal region.

Test names state the scenario and the expectation: crash_between_data_write_and_commit_keeps_previous_generation, not test_commit.

One concept, one name across the whole repository: it must not be called generation here and epoch there.

2. Semantics go in types first: types > names > comments

Take the following example. The same thing written three ways; when the semantics live in the type, an AI that writes it wrong gets an error.

// 1. Semantics in comments: the worst
// dim0: nationality  dim1: gender  dim2: age band
pop[i][j][k]

// 2. Semantics in names: criticized as verbose in the past, machines love it
number_of_people_in_japan[foreigner][woman][teenager_18_to_24]

// 3. Semantics in types: the best, a wrong dimension simply won't compile
number_of_people_in_japan[Nationality::USA][Gender::Woman][AgeBand::Teenager18To24]
Enter fullscreen mode Exit fullscreen mode

3. Branching: write every case out, never write _ =>

enum Medium {
    Rotational,
    SolidState,
    Zoned { zone_size_in_bytes: u32 },
}

fn node_size_in_bytes(medium: &Medium) -> u32 {
    match medium {  // the point is that there is no `_ =>` wildcard arm
        Medium::Rotational                   => 64 * 1024,
        Medium::SolidState                   => 16 * 1024,
        Medium::Zoned { zone_size_in_bytes } => (*zone_size_in_bytes).min(256 * 1024),
    }
}
Enter fullscreen mode Exit fullscreen mode

The key is the _ => that was not written. Add it and the code looks shorter and more "general," but it switches off exactly the compiler's exhaustiveness check. Later, when a new kind of medium is added, the program will quietly walk into the wildcard arm; the behavior will be wrong and nobody will raise an alarm.

The criterion is not "how many branches there are" but "if a case is missed, who finds out first": the compiler, or production.

A few related practices:

  • For encodings read from disk that may gain new values later, make - "unknown" an explicit member (Unrecognized(raw_code)), and match remains exhaustive.
  • Use enum for closed sets, not trait objects. A trait's default method is in fact also a kind of wildcard arm.
  • Don't abstract a trait with only one implementation. Don't abstract "to make swapping easier later."
  • For interactions between two cases, write an exhaustive match on a tuple. In singlefs's code we enforce the following:
fn geometry(candidate: Candidate, map_scope: MapScope) -> Geometry {
    match (candidate, map_scope) {
        (Candidate::LogicalIdentityWriteOrder, MapScope::AllPointers) => Geometry { .. },
        (Candidate::LogicalIdentityWriteOrder, MapScope::Code1Only)   => Geometry { .. },
        (Candidate::MixedPath,                 MapScope::AllPointers) => Geometry { .. },
        (Candidate::MixedPath,                 MapScope::Code1Only)   => Geometry { .. },
        (Candidate::ClassReuseTagged,          MapScope::AllPointers) => Geometry { .. },
        (Candidate::ClassReuseTagged,          MapScope::Code1Only)   => Geometry { .. },
    }
}
Enter fullscreen mode Exit fullscreen mode

Two of the rows actually have identical values, and most people would merge them. Here we do not merge them: the day a fourth candidate is added, the code will not compile until this spot is updated.

4. Types: make illegal states unwritable

#[derive(Clone, Copy, PartialEq, Eq)] pub struct LogicalAddress(pub u64);
#[derive(Clone, Copy, PartialEq, Eq)] pub struct PhysicalAddress(pub u64);
fn read_block(address: PhysicalAddress) -> Block;
// read_block(LogicalAddress(value)) simply won't compile

struct RawNode(Vec<u8>);       // just read from disk, unverified
struct VerifiedNode(Vec<u8>);  // checksum compared and passed
impl RawNode {
    fn verify(self, expected_checksum: Checksum) -> Result<VerifiedNode, CorruptBlock> { /* ... */ }
}
fn walk(node: &VerifiedNode) { /* ... */ }  // want to skip verification? the argument can't even be constructed
Enter fullscreen mode Exit fullscreen mode

The first snippet turns "passing the wrong address" from a runtime bug into a compile-time error. The second is even harsher: it makes "not yet verified" a type too, so "forgetting to verify" simply cannot be written. Verify once at the boundary, then trust the types inside, instead of re-checking at every layer.

A few smaller points:

  • Replace boolean parameters with enums. write(node, true) says nothing at the call site; write(node, Durability::Synced) does.
  • When constructing a struct, write out every field; don't use ..Default::default(). It is a wildcard arm over fields: add a new field, and every place that didn't write it quietly gets the default value.
  • Don't use as for lossy numeric conversions; use try_from and handle the failure. as truncates, drops signs and wraps around, all without any error.

Branching and types are in fact two sides of the same thing: one lets the compiler catch "a missed case," the other makes "an illegal combination" unwritable. Both move the checking from humans to machines.

5. Errors: recoverable ones go into types, invariant violations get assertions

I/O errors, on-disk data corruption and running out of space are recoverable, and go through Result. The error enum is divided by the decision the caller must make (not found, data corrupt, out of space, other I/O error), not expanded one by one by underlying cause.

At the public boundary of a library, don't use a unified error type like anyhow. It lumps "retry, report corruption, report out of space" into one blob, and is a wildcard arm on errors.

A broken invariant is a bug: assert immediately, don't carry bad state forward, and above all don't write it to disk.

Don't write unwrap(); write expect, with a message stating which invariant it relies on.

An expect in experimental code looks like this:

.expect("leaf number is not in the position table: the position table did not keep up after the split")
u64::try_from(distinct_groups.len()).expect("the number of distinct nodes fits in u64")
Enter fullscreen mode Exit fullscreen mode

Each expect is a named assertion. The day it blows up, the message says directly which assumption collapsed.

Longer names, more types, more complete branches. The compiler reports a few more errors. That is what being AI-friendly means on the implementation side.

Finally: What We Lost, and What We Gained

Wasted tokens. One question, three legs, three rounds, sometimes eight; one set of background material runs over two thousand lines.

Wasted time.The format design took a full three weeks, and not a single line of the filesystem itself has been written.

Wasted code. The experimental code may be very long, yet not a single line of implementation code has been written.

Wasted characters. Names written in full, branches written in full, every rejection accompanied by a next step.

What we gained:

  1. A conclusion that has not been verified will never appear in the code.
  2. The code may still be wrong, but when it is wrong, it will be caught.
  3. What I decide is on the record, and when I say something wrong, someone comes to strike it down.

It doesn't matter that I don't have the knowledge: the Agents come to me with evidence and options, and multiple choice is something I can do. Reviewing makes me happy.

I don't write the code, but when it is written wrong, it fails loudly at any time. Simple, linear, multi-branch code should be fun for an Agent to write, I imagine.

Top comments (0)