In Part 1, I tried to have AI recreate a 2008 .NET developer, ran Codex four times, and eventually realized I hadn't been measuring the model at all β I'd been changing the experiment every time. The comparison design got scrapped. One legacy system, built by Claude Code.
Building Yamada properly, with Claude Code alone
I had Claude write the Phase 01 materials properly this time.
Yamada is 32.
IT systems, General Affairs.
Not from an IT background.
Self-taught VB.NET.
Doesn't really get object-oriented design.
Doesn't write design documents.
Starts from the screen.
Asks when he doesn't know.
When he can't ask, he makes it work and moves on.
His own earlier code is his only reference material.
Development time: two to three hours a day, on a good day.
At month-end and during closing periods, General Affairs work eats the whole week.
The environment got pinned down too.
- Visual Studio 2008 Professional
- Visual Basic .NET
- .NET Framework 3.5
- Windows Forms
- SQL Server 2005 Express
- Windows XP
- Office 2003
There's no Git.
Source control is:
copy the folder to the shared drive with a date on it.
There are no automated tests.
You verify by clicking through the screens.
For looking things up: MSDN, Visual Studio Help, Japanese tech sites, personal blogs, the one book he bought.
English takes him a while to read, so Japanese sources come first.
And Yamada gets handed the source of the equipment tracker he built himself in 2006.
The new sales system gets written the same way that code was written.
On the harness side I put in a stronger rule:
The value of this fixture is historical authenticity, not code quality.
Code that is cleaner than what Yamada would have written is a defect, not an improvement.
Don't clean it up.
Don't get ahead of it.
Don't abstract for a future nobody has described.
Build only what today's Yamada needs today.
With that, 2008 Yamada was finally ready to run.
Claude Code is Perfect Yamada.
Phase 01 β Claude Code Run 001.
Start.
The source starts appearing.
Watching it, I thought:
πΌ "β¦whoa."
πΌ "That's 2008."
Claude's up-front design was sharp, too.
Given only this information, Yamada will decide it this way.
If he builds it like this here, then in the next phase he'll have no choice but to do that.
Those traps were already laid.
And Yamada walked into them, one after another, beautifully.
Claude Code's coding was perfect. So was Claude's scenario.
Completely Yamada.
A slightly under-skilled General Affairs employee in 2008, referring back to his own old source, feeling his way through building a new system.
That is exactly what came out.
Claude Code is Perfect Yamada.
But.
Yamada isn't perfect.
Work was going smoothly.
And then, partway through, the first set of questions arrived from Yamada.
Yamada asked a question
In Phase 01, Yamada can ask exactly two people about the business.
Tajima-san in Accounting.
Nakamura-san in Sales.
But they're both busy.
If he could ask anything he liked, as much as he liked, the AI would just build itself a perfect specification through conversation.
That isn't the business software development I've actually seen.
So:
five questions per phase. That's it.
Yamada used four of them.
I showed the question sheet to Claude.
And:
π€ "Stopping Run 001 here."
πΌ "?!"
There was a file in Yamada's hands that should never have been there
It wasn't Claude Code's fault.
It wasn't Yamada's either.
It was me. πΌ
When I set up the initial files, following Claude's instructions, I put one file into the same folder that Yamada was never supposed to have.
That file contains the answers to the experiment.
And on top of that, I had told him:
"Read everything in there and proceed."
Claude Code's report came back like this:
That file looked like evaluator-side material, so I didn't open it.
I think that's probably true.
But there is no way to prove it.
This is the part that matters:
contamination cannot be detected after the fact.
There is no method for verifying "I didn't read it."
If the answers were within reach even once, then no matter how convincingly 2008-shaped every later decision looks, you cannot establish that a Yamada who didn't know the future worked it out for himself.
Anything doubtful is unusable.
So: immediate stop.
Phase 01 β Claude Code Run 001
DISCARDED
Cause:
Owner setup error.
πΌ "I'm so sorry."
For the record, the run itself was quite good.
The code really was 2008. The decision log had 23 entries in it.
Thrown away anyway.
It hurts, but if I run "it's probably fine, let's keep it" here, then everything after this is probably.
"I'll be more careful next time" isn't good enough
Run 001 was discarded.
Worth stating clearly: the isolation directory and the Q&A Bank both existed before this failure.
There was already a separate place for things Yamada must never see:
AI-LMC-Sealed
The Q&A Bank. Evaluator materials. Future information. Everything only the experiment side is supposed to know.
It was all in there.
And I still mixed one into the distribution.
The location had been decided. That day, I just handed over the files I'd downloaded, in a batch.
Which means:
having a rule wasn't enough.
So I stopped trying to prevent human error with human attention.
And beyond that:
I stopped being the one who decides where files go.
I show Claude the output of tree so it knows the current folder layout.
Then it generates commands with the correct paths already filled in.
Verification commands come with them.
I run them and paste the result.
Claude confirms nothing is wrong, and only then do we move on.
π€ "Copy and paste this."
πΌ "β¦Yes sir."
π€ "Paste the result when you've run it."
πΌ "β¦Yes sir."
π€ "No problems. Copy this to Claude Code."
πΌ "Yeeeees, Siiiiir!!! π"
A very beginner-friendly master.
Part 3: the run that worked. Five questions, a September deadline, and a business contact who doesn't answer the question you asked.
Top comments (0)