DEV Community

Cover image for Codex Best Practices: Correct Your AI Twice, Then It Writes the Manual Itself
HIROKI II
HIROKI II

Posted on AI-assisted

Codex Best Practices: Correct Your AI Twice, Then It Writes the Manual Itself

Last week you corrected your AI assistant: the test command is npm test, not npm run test. It fixed it. Today it wrote npm run test again, and the test suite failed again.

That is not you failing at prompting. The rule lived in last week's prompt, and this week you didn't carry it over. The official Codex best practices guide (I verified the October 2, 2026 version) treats this as a systems problem: stop repeating instructions, start depositing them.

The commands mentioned below come from the official docs and were not run for this article; check the current docs before using them.

Try it in any chat window first

You don't need Codex installed. Take the classic vague hand-off — "fix the login page error" — and rewrite it with the official four-part template:

Goal: The login page returns a 500 error on submit; make it work again.
Context: Stack trace is in error.log; login logic lives in src/login/; test command is npm test.
Constraints: Do not touch the payments module; follow the existing code style; no new dependencies.
Done when: npm test passes fully, login redirects to the home page, and the error no longer reproduces.
Enter fullscreen mode Exit fullscreen mode

The one-liner version lets the model decide scope and endpoints. The four-part version tells it where the boundaries and the finish line are. The official docs say it directly: this helps the agent make fewer assumptions and produce work that is easier to review.

You just learned something concrete: say what "done" means before you hand over the task.

Second experiment, zero setup: when a task is too fuzzy to describe, flip the interview and let the model question you:

I want to build a small weekly-report tool for my team, but the requirements are fuzzy.
Do not start yet. Ask me 5 questions, challenge my assumptions,
then turn my answers into a concrete requirements list.
Enter fullscreen mode Exit fullscreen mode

It asks, you answer, and you get a requirements list you had not fully thought through.

The ladder Codex builds in

Rung 1 — Plan mode. Press /plan or Shift+Tab: Codex gathers context, asks clarifying questions, and drafts a plan before it touches code. That is experiment two, built in.

Rung 2 — AGENTS.md. A project manual that loads into every session. /init scaffolds it; you edit it into your real conventions: how to run, how to test, what is off-limits, what done means. Two official disciplines matter: keep it short and accurate rather than long and vague, and when the same mistake happens twice, have Codex run a retrospective and update AGENTS.md. It layers across ~/.codex (personal), the repo root (team), and subdirectories (local) — the closest file wins.

Rung 3 — Configure once. Personal defaults in ~/.codex/config.toml, repo behavior in .codex/config.toml, CLI flags only for one-offs. Tighten the two safety knobs first: approval mode (when it asks you) and sandbox mode (what it can read and write); loosen only after trust. The official guide makes a blunt point: many "AI quality problems" are actually setup problems — wrong directory, missing write access, wrong default model.

Rung 4 — Close the loop. Don't stop at "it says it's done." Have it write tests, run them, pass lint, review the diff — then you look. /review covers PR-style reviews, uncommitted changes, or a single commit. The guide includes a heavyweight data point: "At OpenAI, Codex reviews 100% of PRs."

Rung 5 — Package repetition. Paste the same external context every time? Bring tools in with MCP — one or two that remove a real manual loop, not twenty. Reusing the same prompt? Turn it into a skill that does one job. Workflow stable enough to leave alone? Schedule it. The official one-liner: "skills define the method and scheduled tasks define the schedule" — with the discipline attached: never schedule a workflow you haven't run manually until it is stable.

Run it once

  1. Enter your project, start Codex, and run /init to scaffold AGENTS.md.
  2. Replace the placeholders with your real conventions: how to run, how to test, what is off-limits, what done means.
  3. Hand it a small task with zero explanation and watch whether it follows AGENTS.md.
  4. When it errs, ask for a retrospective, deposit the lesson into AGENTS.md, reopen the session, and verify again.

What you get is not one completed task but a project manual that takes effect automatically. It does not guarantee the AI never errs — but the same mistake will not need a third correction.

Skip for now

  • Don't wire up a pile of MCP tools; start with the single one that saves the most work.
  • Don't schedule workflows you haven't run manually until they are stable.
  • Don't open full permissions; build trust inside the default sandbox first.
  • Don't let parallel tasks share the same files; use Git worktrees when you parallelize.
  • Don't run one chat per whole project; one coherent task per chat, and /compact when it grows.

Sources: Codex best practices guide, quickstart, AGENTS.md, config reference, and the merged manual.

Top comments (0)