I’ve stopped asking one AI to do everything. Claude plans the work and Codex writes the code. Both got better at their jobs once they weren’t doing each other’s.
A recent feature on a community platform I help run shipped in five slices. Each was its own pull request: new database tables, new API routes, a public dashboard, a redesigned member view and some UI polish. It rolled out over about ten days. I wrote zero of the code. Claude planned every slice, Codex wrote every slice, and I reviewed and merged. It was the smoothest stretch of shipping I’ve ever had.
The split
I still use Claude Code for most things. It’s the model I trust most for judgment. But for one job, take this plan and write the code that matches it, exactly, I reach for OpenAI’s Codex CLI.
- Claude plans. I describe what I want, Claude asks questions, and we land on a plan document: the files to touch, the new schema, the test cases, the rollout order.
- Codex executes. It reads the plan and writes the code, with no new design decisions. If the plan says four tables, it makes four tables. If the plan says don’t touch the auth module, it doesn’t.
- Claude reviews. Back in Claude Code, I ask the model that wrote the plan to check the diff against it. Sometimes it spots things I’d miss.
Why split them
For a long time I let Claude do everything. But when the same AI plans and builds, the code drifts from the plan, usually while trying to help. “While I was in there I noticed X” turns into a refactor. A small gap in the plan gets filled with a guess. By the end, the code is sometimes better than the plan, sometimes worse, and almost always different.
Codex doesn’t do that. If the plan is wrong, the code is wrong, so the gap in the plan shows up instead of getting papered over.
That’s also where the pattern strains. If Claude misreads the existing code and plans against a wrong assumption, Codex faithfully builds against it. The code is correct against the spec and incorrect against reality. The fix is better plans, with a more careful read of the existing code first.
So the plans now get a second reviewer before Codex sees them. After Claude writes the spec, I run it through GPT-5.5-pro for an editorial review. I get back two confidence numbers, Claude’s and GPT-5.5-pro’s. If they’re far apart, I look at why. If both are high, it goes to Codex. It’s the same idea as the two-AI code review system I wrote about a few weeks ago, just earlier in the process.
When I don’t bother
For one-off scripts, debugging, or anything where I’m still figuring out what I want, I just use Claude. Writing a plan document for a five-line shell script is silly. The split is worth it when a change touches multiple files, or when other people (or future me) will review it.
The part I didn’t expect
The split changed the shape of the work, not just the speed. The plan stage ends. There’s a document I can put down and pick up the next day, and someone else can read it. Giving two AIs different jobs forced me to write down what each job actually was.
I started doing this because Codex was good at one specific thing. I kept doing it because it made the rest of my work clearer too.
