
Cartoon, made with AI.

Cartoon, made with AI.

Cartoon, made with AI.

Cartoon, made with AI.

Cartoon, made with AI.
Every job
All 21 jobs, side by side
Times are the first try, in seconds. Review rounds are not added in. "Times as long" is Sol's time divided by Opus's.
Tap a job for the short version of what we asked, both times, the hidden tests and what the judge said.
Both missed the same job the same way. t10 said only "Add exports to this tool." Both built JSON and CSV exports, but the hidden tests wanted an option called --export. Both scored 2 of 7. Neither one could ask, so both guessed the same way.
| Job | Opus (s) | Sol (s) | Sol took | Hidden tests, Opus | Hidden tests, Sol | Judge pick |
|---|---|---|---|---|---|---|
| t01 Kanban board web service | 68.7 | 319.0 | 4.6x | 15/15 | 15/15 | Opus (clear) |
| t02 Bug hunt | 56.7 | 402.2 | 7.1x | 10/10 | 10/10 | Sol (clear) |
| t03 Refactor | 121.2 | 373.9 | 3.1x | 12/12 | 12/12 | Sol (clear) |
| t04 Make it fast | 106.5 | 242.5 | 2.3x | 5/5 | 5/5 | Sol (clear) |
| t05 Vulnerable web service | 99.7 | 830.0 | 8.3x | 11/11 | 11/11 | Sol (clear) |
| t06 Database migration | 147.8 | 272.9 | 1.8x | 11/11 | 11/11 | Tie |
| t07 Race condition | 89.0 | 486.3 | 5.5x | 6/6 | 5/6 | Opus (clear) |
| t08 Write the tests | 169.2 | 1023.0 | 6.0x | 9/9 | 9/9 | Tie |
| t09 Port Python to JavaScript | 223.5 | 1274.3 | 5.7x | 300/300 | 300/300 | Sol (clear) |
| t10 Vague request | 57.9 | 262.1 | 4.5x | 2/7 | 2/7 | Opus (clear) |
| t11 Messy money | 67.7 | 374.6 | 5.5x | 16/16 | 16/16 | Opus (clear) |
| t12 Build from a spec | 87.3 | 315.1 | 3.6x | 15/15 | 15/15 | Tie |
| t13 Flaky web service client | 111.6 | 317.9 | 2.8x | 13/13 | 13/13 | Sol (small edge) |
| t14 Big codebase questions | 42.4 | 134.9 | 3.2x | 15/15 | 15/15 | Opus (clear) |
| t15 Red build from logs | 14.3 | 51.1 | 3.6x | 8/8 | 8/8 | Tie |
| t16 Docker setup | 64.9 | 278.5 | 4.3x | 12/12 | 12/12 | Sol (small edge) |
| t17 Fake-success trap | 18.2 | 61.8 | 3.4x | 9/9 | 9/9 | Sol (small edge) |
| t18 Feature with changes | 74.3 | 194.8 | 2.6x | 8/8 | 8/8 | Opus (small edge) |
| t19 Review code changes | 28.1 | 170.1 | 6.1x | 13/13 | 13/13 | Tie |
| t20 Issue to pull request | 53.3 | 180.2 | 3.4x | 8/8 | 8/8 | Opus (clear) |
| t21 Snake game | 118.8 | 530.7 | 4.5x | 11/11 | 10/11 | Opus (clear) |
How we tested (and how to check us)
Same everything
Same task text, same starting code, same machine, same 60-minute limit. Both started each job at the same moment, except t09: Sol's first try hit an OpenAI capacity error, so its counted run started about 17 minutes later. No plug-ins, no saved memory, no house rules.
The exact settings
Claude Code running claude-opus-5-5 with effort high. Codex running gpt-6.1-sol with reasoning effort high.
Hidden tests decide
Every job had tests neither AI could see. "Did it work" means those tests passed, not that the AI said so.
A blind judge from a third company
Google's Gemini 3.1 (preview) read each answer with the names scrubbed, said pass or fail, and picked a winner from each pair, shown as A and B in an order that changed from job to job.
Review rounds
If the judge failed an answer, that AI got the notes and up to 2 more tries. Only Sol needed them: 2 jobs, 1 round each. One of the two was t09, where we think the judge was wrong.
One run per job
Each job ran once per AI. A second run could come out different. Read this as one snapshot, not the final word.
Our bias, out loud
We use both Claude Code and Codex at work, mostly Claude. A Claude Opus 5.5 session built the test kit and wrote the write-up, the same model as one of the two. That is why the judge came from a third company and never saw names. So download the data and check our math. Please.
Two hiccups, set aside
Two problems, neither counted against anyone: Sol's first t09 try stopped when OpenAI said the model was at capacity (run again from a fresh copy; the judge's verdict on that empty try was thrown out with it), and Opus's t01 code was saved empty by a timing bug in our kit (kit fixed, the same run judged again).
Settings change the clock
Other public tests, with other models and settings, have shown Codex faster per job. This is our setup, on small jobs in Node.js and Python, on one day.
No cost numbers
We did not measure cost the same way for both, so there is no cost comparison here.
Questions people ask
Codex vs Claude Code: which one is faster?
In our test, Claude Code with Opus 5.5. On the same 21 jobs, Opus was faster on all 21. A typical job took Opus 74.3 seconds and Sol 315.1 seconds, about 4.2 times as long. All 21 jobs added up: 30 min 21 s for Opus, 2 h 14 min 56 s for Sol.
Is Claude Opus 5.5 smarter than GPT-6.1 Sol?
Not in this test. A blind judge picked Opus 8 times, Sol 8 times, and called 5 ties. On hidden tests, Opus passed 509 of 514 and Sol passed 507 of 514.
Did either AI say it was done when it was not?
Once. Sol said "DONE" on the race-condition job while a hidden test still failed, then fixed it in one review round. Opus: zero. The judge also failed Sol on the Python-to-JavaScript port, but we checked and think the judge was wrong there.
Which settings did you use?
Claude Code running claude-opus-5-5 with effort high, and Codex running gpt-6.1-sol with reasoning effort high. Same prompt, same starting code, same machine, a 60-minute limit, no plug-ins and no house rules.
Did you test GPT-6.1 Sol Ultrafast?
No. This is standard GPT-6.1 Sol on high reasoning effort. Ultrafast would make a fun round 2.
Why do other tests show Codex faster?
Settings change the clock: the effort level, which model sits behind the tool, and the kind of job. Ours were small jobs (the typical job took about 1 to 5 minutes) in Node.js and Python, both on high effort, on one day. Different setup, different clock.
Which one costs less?
We do not know, so we do not say. We did not measure cost the same way for both.
How many times did each job run?
Once per AI. That makes this a snapshot, not a final verdict. Round 2, with harder jobs, is coming.
Who ran this test, and are you biased?
The Call Center Doctors ran it. Yes, a little: we use both tools, mostly Claude, and a Claude Opus 5.5 session built the test kit. That is why the judge came from a third company and never saw names, and why the raw data is here to download.
Who judged the answers?
Google's Gemini 3.1 (preview), blind, with model and company names scrubbed. One re-check (Sol's t07 after its review round) was judged by DeepSeek, because our first judge hit its usage limit. DeepSeek re-judged both original t07 answers too and agreed with Gemini on both.
What is a token?
A small chunk of text, often part of a word. Opus wrote 189,731 of them across all 21 jobs, and Sol wrote 185,772.
What does "output per second of clock time" mean?
Output tokens divided by the whole run time, start to finish, including time spent running tests and commands. Median: 106.6 for Opus, 23.3 for Sol. Read it as how fast the whole job moved, not pure typing speed.
Can I check your numbers?
Yes. Every number on this page is read from one results file. You can download it at the bottom of this page, with the judge's file for every job.
Download the data
Do not take our word for it. Every number on this page comes from one results file. Here it is, with a summary and the judge's file for every job.
- The results file (every number on this page)
- The results summary, as a text file
The judge's files, one per job
- t01 Kanban board web service
- t02 Bug hunt
- t03 Refactor
- t04 Make it fast
- t05 Vulnerable web service
- t06 Database migration
- t07 Race condition
- t08 Write the tests
- t09 Port Python to JavaScript
- t10 Vague request
- t11 Messy money
- t12 Build from a spec
- t13 Flaky web service client
- t14 Big codebase questions
- t15 Red build from logs
- t16 Docker setup
- t17 Fake-success trap
- t18 Feature with changes
- t19 Review code changes
- t20 Issue to pull request
- t21 Snake game
Before posting, we checked every file for private names and passwords.
Round 2 is coming
Harder jobs, same kind of test. OpenAI, if you are reading this: Sol needed 2 h 14 min 56 s for these 21. We would love to try Ultrafast next time.
This test was run by The Call Center Doctors.
Or call The Call Center Doctors: 1-877-766-3765
Follow the next race
Get an email when we publish new AI test results.

