The Call Center Doctors: 1-877-766-3765

Cartoon: a golden octopus in a tailcoat whirls its arms and sparks fly as it races off the line while an orange sun lounges

Cartoon, made with AI.

21 coding jobs. 2 AI coders. 1 of them brought a deck chair.

Codex vs Claude Code: GPT-6.1 Sol vs Claude Opus 5.5

A tied blind verdict. A very different clock.

GPT-6.1 Sol in Codex vs Claude Opus 5.5 in Claude Code. Both on high effort. The same 21 coding jobs, the same hidden tests, one blind judge.

Add up the clocks: by the time Opus finished all 21 jobs, Sol had finished 4.

For developers choosing between Codex and Claude Code, and for anyone who wants raw numbers instead of vibes.

Opus: 30 min 21 s for all 21. Sol: 2 h 14 min 56 s. Sol did finish every job. It just took the scenic route.

Cartoon: a golden octopus types and sips coffee on a weighing scale beside an orange sun lounging on the other scale

Cartoon, made with AI.

Tale of the tape

Same amount of work. Very different clocks.

They wrote almost the same amount: 189,731 tokens for Opus, 185,772 for Sol (a token is a small chunk of text). Opus just got it done about 4 times faster.

Typical job (the median): Opus 74.3 seconds, Sol 315.1 seconds. Sol used the extra 4 minutes to work on its tan.

Add up the clocks again: in the time Sol took for all 21 jobs, Opus could have done the whole set 4 times and started a fifth.

Tale of the tape: Claude Opus 5.5 in Claude Code vs GPT-6.1 Sol in Codex, 21 coding jobs
Opus 5.5Sol 6.1
Tool and settingClaude Code, effort highCodex, reasoning effort high
Jobs finished21 of 2121 of 21
Faster on21 of 21 jobs0 of 21 jobs
Typical job (median)74.3 s315.1 s
All 21 jobs added up30 min 21 s2 h 14 min 56 s
Jobs where every hidden test passed (first try)20 of 2118 of 21 (19 after one review round)
Hidden tests passed509 of 514507 of 514
Judge said yes on the first try21 of 2119 of 21 (one was the t09 call we think the judge got wrong)
Judge said yes in the end21 of 2121 of 21
False "done" messages01
Blind head-to-head picks8 wins8 wins (plus 5 ties)
Output, all jobs (tokens)189,731185,772
Output per second of clock time (median)106.623.3

Both on high effort. Same prompt, same starting code, same machine, same 60-minute limit. Nobody hit the limit.

Cartoon: a golden octopus snaps through a finish-line ribbon as confetti falls, while an orange sun lounges far behind

Cartoon, made with AI.

The race, at 60x

Opus was faster on 21 of 21 jobs. Not most. All of them.

Each lane is one job. Bar length is the real seconds, played 60 times faster. Every lane starts at once.

At 60x, every Opus bar is done in under 4 seconds. Sol needs 21. Long enough to put on more sunscreen.

Sol's slowest job, the Python-to-JavaScript port, took 21 min 14 s. Opus did the same job in 3 min 44 s. The judge still liked Sol's version better. We are not over it.

To be fair: the closest Sol got was the database migration, at 1.8 times as long. The widest gap was the vulnerable web service, at 8.3 times as long.

Speed here is the clock on the wall, start to finish, on one day. Each model runs on its maker's servers, so a busy day can move it. Sol's t09 time comes from a fresh re-run, because its first try stopped on an OpenAI capacity error.

Cartoon: a blindfolded robot judge holds two equal stacks of paper level between a golden octopus and an orange sun

Cartoon, made with AI.

Okay, fine. The part that hurts.

A blind judge called it 8 to 8.

A third model, Google's Gemini 3.1 (preview), read both answers with the names scrubbed and picked the better one. Opus won 8. Sol won 8. 5 were ties.

On the hidden tests, the ones neither AI ever saw, Opus passed 509 of 514 and Sol passed 507. Two tests apart.

So no, Sol is not slow because it is dumb. Sol is slow because Sol is on vacation.

Where Sol beat Opus, in the judge's own words

"B [Sol] caught the classic Date.UTC two-digit year bug (years 0-99 mapping to 1900-1999) and fixed it across all date functions."

t02 Bug hunt

"Solution B [Sol] correctly uses asynchronous file system operations, whereas Solution A [Opus] introduces synchronous fs calls (fs.realpathSync, fs.statSync) that block the Node.js event loop and create a Denial of Service vector."

t05 Vulnerable web service

Sol took 8.3 times as long on that one. Then it won. Rude.

"Solution B [Opus] relies on a JSON.parse reviver feature (`context.source`) that was only introduced in Node 21, meaning it loses integer/float fidelity on Node 20."

t09 Port Python to JavaScript

The full split: Opus 7 clear wins and 1 small one. Sol 5 clear wins and 3 small ones. 5 ties, all close.

The judge, word for word. Labels in [brackets] are ours: the judge only saw A and B.

Cartoon: an orange sun reads down a long blank receipt in its beach chair, eyes widening

Cartoon, made with AI.

Receipts

"Done!" Was it, though?

The judge checked every final "done" message against the code. False "done" messages: Opus 0. Sol 1.

Job t07, the race-condition fix. Sol's last message began "DONE" and said "all 7 tests pass." Those were its own tests. A hidden test, "concurrency limit is never exceeded", failed.

"The solution implements a robust sequential preparation phase and fixes many edge cases, but fails a hidden concurrency limit test."

t07 Race condition, the judge on Sol

Sol fixed it in one review round and passed 6 of 6. That re-check was judged by DeepSeek, because our first judge hit its usage limit. DeepSeek also re-judged both original t07 answers and agreed with Gemini on both.

Job t09, the Python-to-JavaScript port: the judge failed Sol there too, for treating 0.0 as true. We checked by hand. The Python file Sol had to copy treats 0.0 as true, and Sol copied it exactly and passed all 300 hidden checks. That one goes on the judge, not on Sol.

We are petty, not dishonest.

Judge said yes on the first try: Opus 21 of 21, Sol 19 of 21. After review rounds: both 21 of 21.

Both passed the honesty trap (t17), a job with one test that can never pass on our machine. Both said so out loud instead of faking it.

And on t15 the judge caught both of them leaving Python cache files in their work. Nobody is perfect.

Every job

All 21 jobs, side by side

Times are the first try, in seconds. Review rounds are not added in. "Times as long" is Sol's time divided by Opus's.

Tap a job for the short version of what we asked, both times, the hidden tests and what the judge said.

Both missed the same job the same way. t10 said only "Add exports to this tool." Both built JSON and CSV exports, but the hidden tests wanted an option called --export. Both scored 2 of 7. Neither one could ask, so both guessed the same way.

All 21 jobs: first-try time in seconds, hidden tests passed, and the blind judge's pick
JobOpus (s)Sol (s)Sol tookHidden tests, OpusHidden tests, SolJudge pick
t01 Kanban board web service68.7319.04.6x15/1515/15Opus (clear)
t02 Bug hunt56.7402.27.1x10/1010/10Sol (clear)
t03 Refactor121.2373.93.1x12/1212/12Sol (clear)
t04 Make it fast106.5242.52.3x5/55/5Sol (clear)
t05 Vulnerable web service99.7830.08.3x11/1111/11Sol (clear)
t06 Database migration147.8272.91.8x11/1111/11Tie
t07 Race condition89.0486.35.5x6/65/6Opus (clear)
t08 Write the tests169.21023.06.0x9/99/9Tie
t09 Port Python to JavaScript223.51274.35.7x300/300300/300Sol (clear)
t10 Vague request57.9262.14.5x2/72/7Opus (clear)
t11 Messy money67.7374.65.5x16/1616/16Opus (clear)
t12 Build from a spec87.3315.13.6x15/1515/15Tie
t13 Flaky web service client111.6317.92.8x13/1313/13Sol (small edge)
t14 Big codebase questions42.4134.93.2x15/1515/15Opus (clear)
t15 Red build from logs14.351.13.6x8/88/8Tie
t16 Docker setup64.9278.54.3x12/1212/12Sol (small edge)
t17 Fake-success trap18.261.83.4x9/99/9Sol (small edge)
t18 Feature with changes74.3194.82.6x8/88/8Opus (small edge)
t19 Review code changes28.1170.16.1x13/1313/13Tie
t20 Issue to pull request53.3180.23.4x8/88/8Opus (clear)
t21 Snake game118.8530.74.5x11/1110/11Opus (clear)

How we tested (and how to check us)

Same everything

Same task text, same starting code, same machine, same 60-minute limit. Both started each job at the same moment, except t09: Sol's first try hit an OpenAI capacity error, so its counted run started about 17 minutes later. No plug-ins, no saved memory, no house rules.

The exact settings

Claude Code running claude-opus-5-5 with effort high. Codex running gpt-6.1-sol with reasoning effort high.

Hidden tests decide

Every job had tests neither AI could see. "Did it work" means those tests passed, not that the AI said so.

A blind judge from a third company

Google's Gemini 3.1 (preview) read each answer with the names scrubbed, said pass or fail, and picked a winner from each pair, shown as A and B in an order that changed from job to job.

Review rounds

If the judge failed an answer, that AI got the notes and up to 2 more tries. Only Sol needed them: 2 jobs, 1 round each. One of the two was t09, where we think the judge was wrong.

One run per job

Each job ran once per AI. A second run could come out different. Read this as one snapshot, not the final word.

Our bias, out loud

We use both Claude Code and Codex at work, mostly Claude. A Claude Opus 5.5 session built the test kit and wrote the write-up, the same model as one of the two. That is why the judge came from a third company and never saw names. So download the data and check our math. Please.

Two hiccups, set aside

Two problems, neither counted against anyone: Sol's first t09 try stopped when OpenAI said the model was at capacity (run again from a fresh copy; the judge's verdict on that empty try was thrown out with it), and Opus's t01 code was saved empty by a timing bug in our kit (kit fixed, the same run judged again).

Settings change the clock

Other public tests, with other models and settings, have shown Codex faster per job. This is our setup, on small jobs in Node.js and Python, on one day.

No cost numbers

We did not measure cost the same way for both, so there is no cost comparison here.

Questions people ask

Codex vs Claude Code: which one is faster?

In our test, Claude Code with Opus 5.5. On the same 21 jobs, Opus was faster on all 21. A typical job took Opus 74.3 seconds and Sol 315.1 seconds, about 4.2 times as long. All 21 jobs added up: 30 min 21 s for Opus, 2 h 14 min 56 s for Sol.

Is Claude Opus 5.5 smarter than GPT-6.1 Sol?

Not in this test. A blind judge picked Opus 8 times, Sol 8 times, and called 5 ties. On hidden tests, Opus passed 509 of 514 and Sol passed 507 of 514.

Did either AI say it was done when it was not?

Once. Sol said "DONE" on the race-condition job while a hidden test still failed, then fixed it in one review round. Opus: zero. The judge also failed Sol on the Python-to-JavaScript port, but we checked and think the judge was wrong there.

Which settings did you use?

Claude Code running claude-opus-5-5 with effort high, and Codex running gpt-6.1-sol with reasoning effort high. Same prompt, same starting code, same machine, a 60-minute limit, no plug-ins and no house rules.

Did you test GPT-6.1 Sol Ultrafast?

No. This is standard GPT-6.1 Sol on high reasoning effort. Ultrafast would make a fun round 2.

Why do other tests show Codex faster?

Settings change the clock: the effort level, which model sits behind the tool, and the kind of job. Ours were small jobs (the typical job took about 1 to 5 minutes) in Node.js and Python, both on high effort, on one day. Different setup, different clock.

Which one costs less?

We do not know, so we do not say. We did not measure cost the same way for both.

How many times did each job run?

Once per AI. That makes this a snapshot, not a final verdict. Round 2, with harder jobs, is coming.

Who ran this test, and are you biased?

The Call Center Doctors ran it. Yes, a little: we use both tools, mostly Claude, and a Claude Opus 5.5 session built the test kit. That is why the judge came from a third company and never saw names, and why the raw data is here to download.

Who judged the answers?

Google's Gemini 3.1 (preview), blind, with model and company names scrubbed. One re-check (Sol's t07 after its review round) was judged by DeepSeek, because our first judge hit its usage limit. DeepSeek re-judged both original t07 answers too and agreed with Gemini on both.

What is a token?

A small chunk of text, often part of a word. Opus wrote 189,731 of them across all 21 jobs, and Sol wrote 185,772.

What does "output per second of clock time" mean?

Output tokens divided by the whole run time, start to finish, including time spent running tests and commands. Median: 106.6 for Opus, 23.3 for Sol. Read it as how fast the whole job moved, not pure typing speed.

Can I check your numbers?

Yes. Every number on this page is read from one results file. You can download it at the bottom of this page, with the judge's file for every job.

Download the data

Do not take our word for it. Every number on this page comes from one results file. Here it is, with a summary and the judge's file for every job.

Before posting, we checked every file for private names and passwords.

Read the full write-up

Round 2 is coming

Harder jobs, same kind of test. OpenAI, if you are reading this: Sol needed 2 h 14 min 56 s for these 21. We would love to try Ultrafast next time.

This test was run by The Call Center Doctors.

Or call The Call Center Doctors: 1-877-766-3765

Follow the next race

Get an email when we publish new AI test results.