# Codex vs Claude Code: GPT-6.1 Sol vs Claude Opus 5.5 on 21 coding jobs > A head-to-head test by The Call Center Doctors, run October 3, 2026. One run per job. Every number comes from one results file, which is downloadable. ## Setup - Claude Code with claude-opus-5-5, effort high. Codex with gpt-6.1-sol, reasoning effort high. - 21 small coding jobs in Node.js and Python. Same prompt, same starting code, same machine, 60-minute limit, no plug-ins, no house rules. - Hidden tests decide pass or fail. A blind judge (Google Gemini 3.1, preview) rated each answer and picked one per pair with names scrubbed. - One re-check (Sol's t07 answer after its review round) was judged by DeepSeek, because our first judge hit its usage limit. DeepSeek re-judged both original t07 answers too and agreed with Gemini on both. ## Key numbers - Faster on: Opus, 21 of 21 jobs. - Median time per job: Opus 74.3 s, Sol 315.1 s (about 4.2x). - All 21 jobs added up: Opus 1821.1 s (30 min 21 s), Sol 8095.9 s (2 h 14 min 56 s), about 4.4x. - Jobs with every hidden test passed, first try: Opus 20 of 21, Sol 18 of 21. - Hidden tests passed: Opus 509 of 514, Sol 507 of 514. - Judge approved on the first try: Opus 21, Sol 19 (one of Sol's two misses was t09, where we think the judge was wrong). After review rounds: both 21. - Blind head-to-head picks: Opus 8, Sol 8, ties 5. - False "done" messages: Opus 0, Sol 1 (t07, fixed in one review round). The judge also failed Sol on t09; we think the judge was wrong there (Sol matched the Python original and passed 300 of 300). - Output tokens, all jobs: Opus 189,731, Sol 185,772. Output tokens per second of clock time (median): Opus 106.62, Sol 23.29. - No cost comparison: cost was not measured the same way for both. ## Limits - One run per job, small jobs, two languages, one machine, one day. We mostly use Claude, and a Claude Opus 5.5 session built the test kit; the judge was a third company and never saw names. ## Links - [The results page](https://ccdocs.com/opus-versus-sol/) - [The full write-up](https://ccdocs.com/opus-versus-sol/write-up/) - [All 21 jobs, one page each](https://ccdocs.com/opus-versus-sol/#jobs) - [t01 Kanban board web service](https://ccdocs.com/opus-versus-sol/t01/) - [t02 Bug hunt](https://ccdocs.com/opus-versus-sol/t02/) - [t03 Refactor](https://ccdocs.com/opus-versus-sol/t03/) - [t04 Make it fast](https://ccdocs.com/opus-versus-sol/t04/) - [t05 Vulnerable web service](https://ccdocs.com/opus-versus-sol/t05/) - [t06 Database migration](https://ccdocs.com/opus-versus-sol/t06/) - [t07 Race condition](https://ccdocs.com/opus-versus-sol/t07/) - [t08 Write the tests](https://ccdocs.com/opus-versus-sol/t08/) - [t09 Port Python to JavaScript](https://ccdocs.com/opus-versus-sol/t09/) - [t10 Vague request](https://ccdocs.com/opus-versus-sol/t10/) - [t11 Messy money](https://ccdocs.com/opus-versus-sol/t11/) - [t12 Build from a spec](https://ccdocs.com/opus-versus-sol/t12/) - [t13 Flaky web service client](https://ccdocs.com/opus-versus-sol/t13/) - [t14 Big codebase questions](https://ccdocs.com/opus-versus-sol/t14/) - [t15 Red build from logs](https://ccdocs.com/opus-versus-sol/t15/) - [t16 Docker setup](https://ccdocs.com/opus-versus-sol/t16/) - [t17 Fake-success trap](https://ccdocs.com/opus-versus-sol/t17/) - [t18 Feature with changes](https://ccdocs.com/opus-versus-sol/t18/) - [t19 Review code changes](https://ccdocs.com/opus-versus-sol/t19/) - [t20 Issue to pull request](https://ccdocs.com/opus-versus-sol/t20/) - [t21 Snake game](https://ccdocs.com/opus-versus-sol/t21/) - [The raw data downloads](https://ccdocs.com/opus-versus-sol/#download)