Can Claude beat Codex with only 5 prompts?

Jul 16, 2026

I gave myself five prompts each to build the same app in Claude using Opus 4.8, and Codex using GPT 5.5.

The results were NOT what I was expecting!

See which agent built the better app, which one stumbled massively, and look at the actual numbers behind these coding agents.

00:00 - The experiment

00:22 - What I'm building

00:44 - Prompt 1

01:40 - Prompt 2

02:27 - Prompt 3

02:50 - Prompt 4

03:23 - It's not working?

03:45 - Prompt 5

04:10 - Final state

04:48 - My scores

04:54 - Claude telemetry analysis

05:30 - Using the cache effectively

06:35 - Codex telemetry analysis

07:04 - Which is the leaner agent?

#claudecode #codex #codingagents #opentelemetry