Can Claude beat Codex with only 5 prompts?
I gave myself five prompts each to build the same app in Claude using Opus 4.8, and Codex using GPT 5.5.
The results were NOT what I was expecting!
See which agent built the better app, which one stumbled massively, and look at the actual numbers behind these coding agents.
00:00 - The experiment
00:22 - What I'm building
00:44 - Prompt 1
01:40 - Prompt 2
02:27 - Prompt 3
02:50 - Prompt 4
03:23 - It's not working?
03:45 - Prompt 5
04:10 - Final state
04:48 - My scores
04:54 - Claude telemetry analysis
05:30 - Using the cache effectively
06:35 - Codex telemetry analysis
07:04 - Which is the leaner agent?
#claudecode #codex #codingagents #opentelemetry