Moonlight & Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra)
GPT-5.6 Sol Ultra via Codex produced a richer game than Claude Fable 5 but missed a visible rendering bug
“Despite reviewing screenshots during development Codex failed to spot and correct this bug.”
Simon Willison ran the same game-generation prompt through GPT-5.6 Sol Ultra (Codex) and Claude Fable 5, finding GPT-5.6 produced a more faithful, complex game but failed to self-correct a glaring visual bug despite reviewing screenshots during development. The bug was trivially fixed with two plain-English follow-up prompts, highlighting a gap in autonomous visual QA for agentic coding systems. The session cost an estimated $23.28 at full API prices, underscoring the real compute cost of sub-agent-heavy workflows.