Confirm action

Are you sure you want to delete?

Link copied!
AI
Jul 28, 2026 · 2 min read

Claude Opus 5 built its own Starfield in 24 hours — and criticized itself along the way

AffMarketing World
Patric Mirgeschiss
Editor, AffMarketing World
Claude Opus 5 builds a Starfield-style space exploration game

A developer ran Claude Opus 5 through a 24-hour marathon session with one goal: build a playable demo of a space game in the spirit of Starfield. The result speaks for itself — the game is playable directly in the browser.

How it was built

The core condition of the experiment was no pre-made assets and no outside code. Every pixel and every sound in the game came from the model itself. For 3D objects, Claude used Blender through the Model Context Protocol, modeling scene objects on its own rather than just generating textures over ready-made shapes.

Worth calling out separately is the quality-control setup: the developer configured an adversarial subagent with the tongue-in-cheek name “Claude of Duty,” which compared game screenshots against reference shots from AAA space titles like Starfield in real time and flagged the model on mismatches. Effectively, one copy of Claude was building the game while another was picking it apart.

The game world has no hard boundaries either — reach the edge of the visible universe and procedural generation builds new planets on the fly with randomized characteristics, and those planets can actually be landed on and explored rather than just viewed from a distance.

What this signals beyond the demo itself

Pairing two copies of a model, one building and one reviewing, echoes a pattern already common in agentic coding pipelines, where a draft gets routed through a separate review loop that looks at the output with fresh eyes, without the first copy’s own reasoning baked in. What stands out here is the grading criteria: the model was judged on a visual comparison against screenshots from real AAA games, essentially a subjective read on whether something “looks like a real game.”

For developers, that setup is more interesting than the game itself. Twenty-four straight hours of agentic work with real-time self-criticism mainly tests context retention and decision consistency over a long stretch, the exact place where extended agent sessions usually fall apart. A loop like this assembled by a single developer, rather than a team running lab-grade infrastructure, lowers the barrier to running experiments at this scale considerably.

What impresses me most isn’t the game itself — it’s clearly amateur by industry standards — it’s that the model held together through 24 straight hours of development with real-time self-criticism baked in and never lost the thread of the project. That’s a far more telling test of agentic endurance than any synthetic benchmark.

Patric Mirgeschiss
Reviewed by
Patric Mirgeschiss
Editor · AffMarketing World
Published Jul 28, 2026
X Profile →