RC RANDOM CHAOS

Can LLM Agents Play StarCraft? A New Benchmark Says: Barely

· via Hacker News

Original source

Brood War Bench

Hacker News →

Brood War Bench pits AI models against each other in StarCraft: Brood War, played entirely through agents that issue commands rather than direct control. The project grew out of an experiment in which casual human players won simply by telling their agents to “attack.” Across a round-robin of 19 model-and-effort configurations run in parallel on cloud VMs, the verdict is blunt: no model rose above beginner level. Codex-based configs led consistently, but even the strongest agents couldn’t build complex armies, hold off basic pushes, or execute coherent strategies. The author notes a human doing a simple photon rush would beat all of them.

The failure modes are revealing. Older and slower models treated the real-time game like a turn-based one, getting overrun while they deliberated — which is why some lower-effort settings actually did better. Codex found early harassment (sending a lone worker to disrupt) far easier than sustained production, and its habit of spinning up uncommunicative subagents for economy, army, and control led to the classic novice error of feeding units into attacks one at a time. Grok models mostly stalled: one run burned over 11,000 reasoning tokens across 43 minutes while issuing six command batches and never fielding a combat unit. Claude Fable stood out for actually trying to play — building economies and climbing tech trees toward Mutalisks or Templar tech — though ambition often outran execution.

The takeaway is less about who won than about how far agents remain from competent real-time play. The core weakness isn’t strategy but the observe-act loop: keeping up with a game that doesn’t pause for thinking. The author frames the benchmark as nowhere near saturated and expects models to improve, positioning it as an open, playable testbed for agentic reasoning under real-time pressure.

Read the full article

Continue reading at Hacker News →

This is an AI-generated summary. Read the original for the full story.