Overthinking can even exhaust the smartest AI.
43 minutes, 11,138 inference tokens, 6 batches of instructions—0 combat units.
This is the desperate scorecard Grok 4.6 delivered in StarCraft: Brood War.
In the game, Grok, with its reasoning set to the highest level, xHigh, spent the entire match from start to finish pondering deeply—yet failed to produce a single unit capable of entering battle.
At the other end of the same leaderboard, GPT-6 Astra also delivered a flawless performance, winning all 18 matches with a 100% win rate.
Over the past couple of days, a ranking called the Brood War Bench for StarCraft large models has gone viral on Hacker News, a hub for Silicon Valley programmers.

Top 10 on the Brood War Bench leaderboard. Astra runs the highest inference tier, xhigh, with 18 wins out of 18 matches; Grok's three tiers are all at the bottom of the leaderboard.
Peter Steinberger, the creator of OpenClaw and now an employee at OpenAI, still took the time to share and promote his own model: New rankings are out.

AI is breaking century-old mathematical problems one after another, prompting some to cry out that the singularity has arrived—so rapidly that Terence Tao has publicly called for slowing down.
But the author of this list, Ben Swerdlow, begins with a completely different assessment:
Even the undefeated champion GPT-6 Astra can't beat a human beginner.

Someone on Hacker News cut to the chase with one sentence: “They’re not incapable—they just need too much time. If you pause the game while the model is thinking, it can play.”

Some people even mock: "As if these models aren't general intelligence."
The smartest AI on the entire network—why did it die from its own thoughts in outer space?
Is it still far from AGI?
The game never stops waiting for you.
The creation of this ranking was somewhat dramatic.
Initially, developer Ben created a space-themed version operated entirely by agents and invited a few friends to play.
These guys have barely touched StarCraft, yet they fight just as well as veterans who've played for years.
Their secret is ridiculously simple: just type “Attack,” and let AI handle recruiting, assembling, and fighting.
This made Ben wonder: if humans completely stayed silent and let these models play on their own, how long would they last?
Thus, all 19 AI competitors from the GPT, Claude, and Grok families took the stage, launching a 171-match AI showdown.
The competition ran on Freestyle's cloud virtual machines, with one machine dedicated per match, all starting simultaneously; all game data and every step of the AI's decision-making were fully recorded.
For math problems or coding, a large model can zone out for five minutes and then submit everything at once.
But you can't play StarCraft.
There are no turn-based routines here—only a nonstop real-time battlefield.
Every second you think, the farmer (called a worker by English players, the unit responsible for mining) is frantically mining, the barracks are rapidly producing units, and a large army is already charging toward your high ground.
The old-generation models still try to play this real-time strategy game with turn-based tactics, so while their minds are still thinking, their base has already been destroyed.
Under Steinberger's post, a netizen cut to the chase:
Interesting, no wonder AI always tends to make turn-based games.

Grok pondered for 43 minutes
Not a single soldier was deployed.
Grok is precisely the one who takes this turn-based thinking to its extreme.
In game G043, Grok was a giant in thought but a dwarf in action.
It outputs lengthy strategic reasoning, but in a 43-minute game, it issued only six batches of commands, averaging more than seven minutes between each action.
This is not a coincidence.
In G003, Grok built three Marines but never reached the enemy base; in G002, it built two Zealots but similarly failed to cross the map.
The troops have been deployed, but they can't reach the opponent.
At the highest reasoning tier, it has a record of 2 wins and 15 losses—disastrous.
The data curve is clear: by the 11-minute mark, Astra had already built a strong army, while Grok had an average of fewer than seven workers and a few units on the field.
Ironically, Grok’s treasury actually holds a massive deposit far exceeding Astra’s.
The money is in your pocket, your mind is racing, but you just can’t bring yourself to act.

The farmers and enthusiasts of Astra dismantled Grok's main base, leaving Grok with no combat units and 928 crystals still sitting in his account.
In the brutal battlefield of Starfield, thinking too long isn't careful planning—it's AFK waiting to die.
Astra's ultimate skill
Make your opponent think of death
If Grok was killed by overthinking itself, then Astra is designed to force others to overthink, exhausting its opponents.
The most effective tactic of the Codex family, to which Astra belongs, is harassment.
It often sends just one basic mining peasant across the entire map to wander around the enemy’s base. To human players, this peasant—with zero combat power—can be easily driven off by simply recruiting two workers.
But in the AI arena, this move is a game-changing advantage.
The AI on the other side instantly falls into deep thought upon seeing this farmer, spending dozens of seconds frantically calculating, "How should I handle this farmer?"
During these precious seconds, it does nothing, and the base comes to a complete standstill.

Astra's harassment force destroyed Grok's main base, leaving Grok with no combat units and 928 crystal ores still sitting in his account.
In real-time strategy (RTS) games, harassment is used to disrupt the economy; when AI faces AI, the greatest impact of harassment is to directly crash the opponent’s AI.
Of course, Astra also has weaknesses.
Its internal structure resembles a disorganized makeshift team, with those in charge of the economy, military production, and warfare each operating independently.
New recruits are sent straight to the front lines by the AI in charge of warfare, completely unaware of what it means to muster troops.
However, the author also found that as soon as he took personal charge and linked the several sub-agents together, Astra immediately learned to gather forces and choose the right moment.
The Codex family still has a strong, resilient spirit.
In one scenario, Astra's brother, the GPT-5.6 Terra force, was completely wiped out, and the main base was destroyed.
It lifted the last command center (a Terran building that can fly) into the air, drifted it across the map to a corner, and held out for another six minutes before being destroyed.
Fable is the most like playing StarCraft.
Several rounds throughout the event seemed to want the author to push hard for it—actually, it’s the top-performing game, Fable.
Its approach is the most straightforward: it diligently focuses on the economy and steadily advances its technological tree, rarely resorting to shortcuts.
In one game, it produced a Dragoon; in another, it built an entire row of Protoss structures, even constructing the Templar Archives needed to produce High Templars.
At the 11:30 mark of the match, Fable leads Astra in farmers, army, buildings, and technology.
A good student doesn't necessarily come in first.
Although Fable has a win rate of 83.3%, ranking third, it still cannot catch up to Astra.
A table full of Yiju IT was struck down by Claude Opus 5 before it could even be converted into combat power.

Fable’s barracks, heavy factories, and academies were all overrun, and over 2,800 units of crystal ore remained unclaimed on his ledger—yet he was defeated by a line of Opus 5 riflemen.
The pacing is also vastly different: the median match duration for Astra is only 8 minutes and 10 seconds, while for Fable, which favors a slower, more methodical playstyle, this number stretches to 10 minutes and 37 seconds.
Three large models, three souls.
Codex is like a streetwise troublemaker full of tricks, Grok is like a stubborn top student obsessively tackling the first question on an exam, and Fable is the diligent type who follows guides word for word.
The longer you think about it
Not necessarily better if you play better.
Since overthinking can be deadly, is it better to have an empty mind?
Not entirely.
GPT-5.6 Sol, when set to maximum reasoning, has the lowest win rate—even lower than the medium setting.
But with Astra, the deeper you think, the greater your chances of winning—top tier guarantees a 100% win rate.
More interesting is the bill.
When Astra is set to the highest inference tier, it performs only 12.6 operations per minute, costing an average of $10.54 per trade; switching to the lowest tier doubles the operation frequency to 25.7 operations per minute, but each trade then costs $21.07.
Top tier, half the bets, half the spending, maximum winnings. This clearly shows that the new generation of models has evolved.
They learned that thinking comes at a cost and know when to stop overthinking and take action.
Seven years ago, AI defeated professional players.
Why can't you beat beginners today?
Someone will surely ask: Didn't DeepMind's AlphaStar already defeat professional players seven years ago?
That's right. In early 2019, DeepMind announced AlphaStar's record: two 5-0 victories against professional players TLO and MaNa.

In 2019, AlphaStar played the second game against professional player MaNa. Below is the AI’s decision-making process: reading the game state, determining where to attack, deciding which units to produce, and simultaneously estimating the likelihood of victory.
In 10 matches, the professional player didn't win a single game.
Later, DeepMind added human-like limitations to it: it could only view partial maps through a screen camera, and the number of actions it could perform per second was also reduced.
And so, it anonymously climbed the European ladder, reaching the Master rank and surpassing 99.8% of human players.
On one side, the top 0.2% of the ladder; on the other, you can’t even beat a beginner who just learned how to use a light cannon rush.
Has AI technology regressed?
Actually, comparing the two is fundamentally comparing apples to oranges.
AlphaStar plays StarCraft II like a pure exam machine: it developed its StarCraft expertise solely by analyzing massive amounts of replays and engaging in relentless self-play, dedicating its entire existence to mastering just this one thing.
Today, these large models that battled in the original "Brood War" are generalists with no foundational knowledge of StarCraft.
They rely entirely on reading the problem on the spot, invoking tools, and learning as they go—and each step must be fully thought through before acting.
It's not uncommon for specialists to defeat professionals; generalists hoping to cross over will need to wait a bit longer.
At the end of the article, the author presents their conclusion:
Even the AI champion, unbeatable as it may be, can't defend against the laser rush favored by human beginners. They can't organize complex formations or execute long-term strategies.
Over the past few years, the questions we've given to AI have all been turn-based: repeatedly asking, "Are you calculating correctly?"
Interstellar is different. Like autonomous driving and embodied robotics, it exists in a world that never stops, where competitors are always moving.
This test also revealed the most critical weakness of AI as it moves toward autonomous action: expending more thinking tokens does not necessarily result in a more intelligent agent.
The next phase of the Agent competition may no longer be about who thinks deeper, but who can connect observation, decision-making, and execution into a continuously spinning loop within limited time.
The crash in space might be a rehearsal for AI entering the real world.
Reference materials:
https://bw.swerdlow.dev/report?utm_source=chatgpt.com
https://x.com/steipete/status/2101748820237500557
This article is from the WeChat public account "New Intelligence Yuan" (ID: AI_era), author: ASI Revelation, editor: Yuan Yu.
