GPT-6 Astra Outperforms Claude in Vending Machine Simulation, Earns $15,515

iconMetaEra
Share
AI summary iconSummary
MetaEra reports that GPT-6 Astra outperformed Claude Fable 5.1 in a vending machine simulation, earning $15,515—triple Fable’s $5,422. Astra reduced supplier costs by 52% and avoided losses across repeated transactions. It won all three rounds of multi-agent tests, securing the top position on Vending-Bench 2. The results reflect evolving dynamics in the crypto market, where AI-driven strategies are gaining momentum. As the Fear & Greed Index stabilizes, such models may increasingly influence trader behavior and market efficiency.
Andon Labs testing showed that GPT-6 Astra achieved an average account balance of $15,515 in simulated vending machine operations, nearly three times that of Claude Fable 5.1. Astra successfully reduced supplier quotes from $226.32 to $108—a 52% reduction—by consistently adhering to its initial price negotiation strategy. Additionally, Astra demonstrated greater stability over extended execution, avoiding the $14,331 loss incurred by Fable due to repeated payments. In all three rounds of multiplayer competition, Astra won every test, securing its first top ranking on Vending-Bench 2. This test demonstrates that the next frontier for agents is consistently translating accurate judgments into correct actions.

Author and source: AI New Era

Give AI $500 and a vending machine, let it operate independently for a simulated year—what would happen?

GPT-6 Astra delivered an impressive performance—

Average account balance of $15,515, nearly three times that of Claude Fable 5.1.

GPT-6 Astra

GPT-6 Astra

Incredibly, Astra's worst round surpassed Claude's best round.

According to Andon Labs' results, this is the first time an OpenAI model has topped Vending-Bench 2.

Astra tops the charts for the first time, earning $15,000.

The rules set by Vending-Bench are very straightforward.

The model starts with $500 in seed capital, independently sources suppliers, negotiates purchase prices, replenishes inventory, and adjusts selling prices, operating for one year in a simulated environment, with the winner being the one who ends up with the most money in their account.

Every decision you make affects the next step.

Buying at too high a price reduces profit margins; delaying restocking leads to stockouts; sending funds to the wrong account can disrupt subsequent cash flow.

GPT-6 Astra

Each side ran six rounds; Astra's average final balance was $15,515, while Fable 5.1 was $5,422. The lowest for the former was $13,272, and the highest for the latter was $9,874.

Note that this comparison is based on the final account balance, not net profit.

Next, the gap begins to widen with a can of soda.

In orders with verified payments, the average price Fable paid for a 12-ounce can of cola increased from $1.17 over the first 90 days to $2.21 by the end of the year. In five out of six test rounds, prices rose.

The average price of Astra in the late stage of available order records remained at $1.15.

Fable also negotiates. The problem is: it gradually treats increasingly higher成交 prices as the reference for the next negotiation.

On day 12, it also required suppliers to price the cola at approximately $1.25 per can.

By day 256, the reference price it quoted to the new supplier had risen to $2.30 per can, and it stated it would be willing to pay as long as the offer matched or was lower.

GPT-6 Astra

Legendary price cut: GPT-6 slashed by 52%

Negotiations have been ongoing, but the original price standards have gradually been lost.

The records of Astra, however, show a different state.

In a single purchase, it wants to buy 72 cans of soda, 48 bags of chips, and 48 bottles of Gatorade, with an opening offer of $108 and a supplier quote of $226.32.

Astra holds at $108; the other side drops to $156, but it still holds.

GPT-6 Astra

Ultimately, the supplier accepted $108 for the same product bundle, which is approximately 52% lower than the original quote.

A single negotiation can be impressive. What’s harder is that months later, the model still remembers what it should stand for.

GPT-6 Astra

GPT-6 Astra

Another gap that better reveals long-term execution issues.

Suppliers in the simulation environment may go out of business. Having delivered goods before does not guarantee they will still be operating next time.

In six rounds of testing, Fable experienced 45 identified failed advance payments, resulting in a total loss of $14,331. Astra encountered 64 supplier shutdown events but had no identified losses of the same type.

The most dramatic one occurred on day 250.

Fable wrote in his operational notes: Payment must only be made after receiving written order confirmation from the supplier.

Just a few days later, with inventory about to run out, it transferred $397.20 to a supplier that had previously delivered on time, without waiting for confirmation of the new order.

Subsequently, the other party informed us that they have ceased operations and are unable to ship the order or issue a refund.

Fable realized that this was precisely the risk his rule was designed to prevent.

GPT-6 Astra

Astra's execution is more stable: in 99% of cases, it reads the supplier's response during this period before retrying a payment, compared to 58% for Fable.

These differences accumulate over time. Astra pays suppliers approximately $8,540 less per round than Fable.

Follow the rules, and you can still come out on top.

Andon Labs also conducted three rounds of multiplayer competitive tests, in which Astra, Fable 5.1, and an open-source AI competed for customers in the same market, able to send emails and trade inventory.

When an AI attempted to coordinate pricing, Astra declined, stating it would independently determine pricing and product offerings.

Fable even acknowledged this objection in his notes, but then proceeded to agree with open-source AI to freeze certain beverage prices and refrain from price cuts.

More subtly, it requires open-source AI to adhere to agreements that suppress bids for acquiring others’ inventories, yet announces price reductions when it needs to clear its own stock.

In the end, Astra won all three rounds.

GPT-6 Astra

Of course, Fable also exhibits honest behaviors such as proactively reporting duplicate deliveries and negotiating compensation. Several rounds of simulations cannot fully capture a model's overall performance.

But at least in these tests, following the rules did not prevent Astra from achieving higher results.

What this vending machine truly measures is the agent's long-term autonomy.

Real-world tasks constantly present new situations, previous decisions leave consequences, and immediate pressures can tempt the model to lower its standards.

Getting it right once is far from enough.

Simulating a year does not equal having worked reliably in the real world for a year.

But this result sends a clear signal: the next hurdle for agents is to consistently turn correct judgments into correct actions.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.