Too competitive.
Only three days into September, five cutting-edge models have already been released!
Anthropic’s Fable 5.1 and Mythos 5.1, Google’s Gemini 3.8 Flash and Cyber, and Meta’s Muse Spark1.3… RSI has surely arrived.
Last night, just as Google's Gemini 3.8 Flash became the lightweight champion, Meta showed up to challenge it.
Zuck proudly announced on X: Muse Spark 1.3 is officially launching today, delivering an experience so cost-efficient it’s almost free, powered by cutting-edge performance. This is our biggest leap yet in programming and agent workflows!

Alexandr Wang, head of Meta’s Superintelligence Lab, also posted three X posts revealing the core breakthrough behind this "game-changing" development.
He stated that compared to Muse Spark 1.2, Muse Spark 1.3 reduces invalid rounds, decreases tool calls by approximately 20%, reduces token consumption by approximately 25%, and better maintains task requirements over long-duration tasks—“users will clearly feel this leap.”

As soon as Muse Spark 1.3 was released, Gemini 3.8 Flash, which was just launched last night, seemed a bit awkward in comparison.
The benchmark results show that this Muse Spark 1.3 improvement is greater.
It has not only surpassed GPT-5.6 Sol and Fable 5 in benchmark scores, but its Artificial Analysis IQ under Max inference strength has reached 62, firmly placing it in the current top tier.

Over five months, Meta has quietly rolled out four versions of Muse Spark.
Alexander Wang mocks Google: "drinking the exhaust of other models"
Recently, the top tier of AI has become extremely crowded. GPT-5.6 holds the throne, Fable 5.1 has surged ahead, and Gemini 3.8 Flash is overtaking from the inside.
The emergence of Muse Spark 1.3 has thrown everyone into disarray.
On the DeepSWE v1.1 benchmark, which best reflects a model’s true long-range programming capabilities, Muse Spark 1.3 outperformed Opus 5 and GPT-5.6 Sol, and surpassed the newly released Gemini 3.8 Flash, achieving the highest score to date.
Muse Spark 1.3: 75.4 points
Claude Opus 5: 74.0 points
GPT-5.6 Sol: 73.0 points

Moreover, Muse Spark 1.3 also demonstrated impressive performance in several other key developer metrics.
In the SWEAtlas CodeBase QnA, it scored 59.4, surpassing GPT-5.6 Sol and Opus 5.
On Terminal-Bench 2.1, it scored 88.8.
On MRCR 512K-1M, it achieved an astonishing score of 98.1! (For comparison, GPT-5.6 Sol scored only 73.8 on this metric—completely outclassed.)

In terms of agent capabilities, it has nearly matched the state-of-the-art, with only a few percentage points separating it from Opus 5 on benchmarks such as JobBench and OSWorld.

On Artificial Analysis, Muse Spark 1.3 significantly outperforms Gemini 3.8 Flash.
On the Smart Index, it scored 62, tying with Claude Fable 5 and placing among the top tier.

In terms of cost, which is of greatest concern to developers, Muse’s input cost is 8 times lower than Fable 5, and its output cost is nearly 12 times lower, delivering exceptional value for money.
Facing such impressive results, Meta’s Chief AI Officer, Alexandr Wang, sarcastically remarked: “On the Artificial Analysis intelligence index, Gemini is nothing but a tailpipe emission from other models.”
Google is really down on its luck.

Some jokingly said: “Google worked hard for four months and released four Gemini Flash versions, while Meta rolled out four Muse Spark⁺ updates in five months. But just hours after Google took the lead, Meta overtook it—this plot is too thrilling.”

Less is More: A Performance Beast Running for 2 Hours on Just $1
High benchmark scores are impressive, but in today’s AI community, what truly excites developers is something that’s both easy to use and affordable.
Muse Spark 1.3 is called the king of value for money because it not only runs fast but also saves you money.
As Alexandr Wang stated, Muse Spark 1.3 is "so cheap it's negligible," with "cutting-edge performance at a fraction of the cost."


It has taken a completely different path from previous large models that relied on brute force and excessive parameter stacking—it has focused on simplification.
Muse Spark 1.3 was specifically trained for long-horizon reasoning and improved instruction following to reduce unnecessary interaction rounds and produce more concise outputs.

It has become smarter and more efficient, with cleaner code and less unnecessary text.
The direct result of this reduction is a dramatic drop in costs.
Here is a truly striking real-world test case.
Developer @SPAC89 compared Muse Spark 1.3 Ultra Contributor mode with Fable 5.1 xHigh.

The prompt he provided includes a strict "self-improvement rule": if the independent evaluator's score is below 9.5/10, the agent must continue improving the code and retry.
Both models ran for approximately two hours, with astonishing results.
Muse Spark 1.3 is like an endlessly tireless super worker—it just won’t stop, completing a full 20 rounds of self-improvement cycles, launching three agents each time—equivalent to over 60 agent runs within just two hours.
And the total cost to complete this incredibly large and complex long-range reasoning task was less than $1!

The developer exclaimed, "This completely exceeds my expectations—it's truly a work of art!"
In terms of macro pricing, Meta has demonstrated strong commitment. The new model offers increased capacity without any price increase, maintaining the pricing at $1.25 per million tokens for input and $4.25 per million tokens for output.

In terms of pricing, the contributor version of Muse Spark 1.3, "muse-spark-1.3-contributor," sets the floor for foreign model pricing. (The trade-off is that your data will be used to improve Meta's products.)

Mehul Mohan, founder and developer of Codedamn, said:
The contributor tier pricing for Muse Spark 1.3 is absurdly unrealistic—that's what American AI labs call " DeepSeek Pricing moment.

According to Artificial Analysis statistics, among models with equivalent intelligence levels (scores above 59), Muse Spark 1.3 has a single-task cost of just $0.55, making it the absolute optimal solution.
For comparison, Grok 4.6 with the same intelligence score (61) costs $0.94, GPT-5.6 Sol requires $0.95, and Claude Opus 5 is as high as $1.23.
Even more, developer Kartik noted that Muse Spark 1.3 appears to possess a "circular depth" reasoning capability similar to Sol and Fable.
It doesn't generate answers instantly; instead, it performs logical reasoning internally, resulting in remarkably high token efficiency.

Alexandr Wang's Rapid Rise, Zuckerberg's Ultimate Ambition
The emergence of Muse Spark 1.3 would not have been possible without the team behind it.
In July this year, Meta released Muse Spark 1.1, featuring a 1M ultra-long context and computer operation capabilities.
In just under two months, version 1.3 has elevated long-range encoding capabilities to the level of GPT-5.6.
This rapid iteration is precisely the result of Alexandr Wang taking over Meta’s Superintelligence Lab.
What is their ultimate goal? As Zuckerberg and Alexandr Wang repeatedly emphasize: “agents” and “cheaper than negligible.”
In other words, Meta is turning its attention to the next ASI opportunity—“Agent Engineering.”
Meta AI CEO Alexandr Wang explicitly stated in an interview with Axios that all usability improvements are aimed at realizing the ambitious vision repeatedly mentioned by Zuckerberg—to build a personal AI agent available 24/7.

For an agent to handle emails, write code, and analyze data for you 24/7, it must meet two conditions.
1. Extremely intelligent and stable: It must handle extremely long conversation threads and manage multiple workflows in parallel.
When instructions are unclear, it should proactively ask questions; when encountering obstacles, it should seek help from humans; before executing critical actions, it should confirm details. Most importantly, it must not generate hallucinations.
2. Extremely cheap: If the daily cost of running a personal agent reaches tens of dollars, it will forever remain a toy for only a few.
The emergence of Muse Spark 1.3 is essentially paving the way for 24/7 personal agents.
It’s like a seasoned, experienced engineer—you give it a goal, and it can plan, correct itself, and deliver on its own.
To this end, Sun Zhiqing, a Peking University alumnus and Meta researcher, revealed that the Meta team has retrained the previous Avocado model to achieve improved test scaling.

Behind the狂欢: Have We Been Fooled by "Rank Manipulation"?
However, Muse Spark 1.3 has also been questioned.
The renowned AI institution BridgeMind shared a thought-provoking statement:
Gemini 3.8 Flash and Muse Spark 1.3 are being released today. Both appear excellent in benchmark tests.
Muse Spark 1.3 is currently tied with Fable 5 on the Smart Index… I should feel excited. But I don’t.
Because this month’s AI releases were identical, scores soared, but real-world experience improved not at all.
This is frustrating. My trust in benchmarks is decreasing each week, and my trust in labs that manipulate benchmarks is nearly gone.

This statement precisely addresses the current pain points in the large model industry.
Some netizens conducted stress tests on Muse Spark 1.3 and Gemini Flash 3.8z under backend-level 3D constraints, and the results exposed the true limitations of both models:
Muse Spark 1.4 isn't that great, and Gemini Flash 3.5 is just gaming the rankings.

Currently, major labs have fallen into a vicious cycle of training solely to boost benchmark scores. Every model release is accompanied by radar charts and leaderboard screenshots purporting to crush competitors.
DeepSWE, Tau3-Bench, and GPQA have become the focus of intense "practice" by various teams.
Meanwhile, Muse Spark 1.3 has also revealed its true nature.
In Artificial Analysis’s in-depth report, we can also see a hint of its regression.
For example, in the AA-Omniscience test, both the xhigh and max versions saw a decline, as their "abstention rate" increased—they became more likely to refrain from answering when uncertain, rather than making up responses.
This reduced the hallucination rate, but also indirectly suggests that its absolute knowledge base has not improved as dramatically as its benchmark scores indicate.

What is the second half of the large model race really about?
Today’s release of Muse Spark 1.3 and Gemini 3.8 Flash allows us to draw the following conclusions.
First, the "parameter dividend" of foundation models is reaching its peak.
The approach of simply increasing parameter scale to enhance intelligence is experiencing diminishing marginal returns.
In the future, whoever can introduce a reasoning mechanism similar to “circular depth” into their system architecture and achieve extreme simplification in tool invocation will achieve a true leap in user experience.
Moreover, agent engineering is the key to victory.
The competition among large models has shifted from "who can talk better" to "who can get more work done."
The core barrier in the future won't be a single benchmark score, but the ability to realistically execute long-term tasks at extremely low cost.
Models like Muse Spark 1.3, which can complete 60 trial-and-error cycles for less than one dollar, will completely transform business models.
Reference materials:
https://x.com/SPAC89/status/2095284969614475744
https://research.meta.ai/blog/introducing-muse-spark-1-3
https://x.com/AIatMeta/status/2095234385129963666?s=20
https://artificialanalysis.ai/zh/models/muse-spark-1-3
https://x.com/EdwardSun0909/status/2095236891574780275?s=20
This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation.
