Today, 99% of the world's human population cannot outperform an AI in terms of IQ.
In the latest offline IQ test by Tracking AI, multiple versions of GPT-5.6's "full suite" all scored 136.

This is the first time an LLM has pushed IQ past the 130 threshold.
In the distribution of human IQ, 130 is the threshold for "genius," and only about 1% of people worldwide reach this level.
In other words, GPT-5.6 is smarter than 99% of humans.

GPT-5.6 scores a staggering 136, breaking the "genius threshold" for the first time
How valuable is this "IQ"?
In fact, Tracking AI has two sets of questions.
One is a public Mensa Norway-style test, available online for anyone to take, and the model has already scored over 140.
Another set is its own privately compiled "offline question bank"—not publicly disclosed and designed to prevent question leaks, specifically to close the loophole where models might have memorized answers in advance.
GPT-5.6 scored 136 on this offline exam, which is the most difficult and anti-cheating version yet.

On this offline leaderboard, the various versions of GPT-5.6—including the vision version—collectively reached a score of 136, leaving all other competitors far behind.
Closely behind is Claude-5 Fable with 130 points.
Further down, names like GPT-5.6 LUNA Max and Claude-4.8 Opus are still hovering between 117 and 123 points.
Keep in mind that the 130 threshold had never been breached before.
Over the past year, model after model—from o3 to various flagship models—has surged forward, all stalling at the threshold of 130, none truly entering the "genius range."
GPT-5.6 was the first to kick down the door.
And it wasn’t just one player carrying the load—SOL, TERRA, and the entire family surged together to 136, with no visual version left behind.
On Reddit, a developer tested it firsthand and ultimately felt that GPT-5.6 was noticeably smarter than GPT-5.5.

In the following test, GPT-5.6 achieved an outstanding score in the shortest possible time.



A score might be hard to believe—so what does GPT-5.6 look like when it leaves the exam hall and gets down to real work?
Having splits isn't enough—get GPT-5.6 to work.
Developer Amir Bohlooli fed the same physics simulation prompt to both Fable 5 and GPT-5.6 Sol, expecting Fable to dominate—only to be surprised by GPT.
It chose particle fluid simulation, with physics advancing in real time rather than performing fixed calculations per frame, and everything—CSS, interface, and rendering—is packed into a single HTML file.
It also automatically hosts it as a shareable webpage. One sentence, one finished product.

Ramanpal Singh used the same prompt to create a RAG-based customer support ticket system.
Four roles, a management dashboard, embeddable components, plus automatic complaint categorization, sentiment detection, and draft response generation.
This app created five of them at once, for a fraction of the cost of a Fable 5.
The most vivid scene is Claire Vo’s segment.
A few days ago, she was stuck on a bug, thinking she had broken her own code, but after switching to GPT-5.6 Sol, she simply said, "I refuse to believe this can't be solved."
Fixed Sol in one go, and accidentally got other models running too.
Her assessment was spot-on: Fable’s obsession with absolute technical precision trapped itself in its own constraints, while Sol’s practical approach got the job done.

It must be said that there is an entire real-world project between AI being able to solve problems and AI being able to handle actual situations.
Is this AGI?
Some netizens have said, "For 99% of people, this is already AGI."
Looking calmly at it, these 136 points come from a specific offline / Mensa Norway-style test run by Tracking AI.
It primarily measures standardized cognitive abilities such as abstract pattern recognition and logical reasoning.
The issue is that IQ tests were not originally designed for large models.
A Mensa test cannot measure the factual reliability of a model, its ability to call tools, or how dependable it truly is in real-world professional scenarios.

It merely slices off a thin piece of "intelligence" and tells you how bright that piece is.
However, real-world user experiences have provided the other half of the answer: GPT-5.6 appears to be gradually merging the ability to solve problems with the ability to get things done.
In standardized tests, models have often seen the questions thousands of times in their training data; what truly tests their ability are the novel questions they’ve never encountered and can’t copy answers from.
Only those who can stay calm there deserve the title of "intelligence."
Reference materials:
https://x.com/davidpattersonx/status/2077049232490672458
https://trackingai.org/
This article is from the WeChat public account "New Intelligence Yuan," authored by ASI Revelation.
