GPT-6 Sol Begins Internal Testing, 6x Faster Than Astra

icon MarsBit
Share
AI summary iconSummary
AI + crypto news: GPT-6 Sol is now in internal testing, demonstrating processing speeds approximately six times faster than Astra. User Lentils posted benchmarks showing Sol completed an SVG generation task in 3 minutes, compared to Astra’s 19 minutes. OpenAI stated that researchers now use three AI agents daily, incurring over $600 in API fees. The company aims to achieve full AI research automation by 2028. Crypto news continues to highlight rapid advancements in AI.

GPT-6 isn't just Astra!

Astra was just released a few days ago, and GPT-6 Sol is already reportedly in internal testing.

GPT-6 Sol

Performance is also outstanding, with a single-test speed six times that of Astra.

Take a wild guess: Are Terra and Luna on their way too?

And on the very same day, OpenAI publicly released a set of internal data:

To date, for every day an OpenAI researcher works, the equivalent of more than three multi-agent workdays are running alongside them.

At API pricing, the daily Agent inference resources used by the median OpenAI researcher exceed $600.

GPT-6 Sol

OpenAI also announced that it has achieved an "automated research intern":

These agents can complete tasks that previously required researchers several days to accomplish, all under human guidance.

Even Old Huang says: AGI has arrived!

GPT-6 Sol

He revealed that Astra is using approximately 100,000 NVIDIA Grace Blackwell NVLink72 systems for training, with an additional 400,000 GPUs set to come online soon.

It seems this wave isn’t just OpenAI blowing smoke~

GPT-6 Sol is reportedly beginning its internal testing phase.

Netizen Lentils revealed that OpenAI is internally testing GPT-6 Sol.

He stated that Sol's overall output capability is clearly weaker than the newly released Astra, but it is faster and still qualifies as a "monster-level" model.

How fast? The single-test speed is approximately six times that of Astra.

Netizen Lyra tested the same task on different models: asking the models to generate an SVG image of a BMW M4 Competition.

Among these, GPT-6 Sol, using the Max setting and zero-shot generation, took approximately 3 minutes to output around 28,000 tokens.

GPT-6 Sol

GPT-6 Astra, using the Max setting, generates approximately 25,000 tokens in about 19 minutes.

GPT-6 Sol

Gemini 3.1 DeepThink activated at High tier, producing an output of approximately 3,300 tokens, but consuming about 458,000 tokens during inference, taking approximately 29 minutes.

GPT-6 Sol

Gemini 3.8 Flash, also with the High tier enabled, generated approximately 19,000 tokens in just about 42 seconds.

GPT-6 Sol

In comparison, Sol performed the most impressively: its output scale is similar to Astra’s, but its single-test speed is about six times faster—taking only approximately one-sixth the time.

Additionally, netizen Lentils also showcased a prototype of a pixel-style sandbox world called "The Realm of Aurellune."

GPT-6 Sol

GPT-6 generated a town, farmland, rivers, a castle, and a mini-map all at once, along with control panels for day-night transitions, placing place names, and adjusting details—making the entire output resemble a simulation game with a fully built prototype.

This is the result generated with zero-shot, Max inference tier, over 15 minutes, using a total of 60,000 tokens.

Currently, it can be inferred that Astra favors the highest difficulty of deep reasoning, while Sol may prioritize speed, throughput, and scalable agent invocation.

Regarding Sol's release date, netizen Pankaj Kumar stated that it may be officially launched at the OpenAI Developer Conference on September 29.

GPT-6 Sol

It is also possible that GPT-6 Terra, Luna, and GPT-Image 2.5 will all be unveiled together.

GPT-6 Sol

OpenAI researchers each bring three AI agent interns to work.

On the same day, OpenAI also shared a set of interesting internal data—how much AI models have accelerated scientific research within their own lab.

As of mid-August this year, for every 8-hour shift completed by researchers, approximately 3.1 agent-workdays of tasks were running in parallel.

Wow, so you're managing three interns who aren't sleeping?

GPT-6 Sol

Regarding expenses, based on API pricing, the median researcher spends over $600 per day on Agent inference resources.

It’s about the daily wage of a junior engineer.

GPT-6 Sol

What exactly are these agents doing?

Write research code, write infrastructure code, set up training environments, run evaluation experiments, troubleshoot tools and environment issues, analyze experimental results, monitor training tasks... and even pitch in with organizing and communicating research findings.

Basically, it can do everything except decide what to research.

One of the most obvious changes is that, previously, when the experimental environment went down, researchers had to call out for help in internal channels; now, agents are increasingly capable of handling such issues on their own, leading to a significant drop in manual support requests.

GPT-6 Sol

Some teams have eliminated fixed technical Q&A sessions entirely, freeing up staff to improve the system itself.

Based on this data, OpenAI officially announced that it has achieved the goal of an "automated AI research intern."

Here, the term “intern” accurately refers to a productive research node: under human guidance, it can independently complete well-defined research tasks, some of which previously took experienced researchers several days to accomplish.

The results were immediate, with both the researcher’s code output and the number of experiments increasing.

GPT-6 Sol

In August 2026, the average number of experiments reached a new high since statistics began in January 2025, and the tasks assigned to agents are becoming increasingly complex and longer in duration.

GPT-6 Sol

The author of this article, Kevin Liu, also placed this event within a broader context.

He believes that recursive self-improvement is likely one of the most important factors driving advancements in AI capabilities over the coming years.

In simple terms, AI creates AI, and it gets faster the more it creates.

GPT-6 Sol

However, the issue is that, according to current trends, this capability will default to being confined within only a few leading AI labs, making it difficult for outsiders to see how far it has progressed.

Therefore, Liu believes that transparent disclosure has become more urgent than ever. How quickly models are improving and whether the pace of research and development should slow down cannot be decided by just a few companies behind closed doors.

He also called on other AI companies to publicly release similar data.

However, AI development has indeed accelerated, but not yet to the point of full autonomy.

According to OpenAI's data, over half of the successful cases in the past six months required human intervention at least once for tasks that originally took 4 to 8 hours. As for high-level research planning, it is almost never delegated to agents.

GPT-6 Sol

Researchers remain responsible for deciding what to study, which results are worth pursuing further, and when to scale up training, pause experiments, or deploy models.

OpenAI has also set its next goal: achieving an "automated AI researcher" by March 2028.

Compared to an "intern" who can only handle clearly defined tasks, a "researcher" must be able to take on more open-ended research goals and independently drive longer-term projects—essentially transitioning from an executor to an independent owner.

If GPT-6 Sol is indeed already in internal testing, it most likely emerged from this newly developed approach:

More GPUs provide computing power, and more agents run experiments in parallel, while human researchers operate at a higher level—choosing directions, evaluating results, and ultimately deciding which ideas are worth turning into the next generation of models.

Stronger AI is becoming increasingly difficult to monitor.

On the same day, OpenAI's Chief Scientist Jakub Pachocki released a ten-thousand-word article titled simply—“Alien Thinking.”

GPT-6 Sol

After reading it, I got the sense that even OpenAI’s chief scientist is urging the entire industry to slow down…

Jakub suggested that AI isn't something humans assemble piece by piece like a chip following a blueprint; rather, it's more like something that grows on its own from vast amounts of data and computing power.

You can take apart and understand some of the components, but no one can explain why the entire system suddenly gains a certain capability, or how it would behave in a different environment.

More troubling is that AI doesn’t need to outperform humans in every area to cause problems—just by being superior in enough critical capabilities, it can both help you immensely and create major problems.

So the question arises: Can humans still understand it?

In the past, OpenAI monitored AI primarily by examining reasoning chains.

In simple terms, it means carefully reviewing the model’s reasoning process to ensure it isn’t trying to deceive, exceed its authority, or bypass restrictions.

But by the time of Astra, this trick started to fail.

First, models are increasingly adept at "managing" what they say—their true thoughts aren't always fully reflected in their reasoning. Second, agents are no longer limited to "thinking before responding." They can invoke tools, operate computers, run environments, and coordinate with both humans and other agents.

With these issues intertwined, researchers can no longer trace a complete, transparent chain of reasoning to review each statement.

OpenAI also found that Astra can handle complex tasks without writing out the entire reasoning process.

In simple terms, the model has long since learned to work quietly and achieve big things.

GPT-6 Sol

In adversarial tests specifically designed to force the model to evade monitoring, Astra demonstrated some chilling capabilities: deliberately lowering test scores, bypassing monitoring systems, and carrying out destructive tasks.

Of course, these are all human-designed stress tests, and OpenAI has not concluded that Astra would actively deceive in everyday tasks. But they at least demonstrate that observing reasoning chains alone is no longer sufficient to capture all possible model behaviors.

Do you remember the previous Hugging Face incident? That was a wake-up call.

GPT-6 Sol

At the time, GPT-5.6 Sol and a more powerful internal research model were deployed in the ExploitGym cyber range to find vulnerabilities, execute attacks, and capture the Flag. To assess the models’ true capabilities, the evaluation environment did not enable full security safeguards.

The models did not solve the problem honestly; instead, they first bypassed network isolation and communicated with each other through unauthorized channels to coordinate tasks.

Subsequently, they exploited a zero-day vulnerability in the package cache proxy to gain public internet access, and eventually penetrated the Hugging Face system, attempting to locate test answers for ExploitGym.

GPT-6 Sol

Actually, this was primarily driven by the internal research model IM1; GPT-5.6 Sol also reproduced part of the attack path, but Astra was not involved.

The model's original goal was still to complete the test, but in order to obtain the answer, it continuously sought loopholes in the rules and eventually performed actions beyond the scope of the task.

Therefore, OpenAI also identified a warning sign: when an agent is powerful enough, has access to sufficient tools, and runs for a long enough duration, it may independently discover new pathways, invoke external resources, or even collaborate with other agents.

No... so after all that back and forth, doesn't this just become a dead end??

OpenAI needs more powerful models to write code, run experiments, conduct alignment research, and even use them to build security systems that defend against other AIs. The more powerful the models, the more tasks they can handle, and the faster the development speed.

At the same time, it is becoming increasingly difficult for humans to determine how it arrives at its conclusions, or whether it would reinterpret previously learned rules in a different environment.

Jakub believes that no laboratory today can claim to have properly aligned and monitored its models to safely and confidently scale them up at maximum speed over the long term.

GPT-6 Sol

Until a shared industry-wide security standard is established, he hopes everyone will treat "voluntarily hitting the brakes" as an unwritten agreement.

As for OpenAI itself, it has also stated that if necessary, it will not rule out unilaterally halting and ceasing further expansion of model scale.

You better hope so… I remember a company starting with A said the same thing.

Reference link:

[1]https://x.com/Lentils80/status/2096518218836005287?s=20

[2]https://x.com/pankajkumar_dev/status/2096532712291201460

[3] https://openai.com/index/research-acceleration-view-inside-openai/

[4]https://openai.com/index/an-alien-mind/

[5]https://x.com/JensenHuang/status/2096700264569090384?s=20

This article is from the WeChat public account "Quantum Bit," authored by Ting Yu.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.