OpenAI's LLM Harness Solves 9 Theoretical CS Problems, Aims for Scientific Expansion

iconCryptoBriefing
Share
AI summary iconSummary
A team of researchers used a prover-verifier loop with GPT-5.5 Pro and Claude Opus 4.8 to solve nine open problems in theoretical computer science and mathematics. The results were published between June 27-30, 2026. The team, led by Binghui Peng from the University of Maryland, plans to expand the method to all scientific fields. This crypto news highlights the growing intersection of AI and scientific research.

A team of researchers just used a pair of competing large language models to solve nine open problems in theoretical computer science and mathematics. The approach, called an “LLM harness,” uses GPT-5.5 Pro as the solver and Claude Opus 4.8 as the verifier in a prover-verifier loop. The results were published around June 27-30, 2026.

Of the nine problems, four came from the Conference on Learning Theory (COLT) problem list, one from the Foundations of Computer Science (FOCS), and four from commutative algebra.

Advertisement

Omri Weinstein, a former NVIDIA researcher who highlighted the project on June 30, noted that one of the solved problems had been his personal open question for two years.

The research team was led by Binghui Peng from the University of Maryland, alongside Runzhou Tao, Steven Wang, and Hantao Yu. Peng brings a resume that includes stints at Columbia, Google, and Stanford.

How the prover-verifier loop works

In the prover-verifier setup, GPT-5.5 Pro generates candidate proofs or solution approaches, then Claude Opus 4.8 evaluates them for correctness. When the verifier finds flaws, it sends feedback to the prover, which refines its approach. This cycle repeats until the verifier accepts the proof.

This builds on a foundation that OpenAI laid back in July 2024, when the company published a paper on “prover-verifier games” that focused on making LLM outputs more legible and verifiable. By December 2025, the approach had matured enough that GPT-5.2 Pro was already tackling a complex challenge in statistical learning theory. The jump from one problem to nine, across multiple mathematical domains, represents a meaningful scaling of the method’s ambitions.

What this means for researchers and investors

The team has explicitly stated plans to extend this method across various scientific fields.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.