OpenAI, which released 722 AI-generated math manuscripts in one go yesterday, has just retracted three of them.
The reason is simple: a plus or minus sign was written incorrectly.

Just now, OpenAI released its first changelog in the openai/math repository. The changelog shows that OpenAI retracted 3 manuscripts, revised 14, and updated citations for an additional 13. The total number of manuscripts in the repository decreased from 722 to 719.

A single plus-or-minus sign undermined three papers.
The three retracted papers were all related to the Hodge conjecture, one of the Millennium Prize Problems:
Algebraicity of Weil classes on split abelian eightfolds
Algebraicity of Kuga–Satake Correspondences for K3 Surfaces
The rational Hodge conjecture for products of K3 surfaces
The issue is with the first paper. OpenAI stated in its retraction notice that, in a key argument, the paper incorrectly used +1 as the symbol for a certain geometric operation, whereas, according to its own conventions, it should have been -1.

https://github.com/openai/math/blob/main/history.md
One positive and one negative result lead to vastly different outcomes. The count, which should have canceled each other out and ultimately reached zero, became a non-zero number. A classic theorem upon which the paper relies requires precisely that this count be zero. With this condition unmet, the entire subsequent construction loses its foundation.
The other two papers also borrowed this construction and were thus retracted together.
However, OpenAI emphasized the same statement in all three retraction notices: the retractions apply to the proofs, not to the mathematical propositions themselves, which are not necessarily incorrect. The original papers have not been deleted and remain accessible via archival links.
These three papers all belong to Repository Outcome Family 032, which corresponds to the previously most-watched results related to the Hodge Conjecture.
After the retraction, the name of this family changed: from "Hodge and Kuga–Satake Results for All Projective K3 Surfaces" to "The Rational Hodge Conjecture for CM Abelian Varieties."
In other words, the part concerning K3 surfaces has been withdrawn from the list. The core claim of this family—that the Hodge conjecture for rational Hodge classes holds for all complex CM abelian varieties—remains intact, and the new version of the paper only updates the citations.
It is worth noting that the erroneous paper was dated September 18 and was among the earliest manuscripts in the repository. Nearly three weeks passed between the date stamp and its public release, after which it was retracted within one or two days. OpenAI did not disclose who discovered the error.
14 more have been revised
In addition to retracting the paper, OpenAI revised 14 other manuscripts, including patching proofs, correcting the wording of conclusions, and clarifying assumptions and dependencies.

These 14 papers cover Lipschitz heights and the Ashkin–Teller model in statistical physics (4 papers), the Kähler minimal model program and ampleness in complex geometry (6 papers), symplectic geometry (2 papers), general computation for Navier–Stokes fluids (1 paper), and one paper on the BSD formula with outdated references removed.
One revision is particularly representative. Result Family 342 claimed to prove the "tame pushout compatibility" conjecture proposed by Simon Donaldson. The main conclusion remains unchanged and has been formalized in Lean. However, a stronger auxiliary conclusion in the paper was found to be overstated: the revised version narrows it to specific cases and adds a counterexample showing it does not hold in general.
Formalization rate: 42%
The update also added 6 new Lean formalizations and 5 supplementary auxiliary results. According to OpenAI’s metrics, the main results of 300 manuscripts have now been formalized, accounting for approximately 42% of the total 719 manuscripts.

Several entries in the newly added formalizations are significant, such as the 103rd result family, which claims to prove that logarithmic space computation can be fully derandomized (L = RL = BPL).
It should be noted that our previous report of "approximately sixty percent" was based on family-level statistics, meaning that if any paper within a family included formalization, the entire family was counted; OpenAI’s reported 42% is based on individual main results, better reflecting the actual proportion verified line by line by computers.
There was an error because Lean did not review it.
Comparing with the repository directory reveals a pattern: the content that was withdrawn or substantially revised almost all falls within sections that lack Lean formalization.
This also confirms the earlier warning from the mathematical community: results that have not been formalized or peer-reviewed can only be considered "claims" at this stage.
Fortunately, this correction was handled in a standardized manner: the error was clearly stated, the dependency chain was explained, and the previous version was retained for reference, in line with the "change tracking" principle recommended by the Institute for Advanced Study’s Mathematics and Artificial Intelligence Advisory Group.
However, just two days after over 700 manuscripts were made public, the first correction notice has already arrived—and it’s unlikely to be the last.
This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), edited by Panda.
