AI system generates 453 mathematics research manuscripts, including 6 conjectures proposed by AI

icon MarsBit
Share
AI summary iconSummary
A project announcement from Xi’an Jiaotong University reveals that an AI system named AI Has Taste, developed by Zijian Zeng, has generated 453 mathematics research manuscripts and 2,312 pages of content. The AI identified counterexamples to published conjectures and proposed six new ones. The system employs multiple AI agents working in parallel, with human experts verifying key results. This AI development underscores a shift in AI’s role from answering questions to shaping research agendas.

AI mathematics is crossing a very subtle threshold.

In the past, we asked: Can large models solve difficult problems? Can they write rigorous proofs? Can they find counterexamples that humans have missed for decades?

Now, a question that more closely resembles the essence of research itself has emerged: when a proof pathway fails, does the AI know what to do next? Should it continue computing, switch paths, narrow the proposition, or even propose a new, falsifiable mathematical conjecture?

Researchers from Universiti Tunku Abdul Rahman, led by Zai Jian, have built AI Has Taste and are working to turn this into a sustainable research system.

The project's public repository currently contains 453 mathematical research manuscripts totaling 2,312 pages, with six explicitly categorized as "AI-Proposed Conjectures."

The warehouse homepage provides a more direct positioning: From Answer Generation to Research-Agenda Generation—From generating answers to generating research agendas.

Mathematical research

AI refuted the conjecture in the paper, and the original author confirmed this in their reply.

AI Has Taste does not start from the belief that "AI is inherently good at math." Zeng Zijian recalls that at the project’s outset, he was uncertain whether large models could truly advance mathematical research. A chance test became the turning point: he directly fed a conjecture from a research paper to the AI, merely curious about how far the model could go. Instead of continuing to "prove" the conjecture as the paper did, the AI constructed a counterexample, refuting the conjecture outright.

His first reaction was not excitement, but skepticism. He therefore compiled his counterexample and reasoning into a letter and sent it to the original paper’s author for verification. The subsequent reply confirmed that the counterexample was valid, and the corresponding research was later included among the more than 450 mathematical manuscripts currently in existence.

At first, I didn’t believe in AI’s power in mathematics—until I input a conjecture from a paper, and AI immediately provided a counterexample disproving it. I immediately wrote to the original author, who confirmed it was correct. After that, I truly let loose and went all in. — Zeng Zijian

Mathematical research

Mathematical research

This email didn’t change a hypothesis—it changed the scale of the project. If AI can not only restate known knowledge and generate text that appears like a proof, but also actively challenge new conjectures in papers and find counterexamples verified by the original authors, then it has the potential to truly enter the “research loop.” Only then did Zeng Zijian begin gradually agentifying each step—topic selection, proof construction, counterexample search, failure management, independent review, and writing—and enabled multiple machines to run these tasks in parallel over the long term.

True change isn't about "how many articles you've written"

At first glance, 453 manuscripts are certainly eye-catching. But if you reduce this project to merely "AI generating math papers at scale," you'll miss the most noteworthy aspect.

The project repeatedly emphasizes that 453 refers to the number of research manuscripts, not that all 453 original open problems have been fully solved. The outputs include complete proofs, counterexamples, partial results, finite computational certificates, structured reductions, and papers proposing new conjectures. Internal AI review does not equate to external peer review.

The real change occurs along the research pipeline. Traditional AI benchmarks typically begin with a pre-defined problem: humans select the question, AI solves it, and the outcome is deemed a success or failure. AI Has Taste extends the process both forward and backward—AI can screen candidate problems, compare approaches, and actively seek counterexamples; if one approach fails, it can analyze the failure structure, revise the proposition, and pass the new mathematical object to another independent agent to continue the attack.

Mathematical research

The key to AI Has Taste isn't "more prompts," but turning failures into input for the next round of research.

Topic selection, verification, error detection, and writing agents

This system is not "a single model asking and answering itself within one chat box." The public README breaks down the Agent role into four layers:

The Supervisor/Screener is responsible for selecting candidates, comparing nearby results, and identifying the most cost-effective decisive experiment;

The researcher is responsible for proofs, counterexamples, and precise calculations;

The Fresh-context Reviewer specifically targets the proposition, key lemmas, and certificates after freezing;

The Writer/Release Checker is responsible for clearly defining the final approved scope and verifying references, build materials, and release documentation.

Zeng Zijian further deployed the workflow for multi-machine collaboration: different machines can independently handle proof generation, counterexample search, exact computations, and independent peer review. What is shared are the frozen propositions, experiment logs, failure reasons, code, and evidence—not allowing a single agent to generate results and then self-approve them.

Mathematical research

Core design of Multi-Agent: Role separation + Shared evidence + Fresh-context review.

There is a line in the warehouse that perfectly captures this design: Criticism is part of the engine, not an afterword. Criticism is not a step added after finishing a paper, but an integral part of the engine of research.

Mathematicians, roboticists, and automation experts join the verification chain.

As the number of research manuscripts grew from dozens to hundreds, another issue became acute: who will determine whether these Agent-generated outputs are not systemic hallucinations?

Inside AI Has Taste, the fresh-context Reviewer separates "research" and "error detection," but the project does not treat AI self-review as the final step.

As the work progressed, Zeng Zijian also began inviting real scholars from various fields to participate in validating, reviewing, and discussing some of the results.

According to the project team, collaborators and reviewers involved in parts of the work include: Academician Ong Seng Huat, Emeritus Professor at the Institute of Mathematical Sciences, University of Malaya; Academician Kurunathan Ratnavelu, Emeritus Professor at the Institute of Mathematical Sciences, University of Malaya; and Professor Xiong Yonghua from the School of Artificial Intelligence and Automation, China University of Geosciences (Wuhan).

It is important to distinguish between “participating in verification” and “endorsing all results”: the aforementioned scholars did not individually sign off on each of the 453 manuscripts, but rather verified, reviewed, or discussed certain issues, proofs, calculations, or research directions within some of them.

For a highly automated research system, this combination of "internal red teaming + human expert sampling/review" is even more critical than entrusting all reviews to a single agent.

AI pushes the pace of research to its limits; human experts ensure that what appears plausible is refined into what is truly trustworthy at critical junctures.

After failing, I asked a better question.

The project that best exemplifies "research taste" is not a successful outcome, but a failure.

On the fractional Gaussian trace problem, the system initially attempted monotonicity and one-sided pairing approaches, but precise counterexamples kept arising. Typical benchmarks would simply label this as FAILED at this point.

But this workflow did not stop. The agent continued to ask: What exactly did the counterexample break? Is there a stable structure to the failure? Eventually, the research shifted focus to a fixed negative defect, constructing ten corrective structures and proposing a new unified boundedness conjecture. More crucially, this new conjecture was subsequently challenged by another round of precise calculations and independent review. The public archive scans to index 1004; the unified bound remains unproven to this day.

The significance of this is not that "AI has proven another major theorem"—in fact, it hasn't. The significance lies in the fact that while the original problem remains unsolved, the failed approaches have been condensed into a new, clearer, more falsifiable proposition that, if proven true, could reconnect to the original problem. The research did not end in failure; instead, it grew the next question from within that failure.

Past: Humans posed questions, and AI answered. Present: AI not only answers, but also helps determine the next question.

6 New AI Hypotheses

As of the current public snapshot, the repository categorizes six papers under "AI-Proposed Conjectures." These works span combinatorics, graph theory, number theory, and analytic problems, including: the nonsingularity of the infinite binary-kernel matrix for three subsequences of A182510, CDS-colorability of Young diagrams, decomposition of triangle-free graphs with maximum degree at most 6, unary-theta congruence formulas for 18-colored generalized Frobenius partitions, nonnegativity of Newton coefficients for reciprocal-subsum denominators at dyadic stages, and the aforementioned fractional Gaussian trace fixed-defect conjecture.

What truly matters is not whether AI can "come up with a wild guess." It's easy to casually make a guess. The challenge lies in writing the guess as a precisely defined hypothesis, explaining its origin, identifying the evidence that supports it, specifying what counterexamples could disprove it, outlining the implications if it holds true, and providing the calculations and source code for future verification.

This is why the project considers "formulating a hypothesis" a higher-level research action than "generating an answer": answering an existing question validates what is already known, while a hypothesis can redirect future research resources.

Behind the 453 manuscripts lies an early form of "Agent-scaled Research."

AI Has Taste has publicly released 453 research manuscripts, totaling 2,312 pages; among them, the BigWIN2 workflow corresponds to 223 manuscripts, each accompanied by a paired LaTeX and reproduction materials package.

The most counterintuitive aspect of this scale isn't that "AI types quickly." What truly makes mathematical research expensive is context switching: one problem requires graph theory, the next demands number theory, and the one after that might need SAT solving, brute force, exact rational arithmetic, or matrix computations. A human researcher cannot simultaneously maintain hundreds of entirely different lines of thought, but an agent system can break them into independent tasks and save failed paths, local theorems, and reproducible experiments.

Thus, the role of individual researchers is evolving: rather than personally driving every step, they are becoming more like PIs—defining research boundaries, selecting problem pools, deciding which results warrant further investment, when to stop, and which conclusions are worthy of public disclosure.

Why is it called "AI Has Taste"?

"AI has taste" sounds anthropomorphic, but the project's README explicitly defines the boundary: here, "taste" does not refer to consciousness or subjective experience, but rather to "mathematical judgment demonstrated through decision-making."

For example: Both problems can be tackled—which one should you do first? If a path has already made some progress but the marginal returns are diminishing, should you stop? Is a counterexample enough to kill a proposition, or does it suggest that the true theorem should be rephrased? Is a small result sufficient to stand on its own as a paper? These decisions were traditionally embedded in researchers’ experience and difficult to measure with standard benchmarks.

AI Has Taste seeks to create an auditable trail of these decisions: which questions were selected, why the approach was changed, what structures remained after failure, where the next proposition originates, and how new conclusions are challenged again within independent contexts.

How far are we from an autonomous mathematician?

The most commonly exaggerated aspects of this project also need to be clearly stated: 453 manuscripts are not 453 peer-reviewed journal papers; internal Agent reviews are not equivalent to independent peer review by the mathematical community; and limited computational coverage cannot be automatically extended to infinite cases. These limitations are explicitly stated in the repository itself.

But ignoring how the research is organized simply because these results have not yet been fully externally validated would cause you to miss another shift: the boundary of AI’s capabilities is moving from “executing a given task” to “participating in designing the next task.”

The truly scarce resources in mathematical research have never been just computational power and proof length—they also include: which problems are worth spending time on, which failures are worth preserving, and when it’s time to switch topics.

As AI begins to enter these stages, competition in "AI for Math" will no longer be just about who can prove more, but will gradually shift to who has a better research feedback loop and who can extract more promising directions from a large number of failures.

One More Thing

Past AI research narratives resembled a super student: the teacher posed the questions, and it was responsible for solving them.

What AI Has Taste aims to do is more like a research team: humans set the boundaries and make value judgments, while the agent identifies problems, attempts solutions, makes mistakes, overturns its own conclusions, leaves behind partial results, and then proposes the next question.

If this trend continues, the most important human-AI division of labor in future basic research may evolve from “humans think, AI calculates” to a more complex structure: humans will be responsible for goals, values, and ultimate accountability; AI will handle large-scale search, verification, failure management, and generation of candidate research directions; formalized tools and public evidence will ensure results are verifiable.

After AI can solve problems, the next challenge might truly be: Can AI choose the problems?

This article is from the WeChat public account "AI New Era," authored by AI New Era.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.