TMLR editor finds authors unfamiliar with their own AI-written papers

icon MarsBit
Share
AI summary iconSummary
TMLR editors found authors unfamiliar with their own AI-written papers, as reported by MarsBit. The journal tested 10 desk-rejected papers by asking authors basic questions. Most struggled to explain their work, with three unable to describe core elements. The experiment raises concerns about AI-assisted academic writing in crypto news and AI + crypto news.

Edited by Panda

You wrote a paper; now I ask you three questions: What is the problem statement of your paper? What does this symbol represent? In which part of the main text is the conclusion mentioned in the abstract supported?

If the paper was truly written by you, then these questions would be straightforward—even answerable without any preparation.

However, in a recent small-scale trial by the machine learning journal TMLR (Transactions on Machine Learning Research), the authors of three papers could not answer these basic questions.

arXiv

On September 16, TMLR published an article titled “Asking Authors About Their Own Papers” on its official blog, authored by Nihar B. Shah, one of the journal’s co-editors-in-chief at Carnegie Mellon University (CMU). During his two-week editorial rotation, Shah selected 10 submissions that would have otherwise been desk-rejected and instead reached out to each author for a conversation. In the end, all 10 papers were still rejected.

arXiv

https://medium.com/@TmlrOrg/asking-authors-about-their-own-papers-3d2e04e5dee0

Gautam Kamath, co-editor of TMLR, shared on 𝕏 that this was an “audacious experiment” and said it confirmed what many had long suspected: some submitters don’t actually understand what’s in their own papers.

arXiv

Reviewed 10 articles pending rejection; all were rejected.

According to Shah’s account, the trial took place during his tenure as acting editor-in-chief from August 14 to 28, 2026. He informally selected 10 papers that were pending immediate rejection and sent authors a brief message via the review platform OpenReview: “An editor-in-chief would like to speak with you before peer review to better understand your paper; if convenient, please email to suggest a time.”

After the message was sent, things quickly diverged. One author retracted their paper outright; another replied that they were too busy to take a call; eight others scheduled meetings, but one author failed to show up.

The person who ultimately met with Shah is the author of seven papers. They have diverse backgrounds, including undergraduate students, master’s students, PhD candidates, university faculty, and independent researchers. Most of these papers are single-authored, but not all.

arXiv

Shah’s questions fall into two categories: one concerning foundational issues about problem setups, notation, and the claimed results in the paper; the other focusing on detailed questions regarding specific technical expressions, theoretical results, and experimental design choices.

Out of the seven interviews, only one author answered all questions correctly. Three other authors could explain the high-level concepts but struggled when pressed on technical details. The worst case involved the remaining three papers, whose authors couldn’t answer even basic questions—and all three were single-author papers.

Shah wrote that two of the authors appeared to have little substantive understanding of the content of their own paper, and the third could not locate where in the main text the key results claimed in the abstract were presented or supported.

Even the only paper that answered all questions correctly was not accepted. During review, Shah found a major error in one of the paper’s key conclusions, which the authors subsequently acknowledged. TMLR issued a direct rejection for this paper but allowed the authors to resubmit after correcting the error or narrowing the conclusion. The remaining nine papers were directly rejected without the option to resubmit.

Two side events after the meeting

Shah also recorded two other events in the text, which read with a darkly humorous tone.

arXiv

First, the authors of two papers were unable to answer basic questions during the conference, but later sent written responses. Shah submitted these two emails to the AI text detection tool Pangram, and both were flagged as "100% AI-generated."

arXiv

Second: In another conversation, an author attempted to introduce Shah to their analytical approach, but in the process, unintentionally described a complete p-hacking procedure—repeatedly adjusting analytical methods until the data became “significant.”

It should be noted that Shah himself does not reject AI. In the article, he openly admits that, to read eight papers within two weeks, he also used LLMs to aid his understanding and even learned some concepts he was previously unfamiliar with. What concerns him is not whether the author used AI, but whether the author takes responsibility for content published under their own name.

Why do this?

TMLR was founded in 2022 and is operated by the team behind JMLR, known for its review standard that evaluates submissions solely on whether conclusions are supported by evidence, without rejecting papers based on novelty or SOTA status. Its reviewers, action editors, and editors-in-chief are all unpaid volunteers, making it especially strained during this year’s surge in submissions.

According to a June announcement by TMLR, its submission volume has tripled over the past year, with single-author submissions surging thirteenfold—editors have even encountered individuals submitting five papers in a single day. Another set of figures presented by Shah in his latest article provides an even clearer picture: in 2023, TMLR’s direct rejection rate was approximately 6%; today, it has risen to about 53%, meaning more than half of submissions never reach external review.

arXiv

In response to this situation, TMLR rolled out a series of measures intensively this summer.

One approach is to set annual submission quotas per author. Unlike many conferences that impose a fixed limit of “N papers per person,” TMLR employs a “harmonized” quota rule in which the quota consumed by a paper decreases as the number of co-authors increases. Specifically: authors submitting only single-author papers may submit up to 2 papers per year; if all submissions are co-authored by nine authors, the maximum becomes 9 papers per year; active reviewers and action editors have their quotas doubled. This rule takes effect on July 1. TMLR explains in its announcement that the quota is not divided equally among authors to prevent individuals from gaining additional submission slots by adding nominal co-authors.

Second, AI peer review has been introduced. In July, TMLR announced that each submission will be accompanied by an AI-generated review in addition to conventional human reviews. The AI review assesses only the paper’s “reliability”—whether the conclusions are supported by accurate, clear, and convincing evidence—without making subjective judgments or recommending acceptance or rejection; the final decision remains with the action editor. After evaluation, TMLR selected CSPaper’s AI review tool.

Third, include "written clearly" as part of the acceptance criteria. On August 28, TMLR revised its previous "audience interest" standard to explicitly require papers to clearly communicate their findings to readers. The announcement directly stated that current AI-generated writing often tends to be convoluted, overloaded with jargon, and difficult to understand; if a paper is almost entirely generated by AI with minimal human involvement, it is unlikely to meet this standard—at least with the current capabilities of AI systems.

Shah concluded at the end of the article that the interview strengthened the editorial team’s confidence in their existing immediate rejection process, and that the quota system and emphasis on clear writing were also proven to be helpful.

Not just TMLR

TMLR's experience is not an isolated case. This year, several major publication channels in the AI field have been grappling with the same issue.

In early June, the NeurIPS official blog revealed that of the 971 submissions to the NeurIPS 2026 Position Paper track, 273 (28.2%) were flagged by Pangram as having an AI score of 100%.

arXiv

https://blog.neurips.cc/2026/06/02/ai-generated-papers-in-the-neurips-2026-position-paper-track/

This year, the track explicitly requires that papers be "substantially written by humans," with AI permitted only for peripheral edits such as language polishing. NeurIPS also noted that the growth in AI-generated writing is widespread: in the "Evaluation and Datasets" track, the number of papers with a Pangram score of 90% or higher increased more than tenfold from 2025 to 2026. However, NeurIPS acknowledged that detection results are sensitive to parameters; switching to a medium-sized detection window reduces the proportion of papers with AI scores between 90% and 100% from 42.7% to 12.7%.

Earlier in mid-May, Thomas G. Dietterich, head of the arXiv computer science section, announced a clear penalty on social media: if there is definitive evidence that authors failed to review the outputs of large models—such as hallucinated references or chatbot dialogue remnants left in the text—all authors will be banned from submitting to arXiv for one year, and future submissions must first be accepted through formal peer review before being posted. Dietterich emphasized that authorship on a paper means each author is responsible for all content, regardless of how it was generated.

arXiv

The related data is equally striking. According to a Lancet study by Columbia University, the proportion of biomedical papers in PubMed Central containing at least one fabricated citation rose from approximately 4 in 10,000 in 2023 to about 57 in 10,000 by early 2026.

arXiv

https://www.nursing.columbia.edu/news/nearly-3-000-peer-reviewed-medical-papers-have-fake-citations-columbia-nursing-ai-assisted-audit-finds

From checking text to checking people

Shah admitted in the article that the trial took him 20 to 25 hours to process just eight papers, making this approach difficult to scale given today’s submission volumes.

He also mentioned that his team is researching scalable solutions.

Just a few days ago, Shah, along with Justin Payan, Bálint Gyevnár, and Atoosa Kasirzadeh from CMU, published a paper introducing a method called greCAPTCHA—a name derived partly from the GRE graduate entrance exam and partly from CAPTCHA, the verification system that distinguishes humans from machines.

arXiv

https://www.cs.cmu.edu/~nihars/preprints/greCAPTCHA.pdf

The approach is simple: instead of detecting whether text was written by AI, it directly evaluates the author. The author submits a paper along with a statement of their personal contribution, and the system automatically generates questions based on this. The author then answers these questions under supervised exam conditions, after which an evaluation report is generated for reference by journals, employers, or admissions committees. Researchers refer to the targeted skill as the “capacity to verify”—that is, whether the author possesses the knowledge and reasoning ability to critically assess their own contributions.

arXiv

The questions are divided into four categories: identifying data from the paper's true report among multiple versions (pre-set error identification), explaining design choices not justified in the paper, explaining background concepts assumed known to the reader by default, and identifying the specific conditions under which the method fails.

arXiv

The team recruited 31 researchers to take the test, each answering questions about their own paper and a foreign paper. The results showed that the system achieved an AUC of 0.90 in distinguishing between "their own paper" and "a foreign paper," rising to 0.934 after removing multiple-choice questions. The multiple-choice questions themselves had almost no discriminatory power, with an AUC of only 0.595—simply because the answers could be directly found by searching the PDF. Another notable finding was that participants showed no statistically significant difference in scores when answering foreign papers within their field versus those outside their field, indicating that the test measures more than just domain-specific knowledge.

One participant’s experience is particularly telling: before the test, he handed an unfamiliar paper to AI and had it explain the entire document from start to finish, yet he still couldn’t answer in-depth questions—the unfamiliar paper scored 1.3, while his own paper scored 51.

However, this system is far from mature. The lowest false positive rate reported in the paper was approximately 20%, meaning that one in every five genuine authors might be incorrectly flagged. Participants voiced the most complaints about the scoring system: it was unclear how detailed their answers should be, and the grading criteria appeared overly rigid. One author, confronted with a scoring standard that rejected his own answer, exclaimed, “That was clearly my own design.” Some participants noted that in collaborative papers, certain sections were written by co-authors, making it impossible for them to answer questions about those parts. Others expressed concern that such exams would inadvertently assess English proficiency and typing speed, unfairly disadvantaging non-native English speakers and researchers with disabilities.

Conclusion

Shah listed several conclusions in the article, two of which are worth noting:

First is credit allocation. In today’s research ecosystem, academic contributions are primarily recognized through authorship on papers. If authors cannot explain or defend their papers, the signal value of publication as an indicator of “researcher contribution” is significantly diminished.

Second, the peer review cycle. Many journals and conferences invite authors to review others’ papers. If authors cannot even understand their own papers, their ability to serve as reviewers is naturally questionable. In today’s environment, where AI-generated reviews and AI-written papers are flooding in, if this cycle breaks, the foundation of peer review will be severely undermined.

Of course, this experiment itself had clear limitations. A sample size of 10 is very small, was non-formally selected by Shah, and consisted entirely of papers that were already slated for immediate rejection; thus, it can only reveal what kinds of submissions are deemed low-quality, not the overall profile of TMLR submissions. Shah himself acknowledged this: one of the experiment’s goals was precisely to test whether the immediate rejection process is reliable.

But Shah’s experiment at least established a baseline: a paper can be assisted by AI, but the person listed as author must be able to explain what it contains.

This article is from the WeChat public account "Machine Heart" (ID: almosthuman2014), authored by Machine Heart.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.