The largest academic AI conference just let artificial intelligence grade its own homework. AAAI-26, one of the premier venues for AI research, ran a pilot program that generated AI reviews for 22,977 main-track paper submissions, making it the first full-scale deployment of AI-powered peer review at a major academic conference.
The twist: surveys conducted after the pilot found that both authors and program committee members actually preferred the AI-generated reviews. Specifically, they rated them higher on technical accuracy and quality of research suggestions compared to their human counterparts.
How the double-blind AI review system worked
The program operated within a double-blind framework, meaning neither authors nor reviewers knew each other’s identities. AI-generated reviews were slotted in alongside at least two human reviews for each paper, creating a side-by-side comparison that researchers could evaluate without bias toward either source.
One important design choice: the AI reviews identified themselves as AI-generated. This wasn’t a Turing test. The goal was to complement human reviewers, not to trick anyone into thinking a language model was a tenured professor.
The AI reviews were produced using a multi-stage LLM-based system that could process papers in under one day. Participants noted the supplementary nature of the AI reviews and found them to enhance the overall review process rather than dilute it.
Why AI peer review matters beyond academia
A 2023 study from Stanford found that AI-generated text already appeared in a measurable fraction of peer reviews at top machine learning conferences, though in that case it was reviewers quietly using ChatGPT to draft their feedback rather than an officially sanctioned system.
AAAI-26’s pilot took the opposite approach. Instead of pretending AI isn’t already in the room, it formalized the process, built guardrails around it, and studied the results empirically.
The research findings are documented in arXiv paper 2604.13940, with a public release of the full dataset and analysis slated for August 2026.
The 22,977 papers processed in this pilot represent a scale that no previous study of AI-assisted peer review has approached. Prior work operated in simulated environments or with small sample sizes that made it difficult to draw generalizable conclusions.
