AIRA₂ AI Agent Outperforms Previous Models by 9 Percentage Points in ML Benchmark

iconCryptoBriefing
Share
AI summary iconSummary
AI + crypto news: AIRA₂, an autonomous AI research agent from Meta’s FAIR lab, University College London, and the University of Oxford, scored 81.5% on the MLE-bench-30 benchmark after 24 hours, rising to 83.1% after 72 hours. This marks a 9 percentage point jump over prior models. AIRA₂ also ranked 8th in a Kaggle competition with 4,000 teams, securing a gold medal. On-chain news activity highlights growing AI integration in data-driven fields.

An AI agent just beat roughly 3,992 human teams at their own game. AIRA₂, an autonomous AI research agent developed by researchers at Meta’s FAIR lab, University College London, and the University of Oxford, placed 8th out of 4,000 teams in a Kaggle competition focused on AI reasoning, earning a gold medal in the process.

What AIRA actually does

The AIRA line of agents (the name stands for AI Research Agent) is designed to autonomously tackle complex machine learning engineering problems. Rather than just running a single model and hoping for the best, AIRA uses an asynchronous multi-GPU execution strategy running on 8 Nvidia H200 GPUs.

Advertisement

The system employs what the research team calls a “Hidden Consistent Evaluation” protocol, a method to prevent the agent from gaming its own test scores. The agent also uses dynamic ReAct-style operators, meaning it can reason through problems step by step, adjust its approach on the fly, and execute code across multiple processors simultaneously.

On MLE-bench-30, a structured benchmark built from 30 real Kaggle competitions, AIRA₂ achieved a mean percentile rank of 81.5% after 24 hours of computation. Give it 72 hours, and that number climbed to 83.1%. The best prior baseline from other agents sat at 72.7%, making AIRA₂’s improvement roughly 9 percentage points above the competition.

Why this matters beyond bragging rights

AIRA₂ exceeded human state-of-the-art performance on 6 out of 20 tasks in the AIRS-Bench evaluation. It also won gold medals on individual Kaggle tasks where earlier versions of the agent failed to secure any awards at all.

The framework builds on what the team calls the AIRA-dojo approach, designed specifically to address limitations of earlier agents. Those limitations included limited computational throughput, evaluation overfitting, generalization gaps across different problem types, and static operational constraints that prevented the agent from adapting its strategy mid-task.

The competitive AI research landscape

The collaboration between Meta FAIR, UCL, and Oxford is notable for its institutional weight. The research team disseminated its findings through arXiv preprints, the standard channel for cutting-edge AI research.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.