OceanBase Data Agent Outperforms GPT and Claude in International Benchmark

iconMetaEra
Share
AI summary iconSummary
OceanBase's Data Agent solution, built on the GLM-5.2 large model, achieved a 90.62% accuracy rate on the Data Agent Benchmark, outperforming GPT and Claude Opus. The evaluation covered finance, biomedicine, and government sectors, assessing data understanding and result retrieval. OceanBase plans to integrate this technology into DataPilot to enhance its capabilities in on-chain and inflation data analysis. This solution represents a significant step toward building an AI-powered data platform.
The OceanBase team's Data Agent solution, built on the domestic large model GLM-5.2, has topped the international Data Agent Benchmark (DAB) with an accuracy rate of 90.62%, becoming the first entry on the leaderboard to surpass 90%. This solution outperforms multiple other approaches based on overseas models such as GPT and Claude Opus, demonstrating the capability of domestic databases and AI models to jointly support complex data analysis tasks. The DAB evaluation spans multiple domains including finance, biomedical research, and government services, testing the full capability from “understanding data” to “obtaining answers.” OceanBase will integrate these technological capabilities into its DataPilot product to drive the evolution from a database to an AI-powered data platform.

Author and source: AIBase

Recently, the Data Agent solution submitted by the OceanBase team topped the international Data Agent Benchmark (DAB) with an accuracy rate of 90.62%, becoming the first entry on the leaderboard to surpass 90% accuracy. Notably, this achievement was accomplished using OceanBase, a domestic database, built on China’s large model GLM-5.2, outperforming multiple Data Agent solutions based on overseas models such as GPT, Claude Opus, and Claude Fable. The internal code name for this submitted solution is Scout, and its capabilities will be integrated into OceanBase DataPilot.

DAB, developed in collaboration between UC Berkeley’s EPIC Data Lab and Hasura PromptQL, evaluates performance across multiple domains including internet/local services, financial stocks, biomedical research, intellectual property, enterprise operations, government/public administration, and media/entertainment, and supports various databases such as PostgreSQL, MongoDB, SQLite, and DuckDB.

Unlike traditional Text-to-SQL, which primarily tests whether AI can convert natural language into SQL, DAB focuses on whether AI can find the correct answer within a real-world data environment, navigating complex, fragmented, and heterogeneous data. In a complete task, a Data Agent must understand the data, select relevant datasets, plan an analytical path, execute queries and calculations, and validate the results—testing its full capability from "understanding the data" to "obtaining the answer."

Therefore, DAB tests not just the underlying model, but the integrated capability of "model + agent + data system." The model handles understanding and reasoning, the agent manages planning and execution, and the data system supports data discovery, relational computation, and result validation. Whether these three components can work together effectively directly determines whether AI can truly leverage enterprise data.

The solution OceanBase submitted for this evaluation is built around this capability. The system uses DataLens to construct data profiles, identifying fields and data relationships; then plans execution paths based on task complexity to perform data selection, filtering, joining, and computation; after obtaining results, it verifies the computation process and outcomes through evidence tracing and answer validation, adjusting the approach and revalidating when issues are detected, forming a closed loop of "data understanding—planning and execution—verification and correction."

This also explains the significance of the 90.62% score: it’s not merely proof that a domestic model can perform data queries, but rather validation that a domestic database combined with a domestic model can jointly support complex data analysis tasks. Underlying the model with GLM-5.2, the OceanBase solution ultimately outperformed multiple Data Agent solutions built using overseas models such as GPT, Claude Opus, and Claude Fable.

This is also a new validation of OceanBase’s AI capabilities.

In the past, databases primarily handled storing, querying, and processing data. As AI agents emerge as new data consumers, the challenges databases must address are evolving: AI doesn’t just need to retrieve data—it must understand it, connect it, analyze it, and verify whether the results are reliable. As a result, the data foundation for the AI era is shifting from “delivering data to AI” to “helping AI put data to effective use.”

This is precisely the direction in which OceanBase is evolving from a database into an AI data platform. OceanBase DataPilot is the product direction focused on AI-driven data analytics scenarios. For this DAB evaluation, we submitted a solution under the internal code name "Scout"; the associated technical capabilities will subsequently be integrated into the DataPilot product to further enable data understanding, task planning, analysis execution, and result validation—enabling AI to move from “accessing data” to “completing tasks with data.”

From databases to AI data platforms, what has changed is not just the product form, but the relationship between databases and AI: previously, databases were responsible for storing, processing, and retrieving data; now, they must also help AI understand, organize, and analyze data. This top ranking in DAB validates OceanBase’s extension of database capabilities into AI-driven data analysis, signaling that databases are evolving from the data foundation of AI to its data operations engine.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.