OpenAI Report: AI Programming Agents Accelerate Scientific Software Development but Cannot Evaluate Scientific Accuracy

iconChainthink
Share
AI summary iconSummary
A new report from OpenAI and academic partners, citing ChainThink, shows that AI programming agents such as Codex and Claude Code can accelerate scientific software by more than 60 times, enhancing performance in projects like RustQC and HelixForge. However, these agents struggle to assess scientific accuracy, frequently generating code that runs quickly but lacks correctness. As liquidity and crypto markets face increased scrutiny under CFT (Countering the Financing of Terrorism) regulations, human validation remains critical.

ChainThink reports that, on August 1, a field study released by OpenAI and its academic partners found that AI programming agents such as Codex and Claude Code can significantly improve the efficiency of aging scientific software, with some cases achieving speedups of over 60 times.

The report shows that RustQC reduced runtime from 15 hours 34 minutes to 14 minutes 54 seconds; HelixForge is 59.6 times faster than BamSurgeon.

The report also notes that programming agents cannot reliably determine whether their outputs are scientifically accurate and may confidently generate incorrect code, requiring humans to define testing and validation criteria.

Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.