Ramp SWE-Bench Benchmark: Claude Fable 5 Leads with 87.5% Success Rate

iconKuCoinFlash
Share
AI summary iconSummary
Ramp, a fintech unicorn, has released its SWE-Bench benchmark for AI coding agents, featuring 80 backend tasks from real production environments. Anthropic’s Claude Fable 5 scored 87.5%, the highest among 14 models. AI + crypto news highlights Claude Opus 4.8’s cost efficiency at $1.09 per run. Domestic models Kimi K2.6 and GLM 5.1 followed with 72.5% and 71.25%. GPT-5.4 Mini led lightweight models with 58.75%. Ramp noted that trade-offs among performance, speed, and cost are key for deployment. Interest rate news remains a separate focus for traders.
ME AI News: According to monitoring by Beating, fintech unicorn Ramp has released Ramp SWE-Bench, a private benchmark for advanced AI coding agents. Ramp SWE-Bench comprises 80 backend development tasks derived from Ramp’s real production environment, designed to address data leakage and metric saturation issues caused by model pre-training in public evaluation datasets. All test tasks are extracted from real pull requests (PRs) successfully deployed to production via Ramp’s internal AI coding assistant, Inspect. Each task reconstructs the base codebase prior to the PR submission, preserves the merged code and tests as the gold standard, and extracts the engineer’s original intent from Inspect conversations as prompts. Evaluation is conducted in a mini-swe-agent sandbox environment, with pass@1—successfully passing all tests in a single run without breaking existing functionality—as the passing criterion. According to the published benchmark results across 14 models, Anthropic’s newly launched Claude Fable 5 leads with an 87.5% success rate. Claude Opus 4.7 and GPT-5.5 tie for second place, both achieving an 83.75% success rate. Although Claude Opus 4.8 has a lower success rate of 77.5%, its average runtime per task is reduced to 8 minutes and 30 seconds, with a single-run cost of just $1.09—under 40% of Fable 5’s cost—demonstrating exceptional cost efficiency. The test data also reveals trade-offs between price and performance across models. Chinese models Kimi K2.6 and GLM 5.1 show similar success rates of 72.5% and 71.25%, respectively, but Kimi K2.6’s average cost is $0.69—approximately 34% cheaper than GLM 5.1. Among lightweight models, GPT-5.4 Mini leads with a 58.75% success rate, outperforming Claude Haiku 4.5 by about 10 percentage points, while also halving both average cost and steps required. Ramp notes that as models move toward practical deployment, the balance between performance, speed, and cost is becoming a critical consideration for engineering implementation. (Source: BlockBeats)
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information. Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.