The non latin language penalty when using LLMs / agentic workflows is sooo big, that's crazy. I did the experiment with Hebrew this time. A single 6h14 QLoRA run nearly doubled Hebrew tool-call exact-match accuracy, from 43.4% to 83.2%, while English accuracy also improved from 70.5% to 88.4%. Details of the end to end run: - Full pipeline: 7h14 on one NVIDIA L40S GPU - Training: 3,662 steps, 2 epochs, 33.96M tokens - Peak GPU memory: 26,608 MiB - Hebrew valid JSON: 81.0% → 99.6% - Hebrew function-name accuracy: 74.7% → 99.2% - Hebrew–English gap: −27.6 → −5.2 points - Translation yield: 16,272/17,000 rows (95.7%) - Cost: $14.11
abdelShare

Source:Show original
Disclaimer: The information on this page may have been obtained from third parties and does not necessarily reflect the views or opinions of KuCoin. This content is provided for general informational purposes only, without any representation or warranty of any kind, nor shall it be construed as financial or investment advice. KuCoin shall not be liable for any errors or omissions, or for any outcomes resulting from the use of this information.
Investments in digital assets can be risky. Please carefully evaluate the risks of a product and your risk tolerance based on your own financial circumstances. For more information, please refer to our Terms of Use and Risk Disclosure.