Odaily Planet Daily reports that Kimi has open-sourced PerceptionBench, a multimodal large model visual perception evaluation benchmark designed to decompose visual perception into 10 atomic capabilities for independent assessment, covering dimensions such as visual relationships, counting, attributes, depth and 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination detection. The benchmark is built on model failure cases from 42 existing evaluation datasets and includes 3,000 manually verified questions, each assessing only a single visual capability without requiring reasoning or external knowledge.
The evaluation results show that none of the 16 leading multimodal large models achieved an overall accuracy rate above 60%. GPT-5.6-Sol ranked first with an accuracy rate of 59.7%, followed by Kimi K3 (58.5%), Claude-Fable-5 (57.2%), Gemini-3.1-Pro (56.2%), and GPT-5.5 (55.8%). The report notes that visual hallucination remains the weakest capability across all models, indicating significant room for improvement in overall perceptual performance.
