金色财经|7月 28, 2026 14:38
[Kimi Open-Sources PerceptionBench Visual Perception Benchmark, GPT-5.6-Sol Ranks First but No Model Exceeds 60% Accuracy]
According to a report by Jinse Finance, the Kimi Team has announced the open-sourcing of the multimodal large model visual perception evaluation benchmark, PerceptionBench. The benchmark aims to decompose visual perception capabilities into 10 atomic-level abilities for independent evaluation, covering dimensions such as visual relationships, counting, attributes, depth and 3D, localization, comparison, fine-grained recognition, context integration, OCR, and hallucination recognition.
The benchmark is constructed based on failure cases from 42 existing evaluation datasets and includes a total of 3,000 manually verified questions, each designed to test a single visual capability without requiring reasoning or external knowledge.
Evaluation results show that among 16 cutting-edge multimodal large models, none achieved an overall accuracy rate exceeding 60%. GPT-5.6-Sol ranked first with an accuracy rate of 59.7%, followed by Kimi K3 (58.5%), Claude-Fable-5 (57.2%), Gemini-3.1-Pro (56.2%), and GPT-5.5 (55.8%) in the top five. The report highlights that visual hallucination remains the weakest capability across all models, indicating significant room for improvement in overall perception capabilities.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink