Rabdos Vizi Bench
27 problems · Updated on August 6, 2026
For reasoning over scientific and mathematical visualizations.

RankModel
RubricStrict
| Rank | Provider | Model | Reasoning effort | Rubric | Strict |
|---|---|---|---|---|---|
| 1 | Anthropic | Claude Opus 5 | max | 32.7% | 7/27 (25.9%) |
| 2 | Anthropic | Claude Fable 5 | max | 24.6% | 6/27 (22.2%) |
| 3 | OpenAI | GPT-5.6 Sol | max | 12.4% | 2/27 (7.4%) |
| 4 | Gemini 3.6 Flash | high | 9.4% | 2/27 (7.4%) | |
| 5 | Alibaba | Qwen 3.8 Max | max | 5.6% | 0/27 (0.0%) |
| 6 | xAI | Grok 4.5 | xhigh | 4.2% | 0/27 (0.0%) |
| 7 | Moonshot AI | Kimi K3 | max | 3.7% | 0/27 (0.0%) |
| 8 | Meta | Muse Spark 1.1 | xhigh | 1.9% | 0/27 (0.0%) |
Interested in evaluating your model on Rabdos Vizi Bench? Get in touch →