Best AI models for base-rate bias

Ranked by accuracy on the meo “base-rate bias” domain — one of eleven un-leaked, contamination-resistant domains. Per-domain strength reveals where a model genuinely reasons versus where it pattern-matches; the leaders differ sharply by domain.

Ranking 22 models across the meo-tested roster · as of 2026-06-08.

#	Model	Lab	accuracy
1	Anthropic: Claude Opus 4.8	anthropic	10%
2	Owl Alpha	openrouter	10%
3	Z.ai: GLM 5.1	z-ai	10%
4	Google: Gemini 3.1 Flash Lite	google	10%
5	OpenAI: GPT-5.5	openai	7%
6	inclusionAI: Ring-2.6-1T	inclusionai	7%
7	Xiaomi: MiMo-V2.5	xiaomi	7%
8	MiniMax: MiniMax M3	minimax	7%
9	Perceptron: Perceptron Mk1	perceptron	7%
10	Tencent: Hy3 preview	tencent	7%
11	Arcee AI: Trinity Large Thinking	arcee-ai	7%
12	Qwen: Qwen3.7 Max	qwen	5%
13	DeepSeek: DeepSeek V4 Flash	deepseek	5%
14	DeepSeek: DeepSeek V4 Pro	deepseek	5%
15	Qwen: Qwen3.7 Plus	qwen	5%
16	Xiaomi: MiMo-V2.5-Pro	xiaomi	5%
17	Google: Gemini 3.1 Pro Preview	google	2%
18	Google: Gemini 3.5 Flash	google	2%
19	MoonshotAI: Kimi K2.6	moonshotai	2%
20	NVIDIA: Nemotron 3 Ultra	nvidia	2%
21	StepFun: Step 3.7 Flash	stepfun	2%
22	xAI: Grok 4.3	x-ai	0%

← All rankings Methodology & 𝕍 →