Stand: Juni 2026

12 KI-Modelle im wissenschaftlichen Vergleich

Unabhaengiger Benchmark von Claude Opus 4.8, GPT-5.5, Gemini 3.1 Pro, DeepSeek R1 und 8 weiteren Modellen. Basierend auf Arena Elo (6M+ Votes), SWE-bench Pro, GPQA Diamond und HLE.

12
Modelle
6M+
Arena Votes
8
Benchmarks
4
Kategorien
LMSYS Arena verifiziert
Woechentlich aktualisiert
Unabhaengig
SWE-bench Pro

Frontier-Modelle Juni 2026

Sortiert nach KapitalAI Composite Score. Arena Elo Gaps unter 30 Punkten sind statistisch gleichwertig.

Frontier (Elo >1500)

GPT-5.5

OpenAI
92.8
Score
Arena Elo: 1561

100% auf AIME 2026 — perfekt. Fuehrt ARC-AGI-2 (25.8%), beste Latenz (1.8s TTFT). Staerkstes Tool-Use und Third-Party-Integrationen.

64.1%
SWE-Pro
93.2%
GPQA
95.8%
MATH
52.2%
HLE
256K
Context
$5.00/$30.00 per 1M
1.8s TTFT 82 TPS
MathReasoningTool Use Teuerster OutputSchwaecher HLE

Gemini 3.1 Pro

Google DeepMind
91.5
Score
Arena Elo: 1520

#1 GPQA Diamond (94.3%). 2M Token Context — laengster am Markt. Bestes Preis-Leistungs-Verhaeltnis unter Frontier-Modellen.

59.2%
SWE-Pro
94.3%
GPQA
93.7%
MATH
51.4%
HLE
2M
Context
$2.00/$12.00 per 1M
2.4s TTFT 71 TPS
ScienceLong ContextMultimodal Schwaecher CodingNur Google Cloud
Near-Frontier (Elo 1450-1500)

Grok 4.2

xAI
88.4
Score
Arena Elo: 1485

Einziger Zugang zu Echtzeit X/Twitter-Daten. 2M Token Context, 89 TPS — schnellstes Frontier-Modell.

52.8%
SWE-Pro
87.4%
GPQA
88.2%
MATH
54.2%
HLE
2M
Context
$2.00/$10.00 per 1M
1.9s TTFT 89 TPS
Realtime DataLong ContextSpeed Nur via X Premium+

Claude Sonnet 4.6

Anthropic
86.8
Score
Arena Elo: 1478

Default fuer Claude Code. 79.6% SWE-bench Verified, 112 TPS. Beste Balance zwischen Qualitaet, Geschwindigkeit und Preis.

54.8%
SWE-Pro
84.2%
GPQA
84.6%
MATH
44.8%
HLE
200K
Context
$3.00/$15.00 per 1M
1.2s TTFT 112 TPS
Daily CodingFast IterationCost-Effective Nicht Opus-LevelKuerzer Context

DeepSeek R1

DeepSeek
87.6
Score
Arena Elo: 1478

#1 MATH-500 (97.3%). Open Source (MIT), 10x guenstiger als Frontier. Training kostete nur $6M — effizientestes Modell.

49.2%
SWE-Pro
82.1%
GPQA
97.3%
MATH
48.5%
HLE
128K
Context
$0.55/$2.19 per 1M
3.2s TTFT 48 TPS
MathReasoningBudget LangsamerChain-of-Thought verbraucht Tokens
High Performance (Elo 1420-1450)

Qwen 3 235B

Alibaba Cloud
86.2
Score
Arena Elo: 1422

Staerkstes Open-Weight Modell. 119 Sprachen, MCP Tool-Calling. Apache 2.0 — volle kommerzielle Nutzung.

48.8%
SWE-Pro
77.2%
GPQA
91.2%
MATH
42.1%
HLE
128K
Context
$0.80/$2.40 per 1M
2.8s TTFT 52 TPS
MultilingualSelf-HostAgents Schwaecher Deutsch

Llama 4 Maverick

Meta
84.8
Score
Arena Elo: 1417

400B MoE mit nur 17B aktiven Parametern pro Token. 1M Context, nativ multimodal. Llama License (700M MAU Limit).

46.3%
SWE-Pro
74.8%
GPQA
85.5%
MATH
38.2%
HLE
1M
Context
$0.87/$2.60 per 1M
2.6s TTFT 58 TPS
Self-Host1M ContextMultimodal 700M MAU Limit
Cost-Effective (Elo <1420)

Gemini 3.5 Flash

Google DeepMind
82.4
Score
Arena Elo: 1405

Schnellstes Modell: 0.8s TTFT, 185 TPS. Frontier-nah bei 1/30 des Preises. Ideal fuer Batch-Processing.

42.8%
SWE-Pro
78.5%
GPQA
82.1%
MATH
34.2%
HLE
1M
Context
$0.15/$0.60 per 1M
0.8s TTFT 185 TPS
High VolumeLow LatencyBudget Nicht Frontier-Level

Mistral Large 3

Mistral AI
81.4
Score
Arena Elo: 1398

Europaeisch, DSGVO-konform. 93.6% MATH-500. Open Weights verfuegbar. Guenstigstes Frontier-nahes Modell.

42.8%
SWE-Pro
73.1%
GPQA
93.6%
MATH
32.4%
HLE
260K
Context
$0.50/$1.50 per 1M
1.4s TTFT 98 TPS
EU/DSGVOSelf-HostBudget Schwaecher Reasoning

o3-mini (high)

OpenAI
79.8
Score
Arena Elo: 1385

#2 ARC-AGI (28.5%). Spezialisiert auf mathematisches Reasoning. 96.4% MATH-500 bei niedrigen Kosten.

38.2%
SWE-Pro
70.2%
GPQA
96.4%
MATH
31.5%
HLE
200K
Context
$1.10/$4.40 per 1M
4.2s TTFT 32 TPS
MathReasoningSTEM LangsamNicht fuer Coding

Claude Haiku 4.5

Anthropic
76.5
Score
Arena Elo: 1358

Schnellstes Anthropic-Modell. Ideal fuer Classification, Extraction, einfache Tasks bei hohem Volume.

32.1%
SWE-Pro
65.8%
GPQA
72.4%
MATH
24.5%
HLE
200K
Context
$1.00/$5.00 per 1M
0.6s TTFT 165 TPS
ClassificationExtractionHigh Volume Nicht fuer komplexe Tasks

Vollstaendiges Benchmark-Ranking

Alle 12 Modelle mit 8 Benchmarks. Beste Werte pro Kategorie hervorgehoben.

#ModellScoreElo SWE-ProGPQAMATHHLEARC-AGIAIME ContextTPS$/1M
1 Claude Opus 4.8
Anthropic
94.2 1569 69.2% 93.6% 92.4% 57.9% 21.4% 72.5% 200K 67 $5.00/$25.00
2 GPT-5.5
OpenAI
92.8 1561 64.1% 93.2% 95.8% 52.2% 25.8% 100% 256K 82 $5.00/$30.00
3 Gemini 3.1 Pro
Google DeepMind
91.5 1520 59.2% 94.3% 93.7% 51.4% 24.2% 89.7% 2M 71 $2.00/$12.00
4 Grok 4.2
xAI
88.4 1485 52.8% 87.4% 88.2% 54.2% 18.9% 78.5% 2M 89 $2.00/$10.00
5 Claude Sonnet 4.6
Anthropic
86.8 1478 54.8% 84.2% 84.6% 44.8% 16.4% 62.5% 200K 112 $3.00/$15.00
6 DeepSeek R1
DeepSeek
87.6 1478 49.2% 82.1% 97.3% 48.5% 19.2% 92.8% 128K 48 $0.55/$2.19
7 Qwen 3 235B
Alibaba Cloud
86.2 1422 48.8% 77.2% 91.2% 42.1% 15.8% 85.7% 128K 52 $0.80/$2.40
8 Llama 4 Maverick
Meta
84.8 1417 46.3% 74.8% 85.5% 38.2% 14.2% 72.1% 1M 58 $0.87/$2.60
9 Gemini 3.5 Flash
Google DeepMind
82.4 1405 42.8% 78.5% 82.1% 34.2% 12.8% 58.5% 1M 185 $0.15/$0.60
10 Mistral Large 3
Mistral AI
81.4 1398 42.8% 73.1% 93.6% 32.4% 11.2% 44% 260K 98 $0.50/$1.50
11 o3-mini (high)
OpenAI
79.8 1385 38.2% 70.2% 96.4% 31.5% 28.5% 86.5% 200K 32 $1.10/$4.40
12 Claude Haiku 4.5
Anthropic
76.5 1358 32.1% 65.8% 72.4% 24.5% 8.5% 42.8% 200K 165 $1.00/$5.00

Quellen: LMSYS Arena, Artificial Analysis, SWE-bench, BenchLM

Benchmark-Kategorien

Industriestandard-Tests fuer KI-Modelle

SWE-bench Pro

Real-World Coding — 1,865 Tasks, 41 Repos

GPQA Diamond

PhD-Level Science Reasoning

MATH-500

Olympiade-Level Mathematik

HLE

Humanity's Last Exam — Hardest Reasoning

Arena Elo

6M+ Human Votes (LMSYS)

ARC-AGI-2

Abstract Reasoning Corpus

Detaillierter Benchmark-Report

92 Seiten mit allen Testergebnissen, Methodologie und Use-Case Empfehlungen

Kein Spam - Jederzeit abbestellbar