Living TierMaker-style ranking of frontier and open-weight AI models.
Personal architectural evaluations calibrated against real-world developer leverage, reasoning precision, agentic tool workflows, and operational cost. Designed with a structured schema for dynamic online benchmark integration.
Architectural Ranking Preset (Models Sorted Left → Right within Tiers)
Active: balanced Mode
SFrontier & Autonomous Champions(4 models)Ranked Left → Right
#1Anthropic
↑ +5 Elo91
Claude Opus 5
Frontier High-Capability Core
1M tokensView Reason ↗
#2Google
↑ +6 Elo91
Gemini 3.1 Pro
Multimodal Massive Context
2M+ tokensView Reason ↗
#3OpenAI
↑ +6 Elo90
GPT-5.6 Sol
Frontier Deep Reasoning
1M tokensView Reason ↗
#4Anthropic
↑ +7 Elo89
Claude Fable 5
Mythos-Class Autonomous Agent
1M tokensView Reason ↗
AElite Specialists & Fast Frontier(5 models)Ranked Left → Right
#1Google
↑ +5 Elo94
Gemini 3.7 Flash
Ultra-Fast Frontier Agentic
2M tokensView Reason ↗
#2OpenAI
↑ +4 Elo93
GPT-5.6 Terra
Balanced Frontier Workhorse
1M tokensView Reason ↗
#3DeepSeek
↑ +6 Elo93
DeepSeek V4-Pro
Open/Frontier Math & Architecture
1M tokensView Reason ↗
#4xAI
↑ +7 Elo92
Grok 4.6
Frontier Real-Time Reasoning
1M tokensView Reason ↗
#5Anthropic
↑ +7 Elo90
Claude 3.7 Sonnet
Hybrid Extended Reasoning
200k tokensView Reason ↗
BSolid Workhorses & Open Weights(4 models)Ranked Left → Right
#1OpenAI
↑ +4 Elo91
GPT-5.6 Luna
Efficient Fast Agentic
500k tokensView Reason ↗
#2DeepSeek
↑ +5 Elo91
DeepSeek V4-Flash
Ultra-Low-Cost Inference
512k tokensView Reason ↗
#3Meta
↑ +6 Elo89
Muse Glimmer (Llama 4)
Open-Weight Frontier
512k tokensView Reason ↗
#4Alibaba Cloud
↑ +4 Elo88
Qwen 3.8-Max
Multilingual & Code Specialist
1M tokensView Reason ↗
CSituational & Demoted Legacy(3 models)Ranked Left → Right
#1Google
↓ -3 Elo83
Gemini 2.0 Flash
Legacy Fast Multimodal
1M tokensView Reason ↗
#2Anthropic
↓ -5 Elo79
Claude 3.5 Sonnet
Legacy Coding Standard
200k tokensView Reason ↗
#3OpenAI
↓ -5 Elo74
GPT-4o
Legacy Omnimodal Foundation
128k tokensView Reason ↗
DSuperseded & Historic Baselines(1 models)Ranked Left → Right
#1OpenAI
10
GPT-3.5 Turbo
Historic Generative Baseline
16k tokensView Reason ↗
Evaluated Benchmark Reports & Diagrams
Public Telemetry Sources & Methodology References
The public benchmarks, blind human arenas, and independent lab telemetry taken into account to determine ratings and diagram dimensions.
Crowdsourced Blind ELO EvaluationLarge Model Systems Organization (LMSYS)
LMSYS Chatbot Arena Leaderboard
Gold-standard crowdsourced double-blind human preference evaluation covering hundreds of thousands of pairwise model battles under rigorous Bradley-Terry statistical modeling.
End-to-End Software EngineeringPrinceton NLP & SWE-bench Team
SWE-bench Verified Benchmark
Real-world software engineering benchmark assessing model ability to resolve end-to-end GitHub issues across complex open-source Python repositories with rigorous unit test validation.
Contamination-Free Dynamic EvaluationAbacus AI & Collaborative Research Labs
LiveBench AI Benchmark
Continuously updated, contamination-free benchmark evaluating AI models on newly published coding challenges, mathematical theorems, and data analysis tasks released after model training cutoffs.
Live ReasoningLive CodingMathematical Problem Solving
Pragmatic coding benchmark measuring model capability to edit real codebases, adhere to diff formats, and resolve test-driven refactoring tasks without breaking existing test suites.
Polyglot Benchmark ResolutionWhole-file vs Diff Patching AccuracyInstruction Following
Real-time market analytics tracking developer adoption, token consumption volume, and cost trends across hundreds of foundation and open-weights AI models.