Intelligence
Reasoning, knowledge, coding, factuality, instruction following, and reliability.
- →Primary benchmark leaderboards and public repositories
- →Reproducible papers and published run logs
Built on public-evidence data, standardized methodology, and up-to-date research using the latest AI models. A multidimensional view across Intelligence, Multimodality, Agent, Context, and Price.
Reasoning, knowledge, coding, factuality, instruction following, and reliability.
Useful text, vision, image, audio, video, and document capabilities—not merely nominal input support.
Tool use, planning, agentic coding, research, computer use, long-horizon completion, and recovery from errors.
Advertised context capacity plus demonstrated retrieval, comprehension, and long-context reasoning.
Public standard input/output pricing and cost per completed task where comparable data exist.
Overall score = 0.50 × Intelligence + 0.15 × Multimodality + 0.15 × Agent + 0.10 × Context + 0.10 × Price. The scorecard’s I, M, A, C and P columns are RS-normalized research-synthesis scores, not raw percentages from one benchmark.
Primary evidence is original benchmark leaderboards, public repositories, reproducible papers and published run logs. Public model documentation and standard API specifications support coverage; vendor claims and aggregators are cross-checks only. Coverage may be incomplete and non-identical across frontier models, so the leading cluster should be read with roughly ±2–4 points of uncertainty.