SEPTEMBER 2026 EDITION

Methodology

Built on public-evidence data, standardized methodology, and up-to-date research using the latest AI models. A multidimensional view across Intelligence, Multimodality, Agent, Context, and Price.

I

Intelligence

50%

Reasoning, knowledge, coding, factuality, instruction following, and reliability.

  • →Primary benchmark leaderboards and public repositories
  • →Reproducible papers and published run logs
M

Multimodality

15%

Useful text, vision, image, audio, video, and document capabilities—not merely nominal input support.

  • →Public capability documentation
  • →Evidence of useful multimodal performance
A

Agent

15%

Tool use, planning, agentic coding, research, computer use, long-horizon completion, and recovery from errors.

  • →Published agent-task results
  • →Run logs and reproducible task evidence
C

Context

10%

Advertised context capacity plus demonstrated retrieval, comprehension, and long-context reasoning.

  • →Public context specifications
  • →Demonstrated long-context evidence
P

Price

10%

Public standard input/output pricing and cost per completed task where comparable data exist.

  • →Standard public API pricing
  • →Comparable task-economics evidence

Calculation

Overall score = 0.50 × Intelligence + 0.15 × Multimodality + 0.15 × Agent + 0.10 × Context + 0.10 × Price. The scorecard’s I, M, A, C and P columns are RS-normalized research-synthesis scores, not raw percentages from one benchmark.

Evidence policy & limitations

Primary evidence is original benchmark leaderboards, public repositories, reproducible papers and published run logs. Public model documentation and standard API specifications support coverage; vendor claims and aggregators are cross-checks only. Coverage may be incomplete and non-identical across frontier models, so the leading cluster should be read with roughly ±2–4 points of uncertainty.