Skip to content
AZ Labs

AZ Labs rankings · Updated 3 October 2026

The 10 best AI models right now

Ranked purely on independent benchmarks: the Artificial Analysis Intelligence Index, 10 evaluations covering agents, coding, science and knowledge. No vendor claims, no paid placement, no votes.

42 current models tracked10 benchmarks · index v4.3.2100% independent evals

The ranking. 100% Artificial Analysis Intelligence Index.

#1
57.6AA index

Anthropic

Claude Opus 5.5

Proprietary

Hard reasoning, long agentic work, anything that has to be right

Intelligence Index (ranked)57.6 / 100
LMArena (context only)1509

Anthropic's flagship. Leads the index with room to spare and is also the top model on LMArena.

Published evals HLE 61.4% · SciCode 66.9% · AA-LCR 84.7%

$8/1M · 97 tok/s · 1M ctx
#2
56.0AA index

Anthropic

Claude Sonnet 5.5

Intelligence Index (ranked)56.0 / 100
LMArena (context only)not listed

Scores 56 on the Intelligence Index.

Published evals HLE 55.0% · SciCode 61.0% · AA-LCR 82.7%

$4/1M · 141 tok/s
#3
53.4AA index

Anthropic

Claude Fable 5.1

Proprietary

Writing, nuance and long-form work

Intelligence Index (ranked)53.4 / 100
LMArena (context only)1501

Anthropic's premium writing-focused model. Near the top on benchmarks, and among the most expensive per token.

Published evals HLE 59.1% · SciCode 63.1% · AA-LCR 85.3% · Coding Index 81.6

  1. 04

    GPT-6 Astra

    OpenAIProprietary
    52.7
    Intelligence Index (ranked)52.7 / 100
    LMArena (context only)1478
    AA index
    52.7
    setting · Max
    Arena
    1478
    ±8 · 7.2K votes
    Price / 1M
    $20
    $10 in · $50 out
    Speed
    60
    tok/s median
    Context
    1M
    tokens

    OpenAI's frontier reasoning model. OpenAI's strongest model, neck and neck with Fable 5.1 on the index.

    Published evals HLE 54.7% · SciCode 56.5% · AA-LCR 80.7% · Coding Index 76.9

  2. 05

    Gemini 4 Argon

    Google
    52.6
    Intelligence Index (ranked)52.6 / 100
    LMArena (context only)not listed
    AA index
    52.6
    setting · High
    Arena
    —
    not listed
    Price / 1M
    $4
    $2 in · $10 out
    Speed
    —
    tok/s median
    Context
    —
    tokens

    Scores 52.6 on the Intelligence Index.

    Published evals HLE 57.1% · SciCode 61.8% · AA-LCR 79.7%

  3. 06

    GPT-6.1 Sol

    OpenAI
    51.8
    Intelligence Index (ranked)51.8 / 100
    LMArena (context only)not listed
    AA index
    51.8
    setting · Max
    Arena
    —
    not listed
    Price / 1M
    $4
    $2 in · $10 out
    Speed
    55
    tok/s median
    Context
    —
    tokens

    Scores 51.8 on the Intelligence Index.

    Published evals HLE 52.9% · SciCode 54.2% · AA-LCR 83.0%

  4. 07

    Muse Spark 1.3

    MetaProprietary
    48.1
    Intelligence Index (ranked)48.1 / 100
    LMArena (context only)1494
    AA index
    48.1
    setting · Max
    Arena
    1494
    ±7 · 10K votes
    Price / 1M
    $2
    $1.25 in · $4.25 out
    Speed
    162
    tok/s median
    Context
    1M
    tokens

    Near-frontier answers at high speed. Meta's best. One of the fastest models near the top of the board, at a low per-token price.

    Published evals HLE 48.7% · SciCode 58.8% · AA-LCR 83.0% · Coding Index 75.8

  5. 08

    GPT-6 Sol

    OpenAIProprietary
    47.6
    Intelligence Index (ranked)47.6 / 100
    LMArena (context only)1457
    AA index
    47.6
    setting · Max
    Arena
    1457
    ±9 · 4.8K votes
    Price / 1M
    $4
    $2 in · $10 out
    Speed
    130
    tok/s median
    Context
    872K
    tokens

    GPT-6 reasoning at a lower price. The mid-tier GPT-6 model: a few points behind Astra at a fraction of the price.

    Published evals HLE 47.9% · SciCode 57.6% · AA-LCR 83.7%

  6. 09

    Grok 4.7

    SpaceXAIProprietary
    46.4
    Intelligence Index (ranked)46.4 / 100
    LMArena (context only)1439
    AA index
    46.4
    setting · Xhigh
    Arena
    1439
    ±9 · 4.1K votes
    Price / 1M
    $3
    $2 in · $6 out
    Speed
    82
    tok/s median
    Context
    500K
    tokens

    The strongest model in the xAI stack. xAI's newest flagship. LMArena voters rate it well below its benchmark score.

    Published evals HLE 43.1% · SciCode 57.4% · AA-LCR 76.7%

  7. 10

    MiMo-V2.6-Pro

    XiaomiOpen weights
    46.3
    Intelligence Index (ranked)46.3 / 100
    LMArena (context only)1480
    AA index
    46.3
    setting · default
    Arena
    1480
    ±9 · 4.0K votes
    Price / 1M
    $0.54
    $0.43 in · $0.87 out
    Speed
    44
    tok/s median
    Context
    1M
    tokens

    Absurd value for an open model. The top open-weights model on the index, MIT-licensed, and priced below every other model in the top ten.

    Published evals HLE 49.4% · SciCode 60.9% · AA-LCR 86.3%

Category winners

Picked from the benchmark top ten.

Best under $5 per 1M tokens

Claude Sonnet 5.5

#2 at $4

Highest-ranked model with a blended API price under $5 per million tokens.

Fastest

Muse Spark 1.3

162 tok/s

Highest median output speed in this top ten.

Lowest price

MiMo-V2.6-Pro

$0.54 / 1M

Cheapest blended API price (3:1 input to output) in this top ten.

Best open weights

MiMo-V2.6-Pro

#10 overall

Highest-ranked model you can download and self-host.

Benchmarks vs blind votes

The ranking uses the horizontal axis only. The vertical axis shows how LMArena voters rate the same models, so you can see where the two disagree. Hover a point for details.

Top 10 Ranks 11+ Hollow = open weights
14401460148015004045505560Artificial Analysis Intelligence Index (the ranking) →LMArena score →Kimi K3GPT-6 SolGPT-6 Astra

Just missed the cut

The next five current models on the index.

RankModelAA indexArenaPrice / 1MNote
#11Qwen3.8 MaxAlibaba45.41479$3Alibaba's flagship. Strong on benchmarks but one of the slower models in the list.
#12GLM-5.3Z AI44.81480$2.15Z.ai's flagship, MIT-licensed, with no weak spot on either leaderboard.
#13Grok 4.6SpaceXAI44.31453$3The previous Grok flagship, still close behind its successor on the index.
#14Step 5 PreviewStepFun43.7—$1.43StepFun's preview release. A competitive score at a low price; not on LMArena yet.
#15Kimi K3Kimi43.61488$6Moonshot's open-weights flagship. Voters rank it above every other open model, but it is slow.

Scores, prices and speeds: Artificial Analysis Data API via AIMI, fetched 2026-10-03. Arena scores: LMArena, 2026-09-25. 42 current models tracked.

How we rank

  1. 1

    Pick the pool

    Every current model Artificial Analysis publishes, pulled from its Data API into AIMI daily and on every new release. 42 are tracked here; superseded versions are excluded.

  2. 2

    Take each model at its best

    Most models ship several reasoning settings (low, high, max). Each model is scored at its strongest listed setting, and we show which setting that was.

  3. 3

    Sort by the Intelligence Index

    The published score is used as-is. No normalising, no weighting of our own, no blending with other sources.

  4. 4

    Nothing else moves a model

    Scores come to one decimal, so ties are rare; an exact tie ranks the newer release first. Votes, price and speed never change the order.

Artificial Analysis Intelligence Index v4.3.2

Weighted average of 10 independent evaluations run by Artificial Analysis: agents 30%, coding 20%, scientific reasoning 20%, general 30%. Pulled from the Artificial Analysis Data API into AIMI daily and whenever a new model appears. This is the only input to the ranking.

LMArena Text Arena (overall)

Elo-style score from 8.5M+ blind head-to-head votes, leaderboard dated 25 Sep 2026. Shown for context only; it does not affect the ranking.

Read the fine print. "Price / 1M" is the API price per million tokens published by Artificial Analysis, blended 3:1 input to output. It is not a subscription price. Speed is the median output rate. Context windows and licences come from each model's Artificial Analysis page. Arena scores carry a ± margin that is wider for newer models.

What is inside the Intelligence Index

The 10 evaluations behind every score on this page, with the weight each carries in index v4.3.2. Private sets are held back by Artificial Analysis so models cannot be trained on them. Full methodology

CategoryEvaluationWeightWhat it testsSizeScoring
AgentsAA-Briefcase v1.1Private15%Agentic knowledge work that ends in real file deliverables91 tasks, 4 scenariosElo from pairwise comparisons of task success, analysis and presentation
AgentsGDPval-AA v2.110%Economically valuable professional tasks with file outputs220 tasksPairwise Elo by a judge panel
AgentsAutomationBench-AAPrivate5%SaaS workflow automation through REST API tools657 tasksTask completion; zero credit if a guardrail is violated
CodingTerminal-Bench 4.010%Real tasks executed in a terminal66 tasks × 3 runsTest suite pass/fail, pass@1
CodingSciCode10%Scientific Python that must pass every unit test288 subproblems × 3 runsCode execution, pass@1
GeneralAA-OmnisciencePrivate15%Knowledge accuracy and how often the model makes things up6,000 questionsAccuracy (10%) plus 1 − hallucination rate (5%)
GeneralGDP.pdf10%Answers grounded in long PDF documents across 10 domains100 tasks × 5 runsAll-pass headline and mean pass rate
GeneralAA-LCR v1.15%Long-context reasoning over very large inputs100 questions × 3 runsLLM equality checker, pass@1
Scientific reasoningHumanity's Last Exam10%Expert-level questions across academic fields2,158 questionsLLM equality checker, pass@1
Scientific reasoningCritPtPrivate10%Research-level physics problems70 problems × 5 runsOfficial grading server, pass@1

Questions

What is the best AI model right now?

On our 3 October 2026 snapshot, Claude Opus 5.5 is the best model overall. It scores 57.6 on the Artificial Analysis Intelligence Index v4.3.2, ahead of Claude Sonnet 5.5 on 56.0.

What is the ranking based on?

Only the Artificial Analysis Intelligence Index v4.3.2. It is a weighted average of 10 independent evaluations: agents 30% (AA-Briefcase, GDPval-AA, AutomationBench-AA), coding 20% (Terminal-Bench 4.0, SciCode), general 30% (AA-Omniscience, GDP.pdf, AA-LCR) and scientific reasoning 20% (Humanity's Last Exam, CritPt). We take each model at its best reasoning setting and sort by that score.

Where do the numbers come from?

The Artificial Analysis Data API. Our model catalogue, AIMI, pulls it daily and as soon as it spots a new model release, keeps the raw response as evidence, records every score with its date, and publishes the result straight to this page. No scores are typed in by hand.

Why show LMArena votes if they do not count?

Benchmarks measure whether a model gets hard problems right. Blind votes show whether people prefer its answers. The two often disagree: Gemini 3.8 Flash is popular with voters but ranks #19 on benchmarks. So votes are shown as context and kept out of the ranking.

What is the best open-weights AI model?

MiMo-V2.6-Pro, at #10 with 46.3 on the index.

How often is this ranking updated?

Automatically. AIMI checks Artificial Analysis once a day, and within about half an hour of spotting a new model release. When anything changes, this page updates on its own. The date at the top shows when the data was last fetched.

Try models from one API

The AZ Labs Model Gateway puts a growing catalogue behind one OpenAI-compatible endpoint, with live health and latency for every route.