Models

Pick the right model.

Every model we serve, side by side: list prices, context windows, and scores on the benchmarks we trust for agentic coding.

  • Best for high-volume agentic work on a tight budget.

    • 1M context
    • Text only
    Input
    $0.14 / 1M
    Cache read
    $0.028 / 1M
    Output
    $0.28 / 1M
    Intelligence
    34
    DeepSWE
    54.4%* reported by the publisher
    Terminal-Bench
    78.7%
  • Best for fast agentic coding at a low price.

    • 1M context
    • Vision
    Input
    $0.15 / 1M
    Cache read
    $0.028 / 1M
    Output
    $0.60 / 1M
    Intelligence
    39
    DeepSWE
    74.2%* reported by the publisher
    Terminal-Bench
    90.6%* reported by the publisher
  • Best for complex, long-horizon coding.

    • 1M context
    • Text only
    Input
    $1.40 / 1M
    Cache read
    $0.26 / 1M
    Output
    $4.40 / 1M
    Intelligence
    45
    DeepSWE
    69%
    Terminal-Bench
    83.9%
  • Best for the hardest, highest-stakes tasks.

    • 1M context
    • Vision
    Input
    $3.00 / 1M
    Cache read
    $0.30 / 1M
    Output
    $15.00 / 1M
    Intelligence
    44
    DeepSWE
    68.5%
    Terminal-Bench
    85%
  • Best for everyday coding at a flash price.

    • 1M context
    • Vision
    Input
    $0.15 / 1M
    Cache read
    $0.03 / 1M
    Output
    $0.50 / 1M
    Intelligence
    42
    DeepSWE
    63.4%
    Terminal-Bench
    84.3%
  • Best for fast side tasks: scouting, summaries.

    • 256K context
    • Vision
    Input
    $0.15 / 1M
    Cache read
    $0.05 / 1M
    Output
    $1.00 / 1M
    Intelligence
    18
    DeepSWE
    n/a
    Terminal-Bench
    44.9%

About the scores

Scores as of October 3, 2026. Each model's page links every score to its source.

Artificial Analysis Intelligence Index

By Artificial Analysis

An independent composite of reasoning, knowledge, maths, coding and agentic evaluations, run by Artificial Analysis on every model the same way.

Methodology

DeepSWE

By Datacurve

Long-horizon software engineering: real repository tasks an agent has to finish end to end.

Methodology

Terminal-Bench 2.1

By tbench.ai

Hard, realistic tasks an agent completes in a terminal sandbox: building, debugging, configuring.

Methodology

* Reported by the model's publisher on its model card. Publishers run their own harnesses, so these compare less cleanly than the independent measurements, which carry no mark. No number here is estimated: a model without one has not been measured on that benchmark yet.

One key, every model.

Create a key once, then switch models by changing one string in your tool.

Looking for a model we no longer serve? See past models.

Coding models: API pricing and benchmarks · Umans AI