Best for high-volume agentic work on a tight budget.
- 1M context
- Text only
- Input
- $0.14 / 1M
- Cache read
- $0.028 / 1M
- Output
- $0.28 / 1M
- Intelligence
- 34
- DeepSWE
- 54.4%* reported by the publisher
- Terminal-Bench
- 78.7%
Models
Every model we serve, side by side: list prices, context windows, and scores on the benchmarks we trust for agentic coding.
Best for high-volume agentic work on a tight budget.
Best for fast agentic coding at a low price.
Best for complex, long-horizon coding.
Best for the hardest, highest-stakes tasks.
Best for everyday coding at a flash price.
Best for fast side tasks: scouting, summaries.
Scores as of October 3, 2026. Each model's page links every score to its source.
By Artificial Analysis
An independent composite of reasoning, knowledge, maths, coding and agentic evaluations, run by Artificial Analysis on every model the same way.
MethodologyBy Datacurve
Long-horizon software engineering: real repository tasks an agent has to finish end to end.
MethodologyBy tbench.ai
Hard, realistic tasks an agent completes in a terminal sandbox: building, debugging, configuring.
Methodology* Reported by the model's publisher on its model card. Publishers run their own harnesses, so these compare less cleanly than the independent measurements, which carry no mark. No number here is estimated: a model without one has not been measured on that benchmark yet.
Create a key once, then switch models by changing one string in your tool.
Looking for a model we no longer serve? See past models.