---
title: "Best AI Models for Coding: Benchmarks and Prices"
description: "Coding models ranked by published benchmark, with the harness named in every row because the scores are not comparable across them."
source: "https://www.kensink.com/models/coding/"
canonical: "https://www.kensink.com/models/coding/"
---
★ 9 MODELS Model collection

MODEL COLLECTION · CODING AGENTS

# Models for coding agents, and the harness caveat.

Ranked by a published coding benchmark, with the benchmark named in every row because the scores are not comparable across it. Terminal-Bench 2.1 and Terminal-Bench 4.0 are different tests, and a model scoring 88.2 on one is not beating a model scoring 57.7 on the other. That caveat is the most useful thing on this page.

[Talk to us about model choice →](https://www.kensink.com/contact) [Browse by family](https://www.kensink.com/models/)

\[HOW WE BUILT THIS\]

Every model in our index with a published coding benchmark, ordered by score within its own benchmark. Numbers come from the vendor or the model card. We have deliberately not normalised across benchmarks or built a composite score, because doing so would invent a comparison the underlying data does not support. Where we have run our own evals, that is on customer tasks under NDA and is not what this table shows.

| Model | Vendor | Weights | Benchmark | Context | $ / 1M in · out |
| --- | --- | --- | --- | --- | --- |
| 
01

[GLM-5.3](https://www.kensink.com/models/glm/glm-5-3/)

zai-org/GLM-5.3

Level with GPT-5.6 Sol on Terminal-Bench 2.1 at a seventh of the input price, and state of the art on CyberGym vulnerability discovery.

 | Z.ai | Open | 88.2 Terminal-Bench 2.1 | 1M | $1.40 · $4.40 $0.26 cached |
| 

02

[GLM-5.3-Flash](https://www.kensink.com/models/glm/glm-5-3-flash/)

zai-org/GLM-5.3-Flash

MIT at 320B, multimodal, and four points behind its own flagship. The most permissive licence attached to a model this size.

 | Z.ai | Open | 84.3 Terminal-Bench 2.1 | 300K | Self-hosted |
| 

03

[Qwen3.8-Flash-Next](https://www.kensink.com/models/qwen/qwen3-8-flash-next/)

Qwen/Qwen3.8-Flash-Next

Beats Claude Opus by 22 points on AndroidWorld device automation. Frontier scores at roughly 6B inference cost, on a 180B machine.

 | Alibaba | Open | 62.5 SWE-bench Pro | 262K, 1M max | Self-hosted |
| 

04

[Qwen3.8-27B](https://www.kensink.com/models/qwen/qwen3-8-27b/)

Qwen/Qwen3.8-27B

Beats Claude Opus 4.6 Max on SWE-bench Pro and OSWorld in Alibaba's table, at 27B under Apache 2.0. Second most-liked model on Hugging Face.

 | Alibaba | Open | 61.7 SWE-bench Pro | 262K, 1M max | Self-hosted |
| 

05

[Claude Mythos 5.1](https://www.kensink.com/models/claude/fable-5-1/)

Fable 5.1's specs and price, restricted to Project Glasswing participants. Scores 60.9 on Terminal-Bench 4.0 against Fable's 55.8.

 | Anthropic | Closed | 60.9% Terminal-Bench 4.0 | 1M | $10 · $50 $0.25 cached |
| 

06

[GPT-6 Astra](https://www.kensink.com/models/openai-gpt/gpt-6-astra/)

First model rated Critical for cyber under the Preparedness Framework. Records on computer use and terminals, fourth on a third-party intelligence index.

 | OpenAI | Closed | 57.7% Terminal-Bench 4.0 | 1.05M | $10 · $50 $1 cached |
| 

07

[Claude Fable 5.1](https://www.kensink.com/models/claude/fable-5-1/)

Tops the Artificial Analysis Intelligence Index at roughly 66. Cache reads at $0.25, a quarter of Fable 5 and of every other Claude model.

 | Anthropic | Closed | 55.8% Terminal-Bench 4.0 | 1M | $10 · $50 $0.25 cached |
| 

08

[Claude Opus 5](https://www.kensink.com/models/claude/opus-5/)

Anthropic's recommended starting tier and the coding arena leader. 63.1 on the Artificial Analysis index at half Fable pricing.

 | Anthropic | Closed | 52.3% Terminal-Bench 4.0 | 1M | $5 · $25 |
| 

09

[GPT-5.6 Sol](https://www.kensink.com/models/openai-gpt/gpt-5-6/)

The prior OpenAI flagship and still the right default for most production work at two and a half times less than Astra.

 | OpenAI | Closed | 37.3% Terminal-Bench 4.0 | 1.05M | $4 · $20 $0.40 cached |

Every row carries the date we last verified it. Prices are list rates at the standard tier and exclude batch discounts, long-context surcharges and regional variation. Hugging Face download and like counts are pulled from the API rather than retyped, and they measure adoption rather than quality.

\[WHAT THE TABLE DOES NOT SAY\]

## Reading it properly.

01

### Read the benchmark column before the score column.

GLM-5.3 at 88.2 and GPT-6 Astra at 57.7 are on Terminal-Bench 2.1 and Terminal-Bench 4.0 respectively. The 4.0 suite is substantially harder. Sorting a table like this by raw score produces a ranking that looks authoritative and means very little, which is exactly why we name the harness in the row rather than in a footnote.

02

### The open-weight coding gap is roughly a price decision now.

GLM-5.3 runs level with GPT-5.6 Sol on Terminal-Bench 2.1 at $1.40 against $4.00 on input. Qwen3.8-Flash-Next posts 62.5 on SWE-bench Pro, ahead of DeepSeek-V4 at 56.0. For mainstream coding work the frontier premium buys less than it did six months ago, and it still buys real headroom on the hardest tasks.

03

### Watch where the gap opens, not where the headline sits.

GLM-5.3 is level with Sol on Terminal-Bench 2.1 and five points behind Fable 5 on the harder 3.0. That shape, competitive on ordinary work and falling off at the top end, is what a routing rule should encode. A single headline benchmark hides exactly the information you need.

04

### Forecastability beats the token price on agent work.

Z.ai's flat-rate Coding Plan at $18 to $168 a month is unusual and underrated. One runaway agent loop can cost more than a month of normal use, and teams are consistently more damaged by an unpredictable bill than a high one. Very few vendors offer this shape, and it unblocks adoption where a per-token discount does not.

\[METHODOLOGY · K-FRAMEWORK\]

## Integrated through the  
K-Framework.

Every model we integrate runs through the same operating system. Three pillars, sixteen layers, one Compound Growth Loop. The methodology that keeps AI work from rotting after the first ship.

[Read the K-Framework](https://www.kensink.com/k-framework)

01

### Foundations

Direct API integration with the model. No LangChain, no orchestration vendor, no agent framework built on quicksand. Typed contracts, the same way we wire up Postgres.

02

### Amplification

An eval suite built from your real tasks gates every prompt and model change. Quality is measured before it ships, not vibed in a demo.

03

### Judgment

Governance, audit, and oversight wired in from day one. Who called what, with which prompt version, at what cost. Your auditors get answers, not screenshots.

\[OBSERVABILITY\]

## Observability your team can read.

A model in production without observability is roulette. We instrument every integration so engineering and finance can see the same numbers, and so a regression at 3am surfaces before a customer opens a ticket.

Instrumented

### Cost per call

Tokens in, tokens out, dollars spent. Sliced by feature, tenant, and route. Budgets enforced where it matters.

Instrumented

### Latency p50 / p95 / p99

Real distributions, not averages. We know which routes are slow, and why.

Instrumented

### Eval pass rates

The same eval suite that gates a release runs continuously in production. A regression on real traffic surfaces fast.

Instrumented

### Prompt + completion logs

PII scrubbed at the proxy, shipped to your SIEM. Retention controls match your compliance window.

Dashboards your team owns, not ours. At handoff you get the queries, the alerts, and the runbook. We are not in the path to read your metrics.

\[COMMON QUESTIONS\]

## Questions we are getting asked.

Which model should we put behind our coding agent?

Route rather than pick. Send the volume to a cheap tier, the judgement-heavy steps to a frontier model, and measure cost per completed task rather than cost per call. If we had to name one starting point for a team with no constraints, it is Claude Opus 5 for quality and GLM-5.3 for the same work at roughly a seventh of the price.

Are these scores comparable?

Only within a single benchmark. Terminal-Bench 2.1, 3.0 and 4.0 are different difficulty levels, and SWE-bench Pro, DeepSWE and CursorBench measure different things again. We name the harness in every row precisely so nobody reads down the score column and draws a conclusion the data will not carry.

Do vendor benchmarks mean anything?

They are evidence, run by an interested party on a comparison set they chose. Several on this page were measured at maximum reasoning effort, which is not how anyone runs a model in production. Treat them as a shortlist mechanism, then run your own eval suite on your own tasks at the effort level you will actually pay for.

Is an open-weight model good enough for production coding?

For mainstream work, increasingly yes. GLM-5.3 matches GPT-5.6 Sol on Terminal-Bench 2.1 and Qwen3.8-Flash-Next beats Claude Opus on several agentic tasks in Alibaba's numbers. The frontier still leads on the hardest long-horizon work, so the sensible architecture routes by difficulty rather than committing the whole workload to one tier.

Share[](https://twitter.com/intent/tweet?url=https%3A%2F%2Fwww.kensink.com%2Fmodels%2Fcoding%2F&text=Coding)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fwww.kensink.com%2Fmodels%2Fcoding%2F)

[View .md](https://www.kensink.com/models/coding.md)

\[OTHER COLLECTIONS\]

## Same models, different question.

[

COLLECTION

The top 50

View

](https://www.kensink.com/models/top-50/)[

COLLECTION

Open weights

View

](https://www.kensink.com/models/open-weight-2026/)[

COLLECTION

Long context

View

](https://www.kensink.com/models/long-context/)[

COLLECTION

Pricing

View

](https://www.kensink.com/models/pricing/)

DIRECT INTEGRATION · NO FRAMEWORK

## Picking a model  
is a routing decision.

We build behind a vendor-neutral abstraction and route by task difficulty at runtime, with an eval suite that decides rather than a launch post. Eval suite at handoff, full source ownership.

[Start a conversation →](https://www.kensink.com/contact) [Browse model families](https://www.kensink.com/models/)
