Kensink Labs
9 MODELSModel collection
MODEL COLLECTION · CODING AGENTS

Models for coding agents, and the harness caveat.

Ranked by a published coding benchmark, with the benchmark named in every row because the scores are not comparable across it. Terminal-Bench 2.1 and Terminal-Bench 4.0 are different tests, and a model scoring 88.2 on one is not beating a model scoring 57.7 on the other. That caveat is the most useful thing on this page.

[HOW WE BUILT THIS]

Every model in our index with a published coding benchmark, ordered by score within its own benchmark. Numbers come from the vendor or the model card. We have deliberately not normalised across benchmarks or built a composite score, because doing so would invent a comparison the underlying data does not support. Where we have run our own evals, that is on customer tasks under NDA and is not what this table shows.

ModelVendorWeightsBenchmarkContext$ / 1M in · out
01
GLM-5.3
zai-org/GLM-5.3

Level with GPT-5.6 Sol on Terminal-Bench 2.1 at a seventh of the input price, and state of the art on CyberGym vulnerability discovery.

Z.aiOpen88.2Terminal-Bench 2.11M$1.40 · $4.40$0.26 cached
02
GLM-5.3-Flash
zai-org/GLM-5.3-Flash

MIT at 320B, multimodal, and four points behind its own flagship. The most permissive licence attached to a model this size.

Z.aiOpen84.3Terminal-Bench 2.1300KSelf-hosted
03
Qwen3.8-Flash-Next
Qwen/Qwen3.8-Flash-Next

Beats Claude Opus by 22 points on AndroidWorld device automation. Frontier scores at roughly 6B inference cost, on a 180B machine.

AlibabaOpen62.5SWE-bench Pro262K, 1M maxSelf-hosted
04
Qwen3.8-27B
Qwen/Qwen3.8-27B

Beats Claude Opus 4.6 Max on SWE-bench Pro and OSWorld in Alibaba's table, at 27B under Apache 2.0. Second most-liked model on Hugging Face.

AlibabaOpen61.7SWE-bench Pro262K, 1M maxSelf-hosted
05
Claude Mythos 5.1

Fable 5.1's specs and price, restricted to Project Glasswing participants. Scores 60.9 on Terminal-Bench 4.0 against Fable's 55.8.

AnthropicClosed60.9%Terminal-Bench 4.01M$10 · $50$0.25 cached
06
GPT-6 Astra

First model rated Critical for cyber under the Preparedness Framework. Records on computer use and terminals, fourth on a third-party intelligence index.

OpenAIClosed57.7%Terminal-Bench 4.01.05M$10 · $50$1 cached
07
Claude Fable 5.1

Tops the Artificial Analysis Intelligence Index at roughly 66. Cache reads at $0.25, a quarter of Fable 5 and of every other Claude model.

AnthropicClosed55.8%Terminal-Bench 4.01M$10 · $50$0.25 cached
08
Claude Opus 5

Anthropic's recommended starting tier and the coding arena leader. 63.1 on the Artificial Analysis index at half Fable pricing.

AnthropicClosed52.3%Terminal-Bench 4.01M$5 · $25
09
GPT-5.6 Sol

The prior OpenAI flagship and still the right default for most production work at two and a half times less than Astra.

OpenAIClosed37.3%Terminal-Bench 4.01.05M$4 · $20$0.40 cached

Every row carries the date we last verified it. Prices are list rates at the standard tier and exclude batch discounts, long-context surcharges and regional variation. Hugging Face download and like counts are pulled from the API rather than retyped, and they measure adoption rather than quality.

[WHAT THE TABLE DOES NOT SAY]

Reading it properly.

01

Read the benchmark column before the score column.

GLM-5.3 at 88.2 and GPT-6 Astra at 57.7 are on Terminal-Bench 2.1 and Terminal-Bench 4.0 respectively. The 4.0 suite is substantially harder. Sorting a table like this by raw score produces a ranking that looks authoritative and means very little, which is exactly why we name the harness in the row rather than in a footnote.

02

The open-weight coding gap is roughly a price decision now.

GLM-5.3 runs level with GPT-5.6 Sol on Terminal-Bench 2.1 at $1.40 against $4.00 on input. Qwen3.8-Flash-Next posts 62.5 on SWE-bench Pro, ahead of DeepSeek-V4 at 56.0. For mainstream coding work the frontier premium buys less than it did six months ago, and it still buys real headroom on the hardest tasks.

03

Watch where the gap opens, not where the headline sits.

GLM-5.3 is level with Sol on Terminal-Bench 2.1 and five points behind Fable 5 on the harder 3.0. That shape, competitive on ordinary work and falling off at the top end, is what a routing rule should encode. A single headline benchmark hides exactly the information you need.

04

Forecastability beats the token price on agent work.

Z.ai's flat-rate Coding Plan at $18 to $168 a month is unusual and underrated. One runaway agent loop can cost more than a month of normal use, and teams are consistently more damaged by an unpredictable bill than a high one. Very few vendors offer this shape, and it unblocks adoption where a per-token discount does not.

[METHODOLOGY · K-FRAMEWORK]

Integrated through the
K-Framework.

Every model we integrate runs through the same operating system. Three pillars, sixteen layers, one Compound Growth Loop. The methodology that keeps AI work from rotting after the first ship.

Read the K-Framework
01

Foundations

Direct API integration with the model. No LangChain, no orchestration vendor, no agent framework built on quicksand. Typed contracts, the same way we wire up Postgres.

02

Amplification

An eval suite built from your real tasks gates every prompt and model change. Quality is measured before it ships, not vibed in a demo.

03

Judgment

Governance, audit, and oversight wired in from day one. Who called what, with which prompt version, at what cost. Your auditors get answers, not screenshots.

[OBSERVABILITY]

Observability your team can read.

A model in production without observability is roulette. We instrument every integration so engineering and finance can see the same numbers, and so a regression at 3am surfaces before a customer opens a ticket.

Instrumented

Cost per call

Tokens in, tokens out, dollars spent. Sliced by feature, tenant, and route. Budgets enforced where it matters.

Instrumented

Latency p50 / p95 / p99

Real distributions, not averages. We know which routes are slow, and why.

Instrumented

Eval pass rates

The same eval suite that gates a release runs continuously in production. A regression on real traffic surfaces fast.

Instrumented

Prompt + completion logs

PII scrubbed at the proxy, shipped to your SIEM. Retention controls match your compliance window.

Dashboards your team owns, not ours. At handoff you get the queries, the alerts, and the runbook. We are not in the path to read your metrics.

[COMMON QUESTIONS]

Questions we are getting asked.

Which model should we put behind our coding agent?
Route rather than pick. Send the volume to a cheap tier, the judgement-heavy steps to a frontier model, and measure cost per completed task rather than cost per call. If we had to name one starting point for a team with no constraints, it is Claude Opus 5 for quality and GLM-5.3 for the same work at roughly a seventh of the price.
Are these scores comparable?
Only within a single benchmark. Terminal-Bench 2.1, 3.0 and 4.0 are different difficulty levels, and SWE-bench Pro, DeepSWE and CursorBench measure different things again. We name the harness in every row precisely so nobody reads down the score column and draws a conclusion the data will not carry.
Do vendor benchmarks mean anything?
They are evidence, run by an interested party on a comparison set they chose. Several on this page were measured at maximum reasoning effort, which is not how anyone runs a model in production. Treat them as a shortlist mechanism, then run your own eval suite on your own tasks at the effort level you will actually pay for.
Is an open-weight model good enough for production coding?
For mainstream work, increasingly yes. GLM-5.3 matches GPT-5.6 Sol on Terminal-Bench 2.1 and Qwen3.8-Flash-Next beats Claude Opus on several agentic tasks in Alibaba's numbers. The frontier still leads on the hardest long-horizon work, so the sensible architecture routes by difficulty rather than committing the whole workload to one tier.
Share
View .md
DIRECT INTEGRATION · NO FRAMEWORK

Picking a model
is a routing decision.

We build behind a vendor-neutral abstraction and route by task difficulty at runtime, with an eval suite that decides rather than a launch post. Eval suite at handoff, full source ownership.