01GPT-6 AstraFirst model rated Critical for cyber under the Preparedness Framework. Records on computer use and terminals, fourth on a third-party intelligence index. | OpenAI | Closed | 1.05M | $10 · $50$1 cached | Proprietary |
02GPT-5.6 SolThe prior OpenAI flagship and still the right default for most production work at two and a half times less than Astra. | OpenAI | Closed | 1.05M | $4 · $20$0.40 cached | Proprietary |
03GPT-5.6 LunaOpenAI's cheap tier. Classification, routing and extraction, which is where most of an agent's token volume actually goes once you measure it. | OpenAI | Closed | 1.05M | $0.20 · $1.20$0.02 cached | Proprietary |
04GPT-5.6 TerraThe balanced OpenAI tier, and where most of the token volume belongs in a working system. Same context and tool surface as Sol. | OpenAI | Closed | 1.05M | $2 · $12$0.20 cached | Proprietary |
05Claude Fable 5.1Tops the Artificial Analysis Intelligence Index at roughly 66. Cache reads at $0.25, a quarter of Fable 5 and of every other Claude model. | Anthropic | Closed | 1M | $10 · $50$0.25 cached | Proprietary |
06Claude Opus 5Anthropic's recommended starting tier and the coding arena leader. 63.1 on the Artificial Analysis index at half Fable pricing. | Anthropic | Closed | 1M | $5 · $25 | Proprietary |
07Gemini 3.1 ProLeads OCR and visual question answering on the Nanonets document leaderboard, and handles sparse tables at 94% where most pipelines break. | Google | Closed | 1M | Preview | Proprietary |
08Claude Mythos 5.1Fable 5.1's specs and price, restricted to Project Glasswing participants. Scores 60.9 on Terminal-Bench 4.0 against Fable's 55.8. | Anthropic | Closed | 1M | $10 · $50$0.25 cached | Proprietary, invitation only |
09Qwen3.8-FlashThe cheapest usable tier we would put in production, flat across a 1M context with no long-context surcharge. Roughly seventy times cheaper than the frontier on input. | Alibaba | Closed | 1M | $0.14 · $0.42 | Proprietary |
10Qwen3.8-MaxAlibaba's hosted flagship with no public weights, at a fifth of Claude Opus 5 on input. The Beijing endpoint runs 60 to 70% cheaper than Singapore. | Alibaba | Closed | 1M | $2 · $6$0.25 cached | Proprietary |
11GLM-5.3zai-org/GLM-5.3 Level with GPT-5.6 Sol on Terminal-Bench 2.1 at a seventh of the input price, and state of the art on CyberGym vulnerability discovery. | Z.ai | Open | 1M | $1.40 · $4.40$0.26 cached | glm-5.3, bespoke |
12Kimi K3moonshotai/Kimi-K3 The largest open-weight model shipped, at 2.8 trillion parameters, and third on the Artificial Analysis index by third-party trackers. | Moonshot | Open | 1M | Self-hosted | Modified MIT |
13GLM-5.3-Flashzai-org/GLM-5.3-Flash MIT at 320B, multimodal, and four points behind its own flagship. The most permissive licence attached to a model this size. | Z.ai | Open | 300K | Self-hosted | MIT |
14Qwen3.8-27BQwen/Qwen3.8-27B Beats Claude Opus 4.6 Max on SWE-bench Pro and OSWorld in Alibaba's table, at 27B under Apache 2.0. Second most-liked model on Hugging Face. | Alibaba | Open | 262K, 1M max | Self-hosted | Apache 2.0 |
15Qwen3.8-Flash-NextQwen/Qwen3.8-Flash-Next Beats Claude Opus by 22 points on AndroidWorld device automation. Frontier scores at roughly 6B inference cost, on a 180B machine. | Alibaba | Open | 262K, 1M max | Self-hosted | qwen-community-1.0 |