| Frontier, API only7 |
|---|
01Claude Fable 5.1Tops the Artificial Analysis Intelligence Index at roughly 66. Cache reads at $0.25, a quarter of Fable 5 and of every other Claude model. | Anthropic | Closed | Proprietary | 1M | $10 · $50$0.25 cached | 1 Sep 2026 |
02Claude Mythos 5.1Fable 5.1's specs and price, restricted to Project Glasswing participants. Scores 60.9 on Terminal-Bench 4.0 against Fable's 55.8. | Anthropic | Closed | Proprietary, invitation only | 1M | $10 · $50$0.25 cached | 1 Sep 2026 |
03Claude Opus 5Anthropic's recommended starting tier and the coding arena leader. 63.1 on the Artificial Analysis index at half Fable pricing. | Anthropic | Closed | Proprietary | 1M | $5 · $25 | 24 Jul 2026 |
04Gemini 3.1 ProLeads OCR and visual question answering on the Nanonets document leaderboard, and handles sparse tables at 94% where most pipelines break. | Google | Closed | Proprietary | 1M | Preview | 2026 |
05GPT-5.6 SolThe prior OpenAI flagship and still the right default for most production work at two and a half times less than Astra. | OpenAI | Closed | Proprietary | 1.05M | $4 · $20$0.40 cached | 26 Jun 2026 |
06GPT-6 AstraFirst model rated Critical for cyber under the Preparedness Framework. Records on computer use and terminals, fourth on a third-party intelligence index. | OpenAI | Closed | Proprietary | 1.05M | $10 · $50$1 cached | 3 Sep 2026 |
07Grok 4.6Consistently in the top five on public leaderboards. We have not run it on customer tasks, so we do not have a view worth publishing yet. | xAI | Closed | Proprietary | — | — | 2026 |
| Open-weight frontier10 |
|---|
08Qwen3.8-27BQwen/Qwen3.8-27B Beats Claude Opus 4.6 Max on SWE-bench Pro and OSWorld in Alibaba's table, at 27B under Apache 2.0. Second most-liked model on Hugging Face. | Alibaba | Open | Apache 2.0 | 262K, 1M max | Self-hosted | 5 Aug 2026 |
09Kimi K3moonshotai/Kimi-K3 The largest open-weight model shipped, at 2.8 trillion parameters, and third on the Artificial Analysis index by third-party trackers. | Moonshot | Open | Modified MIT | 1M | Self-hosted | 16 Jul 2026 |
10DeepSeek-V4-Prodeepseek-ai/DeepSeek-V4-Pro The DeepSeek flagship, MIT licensed. Second on CyberGym behind GLM-5.3, and a long-standing favourite for self-hosted agent work. | DeepSeek | Open | MIT | — | Self-hosted | 2026 |
11gpt-oss-120bopenai/gpt-oss-120b OpenAI's open-weight release under Apache 2.0. Notable less for its scores than for existing at all, and 5.2M downloads a month. | OpenAI | Open | Apache 2.0 | — | Self-hosted | 2025 |
12Qwen3.8-Flash-NextQwen/Qwen3.8-Flash-Next Beats Claude Opus by 22 points on AndroidWorld device automation. Frontier scores at roughly 6B inference cost, on a 180B machine. | Alibaba | Open | qwen-community-1.0 | 262K, 1M max | Self-hosted | 24 Aug 2026 |
13DeepSeek-V4-Flash-0731deepseek-ai/DeepSeek-V4-Flash-0731 MIT at 304B and 4.5M downloads a month. Widely described as the strongest cost-and-agent option among downloadable weights. | DeepSeek | Open | MIT | — | Self-hosted | 31 Jul 2026 |
14gemma-4-31B-itgoogle/gemma-4-31B-it Google's open-weight line, 8.6M downloads a month. The gemma-4 family dominates the any-to-any category on Hugging Face. | Google | Open | Gemma Terms of Use | — | Self-hosted | 2026 |
15GLM-5.3-Flashzai-org/GLM-5.3-Flash MIT at 320B, multimodal, and four points behind its own flagship. The most permissive licence attached to a model this size. | Z.ai | Open | MIT | 300K | Self-hosted | 25 Aug 2026 |
16GLM-5.3zai-org/GLM-5.3 Level with GPT-5.6 Sol on Terminal-Bench 2.1 at a seventh of the input price, and state of the art on CyberGym vulnerability discovery. | Z.ai | Open | glm-5.3, bespoke | 1M | $1.40 · $4.40$0.26 cached | 25 Aug 2026 |
17Muse Glimmer 30BMeta's move to Apache 2.0 and to a dense architecture, away from the Llama Community Licence and mixture-of-experts. | Meta | Open | Apache 2.0 | — | Self-hosted | Aug 2026 |
| Small and on-device3 |
|---|
18Qwen3-0.6BQwen/Qwen3-0.6B 21M downloads a month, more than any frontier open model. Small enough for on-device, browser and edge-runtime inference. | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2025 |
19Qwen3-8BQwen/Qwen3-8B The workhorse size. Still 13M downloads a month a year after release, and where most self-hosted pilots start. | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2025 |
20Spark-X2.5-4BXHToken/Spark-X2.5-4B Top of the Hugging Face trending list on the day we checked, on 697 likes against 7.2k downloads. Attention arriving well ahead of adoption. | XHToken | Open | See model card | — | Self-hosted | Sep 2026 |
| Vision-language and OCR3 |
|---|
21Unlimited-OCRbaidu/Unlimited-OCR 2.7M downloads and 4.2k likes. Document extraction as a dedicated model rather than a prompt to a general vision-language model. | Baidu | Open | See model card | — | Self-hosted | 2026 |
22Qwen3-VL-8B-InstructQwen/Qwen3-VL-8B-Instruct 14.5M downloads a month. Fits on a single accelerator, which makes it the open-weight default for document and screen work. | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2026 |
23PaddleOCR VL 1.5Tops OmniDocBench at 94.37 overall, ahead of every general model on that test. A parser rather than a reasoner. | Open weights | Open | Apache 2.0 | — | Self-hosted | 2026 |
| Image generation5 |
|---|
24FLUX.1-devblack-forest-labs/FLUX.1-dev The most-liked model on Hugging Face, full stop, at 14,508 likes. Note the licence: the dev weights are non-commercial. | Black Forest Labs | Open | Non-commercial, dev | — | Self-hosted | 2024 |
25FLUX.1-schnellblack-forest-labs/FLUX.1-schnell The commercially usable FLUX. Apache 2.0 where the dev weights are not, which is the version most products actually ship. | Black Forest Labs | Open | Apache 2.0 | — | Self-hosted | 2024 |
26Z-Image-TurboTongyi-MAI/Z-Image-Turbo 5.2k likes and 690k downloads. The Alibaba-adjacent entry in open image generation, and a genuine FLUX alternative. | Tongyi-MAI | Open | See model card | — | Self-hosted | 2026 |
27GPT-Image-2Token-priced rather than per-image, so you cannot quote a cost per picture until you measure your own prompts. | OpenAI | Closed | Proprietary | — | $5 · $30$1.25 cached | Apr 2026 |
28Nano Banana ProPer-image pricing you can put in a quote: $0.134 at 1K or 2K, $0.24 at 4K, half that on batch. SynthID on every output. | Google | Closed | Proprietary | — | $0.134 / image | 2026 |
| Video generation2 |
|---|
29MiniMax-H3MiniMaxAI/MiniMax-H3 5M downloads and 5k likes, with a whole ecosystem of turbo, quantised and LoRA derivatives already on Hugging Face. | MiniMaxAI | Open | See model card | — | Self-hosted | 2026 |
30LTX-2.5Lightricks/LTX-2.5 1.6M downloads and 3.1k likes. The other serious open video model, and the one with the longer track record. | Lightricks | Open | See model card | — | Self-hosted | 2026 |
| Speech to text8 |
|---|
31whisper-large-v3openai/whisper-large-v3 The model that made transcription a commodity. Still 6.2k likes and 5M downloads, and still the answer when audio cannot leave your network. | OpenAI | Open | Apache 2.0 | — | Self-hosted | 2023 |
32speaker-diarization-3.1pyannote/speaker-diarization-3.1 9.1M downloads a month for the job the transcription models do not do: working out who was speaking. | pyannote | Open | MIT | — | Self-hosted | 2023 |
33whisper-large-v3-turboopenai/whisper-large-v3-turbo 6.9M downloads a month, more than the model it distils. The default self-hosted transcription choice on constrained hardware. | OpenAI | Open | MIT | — | Self-hosted | 2024 |
34parakeet-tdt-0.6b-v3nvidia/parakeet-tdt-0.6b-v3 Very fast transcription at 0.6B. The pick when throughput per GPU matters more than the last point of accuracy. | NVIDIA | Open | CC-BY-4.0 | — | Self-hosted | 2026 |
35Qwen3-ASR-1.7BQwen/Qwen3-ASR-1.7B 3.3M downloads a month. Open-weight transcription at a size you can genuinely self-host on modest hardware. | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2026 |
36Voxtral-Mini-4B-Realtimemistralai/Voxtral-Mini-4B-Realtime-2602 Mistral's realtime speech model under Apache 2.0, at 2.2M downloads. The European option where that matters for procurement. | Mistral | Open | Apache 2.0 | — | Self-hosted | 2026 |
37GPT-Realtime-WhisperWhisper rebuilt as a streaming model. Removes the chunk-and-stitch layer every live Whisper deployment has had to write. | OpenAI | Closed | Proprietary | — | $0.017 / min | May 2026 |
38GPT-TranscribeAbout $0.27 an hour, cheaper than the legacy whisper-1 endpoint and stronger. The model most transcription pipelines should default to. | OpenAI | Closed | Proprietary | — | $0.0045 / min | May 2026 |
| Text to speech2 |
|---|
39Kokoro-82Mhexgrad/Kokoro-82M 6.8k likes and 11.5M downloads at 82 million parameters. Proof that speech synthesis does not need a large model. | hexgrad | Open | Apache 2.0 | — | Self-hosted | 2025 |
40Qwen3-TTS-CustomVoiceQwen/Qwen3-TTS-12Hz-1.7B-CustomVoice 2.6M downloads. Voice cloning on open weights, which is a capability and a governance question in equal measure. | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2026 |
| Embeddings4 |
|---|
41all-MiniLM-L6-v2sentence-transformers/all-MiniLM-L6-v2 251 million downloads a month, the most-downloaded model on Hugging Face by a factor of three. Still the default first embedder. | sentence-transformers | Open | Apache 2.0 | — | Self-hosted | 2021 |
42BAAI/bge-m3BAAI/bge-m3 37.7M downloads. Multilingual, multi-granularity, and the usual upgrade from MiniLM once retrieval quality starts to matter. | BAAI | Open | MIT | — | Self-hosted | 2024 |
43embeddinggemma-300mgoogle/embeddinggemma-300m Google's small embedder, trending hard at 2.3M downloads. Built for on-device retrieval where the vectors never leave the phone. | Google | Open | Gemma Terms of Use | — | Self-hosted | 2025 |
44nomic-embed-text-v1.5nomic-ai/nomic-embed-text-v1.5 16.2M downloads, with Matryoshka dimensions so you can trade vector size against recall without re-embedding. | Nomic | Open | Apache 2.0 | — | Self-hosted | 2024 |
| Reranking2 |
|---|
45bge-reranker-v2-m3BAAI/bge-reranker-v2-m3 18M downloads. Reranking is the cheapest quality win in a RAG pipeline and this is the model most teams reach for first. | BAAI | Open | Apache 2.0 | — | Self-hosted | 2024 |
46ms-marco-MiniLM-L6-v2cross-encoder/ms-marco-MiniLM-L6-v2 85.9M downloads a month, the second most-downloaded model on Hugging Face. Old, small, and still extremely hard to beat on cost. | cross-encoder | Open | Apache 2.0 | — | Self-hosted | 2021 |
| Time series2 |
|---|
47TimesFM 3.0google/timesfm-3.0-pytorch Zero-shot forecasting from a pretrained model, no per-series training. Trending at 272k downloads within days of release. | Google | Open | Apache 2.0 | — | Self-hosted | Sep 2026 |
48Chronos-2amazon/chronos-2 23.9M downloads a month. Foundation models for forecasting have quietly become the default, and this is the most used of them. | Amazon | Open | Apache 2.0 | — | Self-hosted | 2025 |
| Foundation backbones2 |
|---|
49bert-base-uncasedgoogle-bert/bert-base-uncased 50.7M downloads a month, seven years on. A reminder that most production NLP is not a frontier model and never was. | Google | Open | Apache 2.0 | — | Self-hosted | 2018 |
50clip-vit-base-patch32openai/clip-vit-base-patch32 20.5M downloads a month. Still the backbone under a large share of image search, moderation and zero-shot classification. | OpenAI | Open | MIT | — | Self-hosted | 2021 |