---
title: "Top 50 AI Models, September 2026"
description: "The fifty models worth knowing in September 2026, grouped by job: frontier APIs, open weights, vision, speech and embeddings, with licences and prices."
source: "https://www.kensink.com/models/top-50/"
canonical: "https://www.kensink.com/models/top-50/"
---
★ 50 MODELS Model collection

MODEL COLLECTION · UPDATED 8 SEP 2026

# The fifty models worth knowing about.

Frontier APIs, open weights, speech, vision, embeddings and forecasting, grouped by the job they do. There is no single number one here, because an embedding model with 251 million downloads and a 27B vision-language model are not on the same axis. What follows is the set we would expect a senior engineer to recognise, with the numbers that put each one on the list.

[Talk to us about model choice →](https://www.kensink.com/contact) [Browse by family](https://www.kensink.com/models/)

\[HOW WE BUILT THIS\]

Built from three Hugging Face API rankings pulled on 8 September 2026 (trending, thirty-day downloads, and all-time likes), the per-task trending leaders across nine pipeline tags, and vendor documentation for the closed models. Download and like counts are a popularity signal, not a quality one: a legacy backbone outranks every frontier model on volume because it is embedded in a decade of tooling. We have kept those rows in rather than quietly dropping them, because that is a true and useful thing about this industry. One entry per model line: cheaper tiers and variants of something already listed appear in the pricing and open-weight collections instead, which is how this stays at fifty rather than sprawling.

| Model | Vendor | Weights | Licence | Context | $ / 1M in · out | Released |
| --- | --- | --- | --- | --- | --- | --- |
| Frontier, API only7 |
| 
01

[Claude Fable 5.1](https://www.kensink.com/models/claude/fable-5-1/)

Tops the Artificial Analysis Intelligence Index at roughly 66. Cache reads at $0.25, a quarter of Fable 5 and of every other Claude model.

 | Anthropic | Closed | Proprietary | 1M | $10 · $50 $0.25 cached | 1 Sep 2026 |
| 

02

[Claude Mythos 5.1](https://www.kensink.com/models/claude/fable-5-1/)

Fable 5.1's specs and price, restricted to Project Glasswing participants. Scores 60.9 on Terminal-Bench 4.0 against Fable's 55.8.

 | Anthropic | Closed | Proprietary, invitation only | 1M | $10 · $50 $0.25 cached | 1 Sep 2026 |
| 

03

[Claude Opus 5](https://www.kensink.com/models/claude/opus-5/)

Anthropic's recommended starting tier and the coding arena leader. 63.1 on the Artificial Analysis index at half Fable pricing.

 | Anthropic | Closed | Proprietary | 1M | $5 · $25 | 24 Jul 2026 |
| 

04

[Gemini 3.1 Pro](https://www.kensink.com/models/gemini/)

Leads OCR and visual question answering on the Nanonets document leaderboard, and handles sparse tables at 94% where most pipelines break.

 | Google | Closed | Proprietary | 1M | Preview | 2026 |
| 

05

[GPT-5.6 Sol](https://www.kensink.com/models/openai-gpt/gpt-5-6/)

The prior OpenAI flagship and still the right default for most production work at two and a half times less than Astra.

 | OpenAI | Closed | Proprietary | 1.05M | $4 · $20 $0.40 cached | 26 Jun 2026 |
| 

06

[GPT-6 Astra](https://www.kensink.com/models/openai-gpt/gpt-6-astra/)

First model rated Critical for cyber under the Preparedness Framework. Records on computer use and terminals, fourth on a third-party intelligence index.

 | OpenAI | Closed | Proprietary | 1.05M | $10 · $50 $1 cached | 3 Sep 2026 |
| 

07

Grok 4.6

Consistently in the top five on public leaderboards. We have not run it on customer tasks, so we do not have a view worth publishing yet.

 | xAI | Closed | Proprietary | — | — | 2026 |
| Open-weight frontier10 |
| 

08

[Qwen3.8-27B](https://www.kensink.com/models/qwen/qwen3-8-27b/)

Qwen/Qwen3.8-27B

Beats Claude Opus 4.6 Max on SWE-bench Pro and OSWorld in Alibaba's table, at 27B under Apache 2.0. Second most-liked model on Hugging Face.

 | Alibaba | Open | Apache 2.0 | 262K, 1M max | Self-hosted | 5 Aug 2026 |
| 

09

[Kimi K3](https://www.kensink.com/models/kimi/k3/)

moonshotai/Kimi-K3

The largest open-weight model shipped, at 2.8 trillion parameters, and third on the Artificial Analysis index by third-party trackers.

 | Moonshot | Open | Modified MIT | 1M | Self-hosted | 16 Jul 2026 |
| 

10

[DeepSeek-V4-Pro](https://www.kensink.com/models/deepseek/)

deepseek-ai/DeepSeek-V4-Pro

The DeepSeek flagship, MIT licensed. Second on CyberGym behind GLM-5.3, and a long-standing favourite for self-hosted agent work.

 | DeepSeek | Open | MIT | — | Self-hosted | 2026 |
| 

11

gpt-oss-120b

openai/gpt-oss-120b

OpenAI's open-weight release under Apache 2.0. Notable less for its scores than for existing at all, and 5.2M downloads a month.

 | OpenAI | Open | Apache 2.0 | — | Self-hosted | 2025 |
| 

12

[Qwen3.8-Flash-Next](https://www.kensink.com/models/qwen/qwen3-8-flash-next/)

Qwen/Qwen3.8-Flash-Next

Beats Claude Opus by 22 points on AndroidWorld device automation. Frontier scores at roughly 6B inference cost, on a 180B machine.

 | Alibaba | Open | qwen-community-1.0 | 262K, 1M max | Self-hosted | 24 Aug 2026 |
| 

13

[DeepSeek-V4-Flash-0731](https://www.kensink.com/models/deepseek/)

deepseek-ai/DeepSeek-V4-Flash-0731

MIT at 304B and 4.5M downloads a month. Widely described as the strongest cost-and-agent option among downloadable weights.

 | DeepSeek | Open | MIT | — | Self-hosted | 31 Jul 2026 |
| 

14

gemma-4-31B-it

google/gemma-4-31B-it

Google's open-weight line, 8.6M downloads a month. The gemma-4 family dominates the any-to-any category on Hugging Face.

 | Google | Open | Gemma Terms of Use | — | Self-hosted | 2026 |
| 

15

[GLM-5.3-Flash](https://www.kensink.com/models/glm/glm-5-3-flash/)

zai-org/GLM-5.3-Flash

MIT at 320B, multimodal, and four points behind its own flagship. The most permissive licence attached to a model this size.

 | Z.ai | Open | MIT | 300K | Self-hosted | 25 Aug 2026 |
| 

16

[GLM-5.3](https://www.kensink.com/models/glm/glm-5-3/)

zai-org/GLM-5.3

Level with GPT-5.6 Sol on Terminal-Bench 2.1 at a seventh of the input price, and state of the art on CyberGym vulnerability discovery.

 | Z.ai | Open | glm-5.3, bespoke | 1M | $1.40 · $4.40 $0.26 cached | 25 Aug 2026 |
| 

17

[Muse Glimmer 30B](https://www.kensink.com/models/meta/muse-glimmer/)

Meta's move to Apache 2.0 and to a dense architecture, away from the Llama Community Licence and mixture-of-experts.

 | Meta | Open | Apache 2.0 | — | Self-hosted | Aug 2026 |
| Small and on-device3 |
| 

18

[Qwen3-0.6B](https://www.kensink.com/models/qwen/)

Qwen/Qwen3-0.6B

21M downloads a month, more than any frontier open model. Small enough for on-device, browser and edge-runtime inference.

 | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2025 |
| 

19

[Qwen3-8B](https://www.kensink.com/models/qwen/)

Qwen/Qwen3-8B

The workhorse size. Still 13M downloads a month a year after release, and where most self-hosted pilots start.

 | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2025 |
| 

20

Spark-X2.5-4B

XHToken/Spark-X2.5-4B

Top of the Hugging Face trending list on the day we checked, on 697 likes against 7.2k downloads. Attention arriving well ahead of adoption.

 | XHToken | Open | See model card | — | Self-hosted | Sep 2026 |
| Vision-language and OCR3 |
| 

21

Unlimited-OCR

baidu/Unlimited-OCR

2.7M downloads and 4.2k likes. Document extraction as a dedicated model rather than a prompt to a general vision-language model.

 | Baidu | Open | See model card | — | Self-hosted | 2026 |
| 

22

[Qwen3-VL-8B-Instruct](https://www.kensink.com/models/qwen/)

Qwen/Qwen3-VL-8B-Instruct

14.5M downloads a month. Fits on a single accelerator, which makes it the open-weight default for document and screen work.

 | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2026 |
| 

23

[PaddleOCR VL 1.5](https://www.kensink.com/models/vision/)

Tops OmniDocBench at 94.37 overall, ahead of every general model on that test. A parser rather than a reasoner.

 | Open weights | Open | Apache 2.0 | — | Self-hosted | 2026 |
| Image generation5 |
| 

24

FLUX.1-dev

black-forest-labs/FLUX.1-dev

The most-liked model on Hugging Face, full stop, at 14,508 likes. Note the licence: the dev weights are non-commercial.

 | Black Forest Labs | Open | Non-commercial, dev | — | Self-hosted | 2024 |
| 

25

FLUX.1-schnell

black-forest-labs/FLUX.1-schnell

The commercially usable FLUX. Apache 2.0 where the dev weights are not, which is the version most products actually ship.

 | Black Forest Labs | Open | Apache 2.0 | — | Self-hosted | 2024 |
| 

26

Z-Image-Turbo

Tongyi-MAI/Z-Image-Turbo

5.2k likes and 690k downloads. The Alibaba-adjacent entry in open image generation, and a genuine FLUX alternative.

 | Tongyi-MAI | Open | See model card | — | Self-hosted | 2026 |
| 

27

[GPT-Image-2](https://www.kensink.com/models/vision/gpt-image-2/)

Token-priced rather than per-image, so you cannot quote a cost per picture until you measure your own prompts.

 | OpenAI | Closed | Proprietary | — | $5 · $30 $1.25 cached | Apr 2026 |
| 

28

[Nano Banana Pro](https://www.kensink.com/models/vision/gemini-3-pro-image/)

Per-image pricing you can put in a quote: $0.134 at 1K or 2K, $0.24 at 4K, half that on batch. SynthID on every output.

 | Google | Closed | Proprietary | — | $0.134 / image | 2026 |
| Video generation2 |
| 

29

MiniMax-H3

MiniMaxAI/MiniMax-H3

5M downloads and 5k likes, with a whole ecosystem of turbo, quantised and LoRA derivatives already on Hugging Face.

 | MiniMaxAI | Open | See model card | — | Self-hosted | 2026 |
| 

30

LTX-2.5

Lightricks/LTX-2.5

1.6M downloads and 3.1k likes. The other serious open video model, and the one with the longer track record.

 | Lightricks | Open | See model card | — | Self-hosted | 2026 |
| Speech to text8 |
| 

31

[whisper-large-v3](https://www.kensink.com/models/whisper-speech/)

openai/whisper-large-v3

The model that made transcription a commodity. Still 6.2k likes and 5M downloads, and still the answer when audio cannot leave your network.

 | OpenAI | Open | Apache 2.0 | — | Self-hosted | 2023 |
| 

32

[speaker-diarization-3.1](https://www.kensink.com/models/whisper-speech/)

pyannote/speaker-diarization-3.1

9.1M downloads a month for the job the transcription models do not do: working out who was speaking.

 | pyannote | Open | MIT | — | Self-hosted | 2023 |
| 

33

[whisper-large-v3-turbo](https://www.kensink.com/models/whisper-speech/)

openai/whisper-large-v3-turbo

6.9M downloads a month, more than the model it distils. The default self-hosted transcription choice on constrained hardware.

 | OpenAI | Open | MIT | — | Self-hosted | 2024 |
| 

34

parakeet-tdt-0.6b-v3

nvidia/parakeet-tdt-0.6b-v3

Very fast transcription at 0.6B. The pick when throughput per GPU matters more than the last point of accuracy.

 | NVIDIA | Open | CC-BY-4.0 | — | Self-hosted | 2026 |
| 

35

[Qwen3-ASR-1.7B](https://www.kensink.com/models/qwen/)

Qwen/Qwen3-ASR-1.7B

3.3M downloads a month. Open-weight transcription at a size you can genuinely self-host on modest hardware.

 | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2026 |
| 

36

Voxtral-Mini-4B-Realtime

mistralai/Voxtral-Mini-4B-Realtime-2602

Mistral's realtime speech model under Apache 2.0, at 2.2M downloads. The European option where that matters for procurement.

 | Mistral | Open | Apache 2.0 | — | Self-hosted | 2026 |
| 

37

[GPT-Realtime-Whisper](https://www.kensink.com/models/whisper-speech/gpt-realtime-whisper/)

Whisper rebuilt as a streaming model. Removes the chunk-and-stitch layer every live Whisper deployment has had to write.

 | OpenAI | Closed | Proprietary | — | $0.017 / min | May 2026 |
| 

38

[GPT-Transcribe](https://www.kensink.com/models/whisper-speech/gpt-transcribe/)

About $0.27 an hour, cheaper than the legacy whisper-1 endpoint and stronger. The model most transcription pipelines should default to.

 | OpenAI | Closed | Proprietary | — | $0.0045 / min | May 2026 |
| Text to speech2 |
| 

39

[Kokoro-82M](https://www.kensink.com/models/whisper-speech/)

hexgrad/Kokoro-82M

6.8k likes and 11.5M downloads at 82 million parameters. Proof that speech synthesis does not need a large model.

 | hexgrad | Open | Apache 2.0 | — | Self-hosted | 2025 |
| 

40

[Qwen3-TTS-CustomVoice](https://www.kensink.com/models/qwen/)

Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

2.6M downloads. Voice cloning on open weights, which is a capability and a governance question in equal measure.

 | Alibaba | Open | Apache 2.0 | — | Self-hosted | 2026 |
| Embeddings4 |
| 

41

[all-MiniLM-L6-v2](https://www.kensink.com/models/embeddings/)

sentence-transformers/all-MiniLM-L6-v2

251 million downloads a month, the most-downloaded model on Hugging Face by a factor of three. Still the default first embedder.

 | sentence-transformers | Open | Apache 2.0 | — | Self-hosted | 2021 |
| 

42

[BAAI/bge-m3](https://www.kensink.com/models/embeddings/)

BAAI/bge-m3

37.7M downloads. Multilingual, multi-granularity, and the usual upgrade from MiniLM once retrieval quality starts to matter.

 | BAAI | Open | MIT | — | Self-hosted | 2024 |
| 

43

[embeddinggemma-300m](https://www.kensink.com/models/embeddings/)

google/embeddinggemma-300m

Google's small embedder, trending hard at 2.3M downloads. Built for on-device retrieval where the vectors never leave the phone.

 | Google | Open | Gemma Terms of Use | — | Self-hosted | 2025 |
| 

44

[nomic-embed-text-v1.5](https://www.kensink.com/models/embeddings/)

nomic-ai/nomic-embed-text-v1.5

16.2M downloads, with Matryoshka dimensions so you can trade vector size against recall without re-embedding.

 | Nomic | Open | Apache 2.0 | — | Self-hosted | 2024 |
| Reranking2 |
| 

45

[bge-reranker-v2-m3](https://www.kensink.com/models/embeddings/)

BAAI/bge-reranker-v2-m3

18M downloads. Reranking is the cheapest quality win in a RAG pipeline and this is the model most teams reach for first.

 | BAAI | Open | Apache 2.0 | — | Self-hosted | 2024 |
| 

46

[ms-marco-MiniLM-L6-v2](https://www.kensink.com/models/embeddings/)

cross-encoder/ms-marco-MiniLM-L6-v2

85.9M downloads a month, the second most-downloaded model on Hugging Face. Old, small, and still extremely hard to beat on cost.

 | cross-encoder | Open | Apache 2.0 | — | Self-hosted | 2021 |
| Time series2 |
| 

47

TimesFM 3.0

google/timesfm-3.0-pytorch

Zero-shot forecasting from a pretrained model, no per-series training. Trending at 272k downloads within days of release.

 | Google | Open | Apache 2.0 | — | Self-hosted | Sep 2026 |
| 

48

Chronos-2

amazon/chronos-2

23.9M downloads a month. Foundation models for forecasting have quietly become the default, and this is the most used of them.

 | Amazon | Open | Apache 2.0 | — | Self-hosted | 2025 |
| Foundation backbones2 |
| 

49

bert-base-uncased

google-bert/bert-base-uncased

50.7M downloads a month, seven years on. A reminder that most production NLP is not a frontier model and never was.

 | Google | Open | Apache 2.0 | — | Self-hosted | 2018 |
| 

50

clip-vit-base-patch32

openai/clip-vit-base-patch32

20.5M downloads a month. Still the backbone under a large share of image search, moderation and zero-shot classification.

 | OpenAI | Open | MIT | — | Self-hosted | 2021 |

Every row carries the date we last verified it. Prices are list rates at the standard tier and exclude batch discounts, long-context surcharges and regional variation. Hugging Face download and like counts are pulled from the API rather than retyped, and they measure adoption rather than quality.

\[WHAT THE TABLE DOES NOT SAY\]

## Reading it properly.

01

### Popularity is not capability, and the table shows both.

all-MiniLM-L6-v2 has 251 million downloads a month, roughly forty times any frontier model, because it is a small embedder wired into a decade of retrieval tooling. bert-base-uncased still does 50 million a month, seven years after release. Most production machine learning is not a frontier model and never was, and any list that hides that is selling something.

02

### The open-weight gap has closed further than most vendor tables admit.

Qwen3.8-27B beats Claude Opus 4.6 Max on agentic coding and computer use in Alibaba's own numbers, at 27 billion parameters under Apache 2.0. GLM-5.3 runs level with GPT-5.6 Sol on Terminal-Bench 2.1 at a seventh of the price. Neither appears in a frontier vendor's comparison table, which is a positioning decision rather than a capability judgement.

03

### This list has a shelf life measured in weeks.

Claude Fable 5.1 shipped on 1 September and GPT-6 Astra on 3 September, two days apart. Between them they reset three benchmark leaderboards. We date every row with when we last checked it, and we re-pull the Hugging Face numbers rather than letting them quietly rot. Treat any comparison here as a starting point for your own evals, which is the only measurement that decides anything.

\[METHODOLOGY · K-FRAMEWORK\]

## Integrated through the  
K-Framework.

Every model we integrate runs through the same operating system. Three pillars, sixteen layers, one Compound Growth Loop. The methodology that keeps AI work from rotting after the first ship.

[Read the K-Framework](https://www.kensink.com/k-framework)

01

### Foundations

Direct API integration with the model. No LangChain, no orchestration vendor, no agent framework built on quicksand. Typed contracts, the same way we wire up Postgres.

02

### Amplification

An eval suite built from your real tasks gates every prompt and model change. Quality is measured before it ships, not vibed in a demo.

03

### Judgment

Governance, audit, and oversight wired in from day one. Who called what, with which prompt version, at what cost. Your auditors get answers, not screenshots.

\[OBSERVABILITY\]

## Observability your team can read.

A model in production without observability is roulette. We instrument every integration so engineering and finance can see the same numbers, and so a regression at 3am surfaces before a customer opens a ticket.

Instrumented

### Cost per call

Tokens in, tokens out, dollars spent. Sliced by feature, tenant, and route. Budgets enforced where it matters.

Instrumented

### Latency p50 / p95 / p99

Real distributions, not averages. We know which routes are slow, and why.

Instrumented

### Eval pass rates

The same eval suite that gates a release runs continuously in production. A regression on real traffic surfaces fast.

Instrumented

### Prompt + completion logs

PII scrubbed at the proxy, shipped to your SIEM. Retention controls match your compliance window.

Dashboards your team owns, not ours. At handoff you get the queries, the alerts, and the runbook. We are not in the path to read your metrics.

\[COMMON QUESTIONS\]

## Questions we are getting asked.

Why is there no single ranked list from one to fifty?

Because it would be fiction. Ranking an embedding model against a video generator against a frontier reasoning model requires a common axis that does not exist. Grouping by job is the honest structure, and it is also the one that matches how a routing decision actually gets made.

How did you decide what counts as making a name?

Three Hugging Face signals for open weights (trending score, thirty-day downloads, all-time likes) and vendor documentation plus leaderboard presence for closed models. Where a model is trending on very few downloads, as several were on the day we checked, we say so in the row rather than implying adoption that has not happened yet.

Why include models you have not written about?

Because a list that only contains our own pages would be marketing rather than a reference. Grok 4.6 is on here with a row that says plainly we have not run it on customer tasks and do not have a view worth publishing. That is more useful than leaving it out.

How often is this updated?

The Hugging Face figures are refreshed from the API rather than retyped, and every row carries the date we last verified it. Given two frontier releases landed within a week of each other, assume anything more than a month old needs re-checking before you make a decision on it.

Share[](https://twitter.com/intent/tweet?url=https%3A%2F%2Fwww.kensink.com%2Fmodels%2Ftop-50%2F&text=The%20top%2050)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fwww.kensink.com%2Fmodels%2Ftop-50%2F)

[View .md](https://www.kensink.com/models/top-50.md)

\[OTHER COLLECTIONS\]

## Same models, different question.

[

COLLECTION

Open weights

View

](https://www.kensink.com/models/open-weight-2026/)[

COLLECTION

Coding

View

](https://www.kensink.com/models/coding/)[

COLLECTION

Long context

View

](https://www.kensink.com/models/long-context/)[

COLLECTION

Pricing

View

](https://www.kensink.com/models/pricing/)

DIRECT INTEGRATION · NO FRAMEWORK

## Picking a model  
is a routing decision.

We build behind a vendor-neutral abstraction and route by task difficulty at runtime, with an eval suite that decides rather than a launch post. Eval suite at handoff, full source ownership.

[Start a conversation →](https://www.kensink.com/contact) [Browse model families](https://www.kensink.com/models/)
