Kensink Labs
50 MODELSModel collection
MODEL COLLECTION · UPDATED 8 SEP 2026

The fifty models worth knowing about.

Frontier APIs, open weights, speech, vision, embeddings and forecasting, grouped by the job they do. There is no single number one here, because an embedding model with 251 million downloads and a 27B vision-language model are not on the same axis. What follows is the set we would expect a senior engineer to recognise, with the numbers that put each one on the list.

[HOW WE BUILT THIS]

Built from three Hugging Face API rankings pulled on 8 September 2026 (trending, thirty-day downloads, and all-time likes), the per-task trending leaders across nine pipeline tags, and vendor documentation for the closed models. Download and like counts are a popularity signal, not a quality one: a legacy backbone outranks every frontier model on volume because it is embedded in a decade of tooling. We have kept those rows in rather than quietly dropping them, because that is a true and useful thing about this industry. One entry per model line: cheaper tiers and variants of something already listed appear in the pricing and open-weight collections instead, which is how this stays at fifty rather than sprawling.

ModelVendorWeightsLicenceContext$ / 1M in · outReleased
Frontier, API only7
01
Claude Fable 5.1

Tops the Artificial Analysis Intelligence Index at roughly 66. Cache reads at $0.25, a quarter of Fable 5 and of every other Claude model.

AnthropicClosedProprietary1M$10 · $50$0.25 cached1 Sep 2026
02
Claude Mythos 5.1

Fable 5.1's specs and price, restricted to Project Glasswing participants. Scores 60.9 on Terminal-Bench 4.0 against Fable's 55.8.

AnthropicClosedProprietary, invitation only1M$10 · $50$0.25 cached1 Sep 2026
03
Claude Opus 5

Anthropic's recommended starting tier and the coding arena leader. 63.1 on the Artificial Analysis index at half Fable pricing.

AnthropicClosedProprietary1M$5 · $2524 Jul 2026
04
Gemini 3.1 Pro

Leads OCR and visual question answering on the Nanonets document leaderboard, and handles sparse tables at 94% where most pipelines break.

GoogleClosedProprietary1MPreview2026
05
GPT-5.6 Sol

The prior OpenAI flagship and still the right default for most production work at two and a half times less than Astra.

OpenAIClosedProprietary1.05M$4 · $20$0.40 cached26 Jun 2026
06
GPT-6 Astra

First model rated Critical for cyber under the Preparedness Framework. Records on computer use and terminals, fourth on a third-party intelligence index.

OpenAIClosedProprietary1.05M$10 · $50$1 cached3 Sep 2026
07
Grok 4.6

Consistently in the top five on public leaderboards. We have not run it on customer tasks, so we do not have a view worth publishing yet.

xAIClosedProprietary2026
Open-weight frontier10
08
Qwen3.8-27B
Qwen/Qwen3.8-27B

Beats Claude Opus 4.6 Max on SWE-bench Pro and OSWorld in Alibaba's table, at 27B under Apache 2.0. Second most-liked model on Hugging Face.

AlibabaOpenApache 2.0262K, 1M maxSelf-hosted5 Aug 2026
09
Kimi K3
moonshotai/Kimi-K3

The largest open-weight model shipped, at 2.8 trillion parameters, and third on the Artificial Analysis index by third-party trackers.

MoonshotOpenModified MIT1MSelf-hosted16 Jul 2026
10
DeepSeek-V4-Pro
deepseek-ai/DeepSeek-V4-Pro

The DeepSeek flagship, MIT licensed. Second on CyberGym behind GLM-5.3, and a long-standing favourite for self-hosted agent work.

DeepSeekOpenMITSelf-hosted2026
11
gpt-oss-120b
openai/gpt-oss-120b

OpenAI's open-weight release under Apache 2.0. Notable less for its scores than for existing at all, and 5.2M downloads a month.

OpenAIOpenApache 2.0Self-hosted2025
12
Qwen3.8-Flash-Next
Qwen/Qwen3.8-Flash-Next

Beats Claude Opus by 22 points on AndroidWorld device automation. Frontier scores at roughly 6B inference cost, on a 180B machine.

AlibabaOpenqwen-community-1.0262K, 1M maxSelf-hosted24 Aug 2026
13
DeepSeek-V4-Flash-0731
deepseek-ai/DeepSeek-V4-Flash-0731

MIT at 304B and 4.5M downloads a month. Widely described as the strongest cost-and-agent option among downloadable weights.

DeepSeekOpenMITSelf-hosted31 Jul 2026
14
gemma-4-31B-it
google/gemma-4-31B-it

Google's open-weight line, 8.6M downloads a month. The gemma-4 family dominates the any-to-any category on Hugging Face.

GoogleOpenGemma Terms of UseSelf-hosted2026
15
GLM-5.3-Flash
zai-org/GLM-5.3-Flash

MIT at 320B, multimodal, and four points behind its own flagship. The most permissive licence attached to a model this size.

Z.aiOpenMIT300KSelf-hosted25 Aug 2026
16
GLM-5.3
zai-org/GLM-5.3

Level with GPT-5.6 Sol on Terminal-Bench 2.1 at a seventh of the input price, and state of the art on CyberGym vulnerability discovery.

Z.aiOpenglm-5.3, bespoke1M$1.40 · $4.40$0.26 cached25 Aug 2026
17
Muse Glimmer 30B

Meta's move to Apache 2.0 and to a dense architecture, away from the Llama Community Licence and mixture-of-experts.

MetaOpenApache 2.0Self-hostedAug 2026
Small and on-device3
18
Qwen3-0.6B
Qwen/Qwen3-0.6B

21M downloads a month, more than any frontier open model. Small enough for on-device, browser and edge-runtime inference.

AlibabaOpenApache 2.0Self-hosted2025
19
Qwen3-8B
Qwen/Qwen3-8B

The workhorse size. Still 13M downloads a month a year after release, and where most self-hosted pilots start.

AlibabaOpenApache 2.0Self-hosted2025
20
Spark-X2.5-4B
XHToken/Spark-X2.5-4B

Top of the Hugging Face trending list on the day we checked, on 697 likes against 7.2k downloads. Attention arriving well ahead of adoption.

XHTokenOpenSee model cardSelf-hostedSep 2026
Vision-language and OCR3
21
Unlimited-OCR
baidu/Unlimited-OCR

2.7M downloads and 4.2k likes. Document extraction as a dedicated model rather than a prompt to a general vision-language model.

BaiduOpenSee model cardSelf-hosted2026
22
Qwen3-VL-8B-Instruct
Qwen/Qwen3-VL-8B-Instruct

14.5M downloads a month. Fits on a single accelerator, which makes it the open-weight default for document and screen work.

AlibabaOpenApache 2.0Self-hosted2026
23
PaddleOCR VL 1.5

Tops OmniDocBench at 94.37 overall, ahead of every general model on that test. A parser rather than a reasoner.

Open weightsOpenApache 2.0Self-hosted2026
Image generation5
24
FLUX.1-dev
black-forest-labs/FLUX.1-dev

The most-liked model on Hugging Face, full stop, at 14,508 likes. Note the licence: the dev weights are non-commercial.

Black Forest LabsOpenNon-commercial, devSelf-hosted2024
25
FLUX.1-schnell
black-forest-labs/FLUX.1-schnell

The commercially usable FLUX. Apache 2.0 where the dev weights are not, which is the version most products actually ship.

Black Forest LabsOpenApache 2.0Self-hosted2024
26
Z-Image-Turbo
Tongyi-MAI/Z-Image-Turbo

5.2k likes and 690k downloads. The Alibaba-adjacent entry in open image generation, and a genuine FLUX alternative.

Tongyi-MAIOpenSee model cardSelf-hosted2026
27
GPT-Image-2

Token-priced rather than per-image, so you cannot quote a cost per picture until you measure your own prompts.

OpenAIClosedProprietary$5 · $30$1.25 cachedApr 2026
28
Nano Banana Pro

Per-image pricing you can put in a quote: $0.134 at 1K or 2K, $0.24 at 4K, half that on batch. SynthID on every output.

GoogleClosedProprietary$0.134 / image2026
Video generation2
29
MiniMax-H3
MiniMaxAI/MiniMax-H3

5M downloads and 5k likes, with a whole ecosystem of turbo, quantised and LoRA derivatives already on Hugging Face.

MiniMaxAIOpenSee model cardSelf-hosted2026
30
LTX-2.5
Lightricks/LTX-2.5

1.6M downloads and 3.1k likes. The other serious open video model, and the one with the longer track record.

LightricksOpenSee model cardSelf-hosted2026
Speech to text8
31
whisper-large-v3
openai/whisper-large-v3

The model that made transcription a commodity. Still 6.2k likes and 5M downloads, and still the answer when audio cannot leave your network.

OpenAIOpenApache 2.0Self-hosted2023
32
speaker-diarization-3.1
pyannote/speaker-diarization-3.1

9.1M downloads a month for the job the transcription models do not do: working out who was speaking.

pyannoteOpenMITSelf-hosted2023
33
whisper-large-v3-turbo
openai/whisper-large-v3-turbo

6.9M downloads a month, more than the model it distils. The default self-hosted transcription choice on constrained hardware.

OpenAIOpenMITSelf-hosted2024
34
parakeet-tdt-0.6b-v3
nvidia/parakeet-tdt-0.6b-v3

Very fast transcription at 0.6B. The pick when throughput per GPU matters more than the last point of accuracy.

NVIDIAOpenCC-BY-4.0Self-hosted2026
35
Qwen3-ASR-1.7B
Qwen/Qwen3-ASR-1.7B

3.3M downloads a month. Open-weight transcription at a size you can genuinely self-host on modest hardware.

AlibabaOpenApache 2.0Self-hosted2026
36
Voxtral-Mini-4B-Realtime
mistralai/Voxtral-Mini-4B-Realtime-2602

Mistral's realtime speech model under Apache 2.0, at 2.2M downloads. The European option where that matters for procurement.

MistralOpenApache 2.0Self-hosted2026
37
GPT-Realtime-Whisper

Whisper rebuilt as a streaming model. Removes the chunk-and-stitch layer every live Whisper deployment has had to write.

OpenAIClosedProprietary$0.017 / minMay 2026
38
GPT-Transcribe

About $0.27 an hour, cheaper than the legacy whisper-1 endpoint and stronger. The model most transcription pipelines should default to.

OpenAIClosedProprietary$0.0045 / minMay 2026
Text to speech2
39
Kokoro-82M
hexgrad/Kokoro-82M

6.8k likes and 11.5M downloads at 82 million parameters. Proof that speech synthesis does not need a large model.

hexgradOpenApache 2.0Self-hosted2025
40
Qwen3-TTS-CustomVoice
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

2.6M downloads. Voice cloning on open weights, which is a capability and a governance question in equal measure.

AlibabaOpenApache 2.0Self-hosted2026
Embeddings4
41
all-MiniLM-L6-v2
sentence-transformers/all-MiniLM-L6-v2

251 million downloads a month, the most-downloaded model on Hugging Face by a factor of three. Still the default first embedder.

sentence-transformersOpenApache 2.0Self-hosted2021
42
BAAI/bge-m3
BAAI/bge-m3

37.7M downloads. Multilingual, multi-granularity, and the usual upgrade from MiniLM once retrieval quality starts to matter.

BAAIOpenMITSelf-hosted2024
43
embeddinggemma-300m
google/embeddinggemma-300m

Google's small embedder, trending hard at 2.3M downloads. Built for on-device retrieval where the vectors never leave the phone.

GoogleOpenGemma Terms of UseSelf-hosted2025
44
nomic-embed-text-v1.5
nomic-ai/nomic-embed-text-v1.5

16.2M downloads, with Matryoshka dimensions so you can trade vector size against recall without re-embedding.

NomicOpenApache 2.0Self-hosted2024
Reranking2
45
bge-reranker-v2-m3
BAAI/bge-reranker-v2-m3

18M downloads. Reranking is the cheapest quality win in a RAG pipeline and this is the model most teams reach for first.

BAAIOpenApache 2.0Self-hosted2024
46
ms-marco-MiniLM-L6-v2
cross-encoder/ms-marco-MiniLM-L6-v2

85.9M downloads a month, the second most-downloaded model on Hugging Face. Old, small, and still extremely hard to beat on cost.

cross-encoderOpenApache 2.0Self-hosted2021
Time series2
47
TimesFM 3.0
google/timesfm-3.0-pytorch

Zero-shot forecasting from a pretrained model, no per-series training. Trending at 272k downloads within days of release.

GoogleOpenApache 2.0Self-hostedSep 2026
48
Chronos-2
amazon/chronos-2

23.9M downloads a month. Foundation models for forecasting have quietly become the default, and this is the most used of them.

AmazonOpenApache 2.0Self-hosted2025
Foundation backbones2
49
bert-base-uncased
google-bert/bert-base-uncased

50.7M downloads a month, seven years on. A reminder that most production NLP is not a frontier model and never was.

GoogleOpenApache 2.0Self-hosted2018
50
clip-vit-base-patch32
openai/clip-vit-base-patch32

20.5M downloads a month. Still the backbone under a large share of image search, moderation and zero-shot classification.

OpenAIOpenMITSelf-hosted2021

Every row carries the date we last verified it. Prices are list rates at the standard tier and exclude batch discounts, long-context surcharges and regional variation. Hugging Face download and like counts are pulled from the API rather than retyped, and they measure adoption rather than quality.

[WHAT THE TABLE DOES NOT SAY]

Reading it properly.

01

Popularity is not capability, and the table shows both.

all-MiniLM-L6-v2 has 251 million downloads a month, roughly forty times any frontier model, because it is a small embedder wired into a decade of retrieval tooling. bert-base-uncased still does 50 million a month, seven years after release. Most production machine learning is not a frontier model and never was, and any list that hides that is selling something.

02

The open-weight gap has closed further than most vendor tables admit.

Qwen3.8-27B beats Claude Opus 4.6 Max on agentic coding and computer use in Alibaba's own numbers, at 27 billion parameters under Apache 2.0. GLM-5.3 runs level with GPT-5.6 Sol on Terminal-Bench 2.1 at a seventh of the price. Neither appears in a frontier vendor's comparison table, which is a positioning decision rather than a capability judgement.

03

This list has a shelf life measured in weeks.

Claude Fable 5.1 shipped on 1 September and GPT-6 Astra on 3 September, two days apart. Between them they reset three benchmark leaderboards. We date every row with when we last checked it, and we re-pull the Hugging Face numbers rather than letting them quietly rot. Treat any comparison here as a starting point for your own evals, which is the only measurement that decides anything.

[METHODOLOGY · K-FRAMEWORK]

Integrated through the
K-Framework.

Every model we integrate runs through the same operating system. Three pillars, sixteen layers, one Compound Growth Loop. The methodology that keeps AI work from rotting after the first ship.

Read the K-Framework
01

Foundations

Direct API integration with the model. No LangChain, no orchestration vendor, no agent framework built on quicksand. Typed contracts, the same way we wire up Postgres.

02

Amplification

An eval suite built from your real tasks gates every prompt and model change. Quality is measured before it ships, not vibed in a demo.

03

Judgment

Governance, audit, and oversight wired in from day one. Who called what, with which prompt version, at what cost. Your auditors get answers, not screenshots.

[OBSERVABILITY]

Observability your team can read.

A model in production without observability is roulette. We instrument every integration so engineering and finance can see the same numbers, and so a regression at 3am surfaces before a customer opens a ticket.

Instrumented

Cost per call

Tokens in, tokens out, dollars spent. Sliced by feature, tenant, and route. Budgets enforced where it matters.

Instrumented

Latency p50 / p95 / p99

Real distributions, not averages. We know which routes are slow, and why.

Instrumented

Eval pass rates

The same eval suite that gates a release runs continuously in production. A regression on real traffic surfaces fast.

Instrumented

Prompt + completion logs

PII scrubbed at the proxy, shipped to your SIEM. Retention controls match your compliance window.

Dashboards your team owns, not ours. At handoff you get the queries, the alerts, and the runbook. We are not in the path to read your metrics.

[COMMON QUESTIONS]

Questions we are getting asked.

Why is there no single ranked list from one to fifty?
Because it would be fiction. Ranking an embedding model against a video generator against a frontier reasoning model requires a common axis that does not exist. Grouping by job is the honest structure, and it is also the one that matches how a routing decision actually gets made.
How did you decide what counts as making a name?
Three Hugging Face signals for open weights (trending score, thirty-day downloads, all-time likes) and vendor documentation plus leaderboard presence for closed models. Where a model is trending on very few downloads, as several were on the day we checked, we say so in the row rather than implying adoption that has not happened yet.
Why include models you have not written about?
Because a list that only contains our own pages would be marketing rather than a reference. Grok 4.6 is on here with a row that says plainly we have not run it on customer tasks and do not have a view worth publishing. That is more useful than leaving it out.
How often is this updated?
The Hugging Face figures are refreshed from the API rather than retyped, and every row carries the date we last verified it. Given two frontier releases landed within a week of each other, assume anything more than a month old needs re-checking before you make a decision on it.
Share
View .md
DIRECT INTEGRATION · NO FRAMEWORK

Picking a model
is a routing decision.

We build behind a vendor-neutral abstraction and route by task difficulty at runtime, with an eval suite that decides rather than a launch post. Eval suite at handoff, full source ownership.