Kensink Labs
45 MODELSModel collection
MODEL COLLECTION · OPEN WEIGHTS · 2026

Open-weight models, sorted by what the licence lets you do.

Every model here can be downloaded and run on your own hardware. That is where the similarity ends, because open weights and a permissive licence are different things and the difference is what decides whether you can ship. MIT at 320 billion parameters and a bespoke vendor licence at 753 billion are not the same offer, and the table says which is which.

[HOW WE BUILT THIS]

Selected as every downloadable model in our index, with licences taken from the model card and the Hugging Face API rather than from a launch post. Where a card does not name a standard licence we write what it says rather than guessing at an equivalent. Download and like counts are from the Hugging Face API on 8 September 2026 and are a proxy for adoption, not for quality.

ModelVendorLicenceParamsContextLikesDownloads / mo
Open-weight frontier10
Qwen3.8-27B
Qwen/Qwen3.8-27B

Beats Claude Opus 4.6 Max on SWE-bench Pro and OSWorld in Alibaba's table, at 27B under Apache 2.0. Second most-liked model on Hugging Face.

AlibabaApache 2.027.8B27.8B dense active262K, 1M max14.2k6.4M
Kimi K3
moonshotai/Kimi-K3

The largest open-weight model shipped, at 2.8 trillion parameters, and third on the Artificial Analysis index by third-party trackers.

MoonshotModified MIT2.8T1M11.2k2.4M
DeepSeek-V4-Pro
deepseek-ai/DeepSeek-V4-Pro

The DeepSeek flagship, MIT licensed. Second on CyberGym behind GLM-5.3, and a long-standing favourite for self-hosted agent work.

DeepSeekMIT5.5k625.9k
gpt-oss-120b
openai/gpt-oss-120b

OpenAI's open-weight release under Apache 2.0. Notable less for its scores than for existing at all, and 5.2M downloads a month.

OpenAIApache 2.0120B5.2k5.2M
Qwen3.8-Flash-Next
Qwen/Qwen3.8-Flash-Next

Beats Claude Opus by 22 points on AndroidWorld device automation. Frontier scores at roughly 6B inference cost, on a 180B machine.

Alibabaqwen-community-1.0180B6B active262K, 1M max5.0k474.7k
DeepSeek-V4-Flash-0731
deepseek-ai/DeepSeek-V4-Flash-0731

MIT at 304B and 4.5M downloads a month. Widely described as the strongest cost-and-agent option among downloadable weights.

DeepSeekMIT304B3.9k4.5M
gemma-4-31B-it
google/gemma-4-31B-it

Google's open-weight line, 8.6M downloads a month. The gemma-4 family dominates the any-to-any category on Hugging Face.

GoogleGemma Terms of Use31B3.7k8.6M
GLM-5.3-Flash
zai-org/GLM-5.3-Flash

MIT at 320B, multimodal, and four points behind its own flagship. The most permissive licence attached to a model this size.

Z.aiMIT320B18B active300K2.1k784.0k
GLM-5.3
zai-org/GLM-5.3

Level with GPT-5.6 Sol on Terminal-Bench 2.1 at a seventh of the input price, and state of the art on CyberGym vulnerability discovery.

Z.aiglm-5.3, bespoke753B8 experts / token active1M1.7k442.1k
Muse Glimmer 30B

Meta's move to Apache 2.0 and to a dense architecture, away from the Llama Community Licence and mixture-of-experts.

MetaApache 2.030B
Small and on-device4
Qwen3-0.6B
Qwen/Qwen3-0.6B

21M downloads a month, more than any frontier open model. Small enough for on-device, browser and edge-runtime inference.

AlibabaApache 2.00.6B1.6k21.3M
Qwen3-8B
Qwen/Qwen3-8B

The workhorse size. Still 13M downloads a month a year after release, and where most self-hosted pilots start.

AlibabaApache 2.08B1.4k13.0M
Spark-X2.5-4B
XHToken/Spark-X2.5-4B

Top of the Hugging Face trending list on the day we checked, on 697 likes against 7.2k downloads. Attention arriving well ahead of adoption.

XHTokenSee model card4B6997.2k
MiniCPM5-2B
openbmb/MiniCPM5-2B

Trending on 159 likes with 13 downloads, which is what a release looks like in its first hours. Worth watching rather than deploying.

OpenBMBSee model card2B16713
Vision-language and OCR4
Unlimited-OCR
baidu/Unlimited-OCR

2.7M downloads and 4.2k likes. Document extraction as a dedicated model rather than a prompt to a general vision-language model.

BaiduSee model card4.2k2.7M
Qwen3-VL-8B-Instruct
Qwen/Qwen3-VL-8B-Instruct

14.5M downloads a month. Fits on a single accelerator, which makes it the open-weight default for document and screen work.

AlibabaApache 2.08B1.1k14.5M
DeepSeek-V4-Flash-Vision-Exp
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp

DeepSeek's experimental vision checkpoint at 305B, trending within days of release. Experimental is in the name, so treat it that way.

DeepSeekMIT305B789251.6k
PaddleOCR VL 1.5

Tops OmniDocBench at 94.37 overall, ahead of every general model on that test. A parser rather than a reasoner.

Open weightsApache 2.0
Image generation4
FLUX.1-dev
black-forest-labs/FLUX.1-dev

The most-liked model on Hugging Face, full stop, at 14,508 likes. Note the licence: the dev weights are non-commercial.

Black Forest LabsNon-commercial, dev14.5k774.3k
FLUX.1-schnell
black-forest-labs/FLUX.1-schnell

The commercially usable FLUX. Apache 2.0 where the dev weights are not, which is the version most products actually ship.

Black Forest LabsApache 2.05.7k717.8k
Z-Image-Turbo
Tongyi-MAI/Z-Image-Turbo

5.2k likes and 690k downloads. The Alibaba-adjacent entry in open image generation, and a genuine FLUX alternative.

Tongyi-MAISee model card5.2k689.5k
Krea-2-Turbo
krea/Krea-2-Turbo

Top of the Hugging Face trending list for text-to-image on the day we checked, ahead of FLUX on momentum if not on total likes.

KreaSee model card1.1k76.2k
Video generation3
MiniMax-H3
MiniMaxAI/MiniMax-H3

5M downloads and 5k likes, with a whole ecosystem of turbo, quantised and LoRA derivatives already on Hugging Face.

MiniMaxAISee model card33B5.0k5.0M
LTX-2.5
Lightricks/LTX-2.5

1.6M downloads and 3.1k likes. The other serious open video model, and the one with the longer track record.

LightricksSee model card3.1k1.6M
Wan2.2-I2V-A14B
Wan-AI/Wan2.2-I2V-A14B

Image to video at 14B under Apache 2.0, with a large ComfyUI following. The permissively licensed option in open video.

AlibabaApache 2.014B79410.2k
Speech to text6
whisper-large-v3
openai/whisper-large-v3

The model that made transcription a commodity. Still 6.2k likes and 5M downloads, and still the answer when audio cannot leave your network.

OpenAIApache 2.06.2k5.0M
speaker-diarization-3.1
pyannote/speaker-diarization-3.1

9.1M downloads a month for the job the transcription models do not do: working out who was speaking.

pyannoteMIT3.4k9.1M
whisper-large-v3-turbo
openai/whisper-large-v3-turbo

6.9M downloads a month, more than the model it distils. The default self-hosted transcription choice on constrained hardware.

OpenAIMIT3.3k6.9M
parakeet-tdt-0.6b-v3
nvidia/parakeet-tdt-0.6b-v3

Very fast transcription at 0.6B. The pick when throughput per GPU matters more than the last point of accuracy.

NVIDIACC-BY-4.00.6B1.1k745.8k
Qwen3-ASR-1.7B
Qwen/Qwen3-ASR-1.7B

3.3M downloads a month. Open-weight transcription at a size you can genuinely self-host on modest hardware.

AlibabaApache 2.01.7B1.1k3.3M
Voxtral-Mini-4B-Realtime
mistralai/Voxtral-Mini-4B-Realtime-2602

Mistral's realtime speech model under Apache 2.0, at 2.2M downloads. The European option where that matters for procurement.

MistralApache 2.04B9732.2M
Text to speech3
Kokoro-82M
hexgrad/Kokoro-82M

6.8k likes and 11.5M downloads at 82 million parameters. Proof that speech synthesis does not need a large model.

hexgradApache 2.082M6.8k11.5M
Qwen3-TTS-CustomVoice
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

2.6M downloads. Voice cloning on open weights, which is a capability and a governance question in equal measure.

AlibabaApache 2.01.7B1.9k2.6M
VoxCPM2
openbmb/VoxCPM2

1.6k likes on 399k downloads, and rising fast. One of the stronger recent open synthesis releases.

OpenBMBSee model card1.6k399.4k
Embeddings4
all-MiniLM-L6-v2
sentence-transformers/all-MiniLM-L6-v2

251 million downloads a month, the most-downloaded model on Hugging Face by a factor of three. Still the default first embedder.

sentence-transformersApache 2.022.7M5.6k251.4M
BAAI/bge-m3
BAAI/bge-m3

37.7M downloads. Multilingual, multi-granularity, and the usual upgrade from MiniLM once retrieval quality starts to matter.

BAAIMIT3.5k37.7M
embeddinggemma-300m
google/embeddinggemma-300m

Google's small embedder, trending hard at 2.3M downloads. Built for on-device retrieval where the vectors never leave the phone.

GoogleGemma Terms of Use300M1.9k2.3M
nomic-embed-text-v1.5
nomic-ai/nomic-embed-text-v1.5

16.2M downloads, with Matryoshka dimensions so you can trade vector size against recall without re-embedding.

NomicApache 2.090216.2M
Reranking3
bge-reranker-v2-m3
BAAI/bge-reranker-v2-m3

18M downloads. Reranking is the cheapest quality win in a RAG pipeline and this is the model most teams reach for first.

BAAIApache 2.01.2k18.0M
ms-marco-MiniLM-L6-v2
cross-encoder/ms-marco-MiniLM-L6-v2

85.9M downloads a month, the second most-downloaded model on Hugging Face. Old, small, and still extremely hard to beat on cost.

cross-encoderApache 2.031285.9M
Qwen3-Reranker-4B
Qwen/Qwen3-Reranker-4B

The larger reranker option, with a 0.6B sibling for latency-bound paths. Sits after retrieval and before generation.

AlibabaApache 2.04B1552.4M
Time series2
TimesFM 3.0
google/timesfm-3.0-pytorch

Zero-shot forecasting from a pretrained model, no per-series training. Trending at 272k downloads within days of release.

GoogleApache 2.00.3B565271.7k
Chronos-2
amazon/chronos-2

23.9M downloads a month. Foundation models for forecasting have quietly become the default, and this is the most used of them.

AmazonApache 2.043323.9M
Foundation backbones2
bert-base-uncased
google-bert/bert-base-uncased

50.7M downloads a month, seven years on. A reminder that most production NLP is not a frontier model and never was.

GoogleApache 2.0110M3.0k50.7M
clip-vit-base-patch32
openai/clip-vit-base-patch32

20.5M downloads a month. Still the backbone under a large share of image search, moderation and zero-shot classification.

OpenAIMIT1.2k20.5M

Every row carries the date we last verified it. Prices are list rates at the standard tier and exclude batch discounts, long-context surcharges and regional variation. Hugging Face download and like counts are pulled from the API rather than retyped, and they measure adoption rather than quality.

[WHAT THE TABLE DOES NOT SAY]

Reading it properly.

01

Read the licence per checkpoint, not per family.

Qwen3.8-27B is Apache 2.0 and Qwen3.8-Flash-Next, released three weeks later in the same family, is qwen-community-1.0. GLM-5.3 carries a bespoke Z.ai licence while GLM-5.3-Flash is MIT. FLUX.1-dev is non-commercial and FLUX.1-schnell is Apache 2.0. Assuming a family licence is how a compliance problem gets built into a product.

02

MIT at 320 billion parameters is the year's most underrated release.

GLM-5.3-Flash is the most permissively licensed model of its size we are aware of: no field-of-use restriction, no user threshold, no acceptable-use policy that can be revised after you have shipped. It also has more downloads and more likes than its larger, more restrictively licensed flagship, which tells you what practitioners actually optimise for.

03

Self-hosting has to beat the hosted tier, not the frontier.

The comparison teams make is against GPT-6 Astra at $10 per million. The comparison that matters is against Qwen3.8-Flash at $0.14 or GLM-5.3 at $1.40, because those are the same models served by people whose job it is. Once you count GPU hours, on-call and the engineer maintaining the serving stack, self-hosting usually has to win on privacy, latency or genuine sustained volume rather than on cost.

04

You inherit the safety layer along with the weights.

No vendor classifier between prompt and model, no refusal path, no abuse monitoring. Input filtering, output review, rate limiting and audit logging all become yours to build. This is the most common reason a self-hosted pilot stalls before production, and it is worth scoping at the start rather than discovering at the security review.

[METHODOLOGY · K-FRAMEWORK]

Integrated through the
K-Framework.

Every model we integrate runs through the same operating system. Three pillars, sixteen layers, one Compound Growth Loop. The methodology that keeps AI work from rotting after the first ship.

Read the K-Framework
01

Foundations

Direct API integration with the model. No LangChain, no orchestration vendor, no agent framework built on quicksand. Typed contracts, the same way we wire up Postgres.

02

Amplification

An eval suite built from your real tasks gates every prompt and model change. Quality is measured before it ships, not vibed in a demo.

03

Judgment

Governance, audit, and oversight wired in from day one. Who called what, with which prompt version, at what cost. Your auditors get answers, not screenshots.

[OBSERVABILITY]

Observability your team can read.

A model in production without observability is roulette. We instrument every integration so engineering and finance can see the same numbers, and so a regression at 3am surfaces before a customer opens a ticket.

Instrumented

Cost per call

Tokens in, tokens out, dollars spent. Sliced by feature, tenant, and route. Budgets enforced where it matters.

Instrumented

Latency p50 / p95 / p99

Real distributions, not averages. We know which routes are slow, and why.

Instrumented

Eval pass rates

The same eval suite that gates a release runs continuously in production. A regression on real traffic surfaces fast.

Instrumented

Prompt + completion logs

PII scrubbed at the proxy, shipped to your SIEM. Retention controls match your compliance window.

Dashboards your team owns, not ours. At handoff you get the queries, the alerts, and the runbook. We are not in the path to read your metrics.

[COMMON QUESTIONS]

Questions we are getting asked.

What is the difference between open weights and open source?
Almost everything on this page is open weights: you can download the parameters and run them. Very few release the training data or the full training code, so they are not open source in the sense the term normally carries. The practical question is not the label, it is what the licence permits, which is why that is the column we lead with.
Which of these can we use commercially without a review?
The MIT and Apache 2.0 rows: GLM-5.3-Flash, Qwen3.8-27B, whisper-large-v3-turbo, Kokoro-82M, FLUX.1-schnell and the embedding and reranking models. Anything marked with a vendor-specific licence needs an actual read, and FLUX.1-dev specifically is non-commercial despite being the most-liked model on Hugging Face.
How much hardware do we need?
It ranges over four orders of magnitude on this page. Kokoro-82M is 82 million parameters and runs almost anywhere; Kimi K3 is 2.8 trillion and is a serious cluster. The useful rule is to size against your p95 context length and concurrency rather than a single-request demo, because long context is where these deployments run out of memory in production having been fine in testing.
Do the published benchmarks apply to the quantised builds?
No, and this is the trap we see most often. Published numbers are BF16 or the original release format, and almost nobody deploys that. A 4-bit GGUF is a different model with different failure modes, and the losses concentrate in long-context and tool-calling behaviour rather than in short prompts. Evaluate the exact artefact going to production.
Are the uncensored variants on the trending list safe to use?
They are the same architectures with the safety training removed, and in a product you inherit both the behaviour and the exposure with no vendor to escalate to. If you use community derivatives at all, pin exact repository revisions and verify checksums, because a model repo is a supply-chain dependency like any other.
Share
View .md
DIRECT INTEGRATION · NO FRAMEWORK

Picking a model
is a routing decision.

We build behind a vendor-neutral abstraction and route by task difficulty at runtime, with an eval suite that decides rather than a launch post. Eval suite at handoff, full source ownership.