---
title: "GLM-5.3-Flash: MIT at 320B, Benchmarks, Hardware and Licence"
description: "A senior lab's read on GLM-5.3-Flash: MIT licensed at 320B with 18B active, and why the licence matters more than the gap to its own flagship."
source: "https://www.kensink.com/models/glm/glm-5-3-flash/"
canonical: "https://www.kensink.com/models/glm/glm-5-3-flash/"
---
★ MIT LICENCE · 320B · 18B ACTIVE Z.ai Model brief

Z.AI · GLM-5.3-FLASH · MULTIMODAL MoE · 25 AUG 2026

# GLM-5.3-Flash. MIT, at 320 billion parameters.

The most permissive licence anyone has attached to a model this size. Three hundred and twenty billion parameters with eighteen billion active per token, multimodal, a 300K context, and 84.3 on Terminal-Bench 2.1, four points behind its own 753B flagship. For a commercial deployment the licence is doing more work here than any benchmark on the page.

Open weights MIT licence Self-hosted Eval pipelines

[Use it on your project →](https://www.kensink.com/contact) [Z.ai documentation ↗](https://huggingface.co/zai-org/GLM-5.3-Flash)

Released

25 Aug 2026

Model ID

zai-org/GLM-5.3-Flash

Licence

MIT

Parameters

320B / 18B active

Context

300K tokens

Max output

163,840 tokens

Modalities

Text, image → text

Active

18B per token

\[TL;DR FOR CEO + CTO\]

## What to know.

-   01
    
    ### MIT at 320B is the headline, and it is not close.
    
    MIT is about as permissive as software licensing gets: use it, modify it, ship it, sell it, no field-of-use restriction and no user threshold. Attached to a 320-billion-parameter multimodal model, that is a genuinely unusual thing to be able to write down, and it removes an entire category of commercial risk.
    
-   02
    
    ### Four points behind its own flagship, at less than half the size.
    
    Terminal-Bench 2.1 comes in at 84.3 against 88.2 for the 753B GLM-5.3. For that you get a model you can actually deploy, a licence your legal team already knows, and image input the flagship does not have. On most commercial work that trade is not a close call.
    
-   03
    
    ### It is more downloaded and more liked than the flagship.
    
    784k downloads and 2,131 likes against 442k and 1,746. The community has already voted, and it voted for the smaller, permissively licensed, multimodal one. That is a useful signal about what practitioners actually deploy rather than what benchmarks best.
    
-   04
    
    ### Eighteen billion active per token.
    
    Compute per token sits near an 18B dense model while the full 320B must stay resident. Cheap tokens, substantial machine, and a break-even that depends on utilisation rather than on peak throughput. The usual sparse trade, at a size that is reachable rather than aspirational.
    
-   05
    
    ### Explicit reasoning control.
    
    reasoning\_effort at low, high, or max. Set it per call site rather than globally: max on planning and hard debugging, low on the mechanical turns. On a long agent run that single setting moves cost more than the choice between these two models does.
    

\[SOFTWARE DEVELOPMENT IMPACT\]

## What it changes for the team building with it.

The comparison that decides this is internal to the family. Flash against the 753B flagship on licence, hardware and capability, and then both against a frontier API.

| Dimension | vs GLM-5.3 (753B) | vs a frontier API |
| --- | --- | --- |
| 
Licence

 | MIT against a bespoke glm-5.3 licence. This is the difference that matters most and it points one way: MIT needs no review, carries no field-of-use restriction, and cannot be reinterpreted against you later. On client work it is usually decisive on its own. | An API has commercial terms you accept and a counterparty you can escalate to. MIT weights have neither and need neither, which is a simpler position than either alternative. |
| 

Hardware

 | 320B resident against 753B, with 18B active. Still a serious deployment, but it is the difference between reachable and multi-node-only. Combined with the FP8 shipping format, this is the GLM most teams can realistically self-host. | An API needs no hardware. Self-hosting has to win on privacy, latency or sustained volume, and Z.ai's own hosted rates at $1.40 / $4.40 are low enough that the bar is higher than it would be against a frontier vendor. |
| 

Capability

 | 84.3 against 88.2 on Terminal-Bench 2.1, so roughly four points for less than half the parameters. Against that, Flash adds image input, which the flagship does not have. Whether four points outweigh a modality depends entirely on what you are building. | Behind the frontier on the hardest work, as its flagship is. It closes most of the gap on mainstream coding and does not close it at the top end. Route the hardest calls up and keep the volume here. |
| 

Multimodal

 | Image and text in, where the 753B flagship is text only. For document work, screenshots in a coding loop, or anything with a visual input, Flash is not the compromise choice, it is the only choice in this family. | The frontier models still lead on dense document and chart reading. For ordinary screenshots and diagrams inside an agent loop, an MIT-licensed model you host yourself is a very different proposition from a per-image API call. |

This is the GLM we reach for first. The licence removes a review, the size makes self-hosting realistic, and image input covers cases the flagship cannot. We route up to the 753B or to a frontier model when an eval on customer tasks shows the gap is real, and not before.

\[WHAT IS NEW\]

## The features that ship with it.

01

### MIT at this scale

The permissive-licence frontier moved. Apache 2.0 at 27B from Qwen was already notable; MIT at 320B is a different statement. No field-of-use restriction, no acceptable-use policy that can be revised, no user threshold that triggers a negotiation.

02

### Manifold-Constrained Hyper-Connections

The architecture pairs sparse and linear attention with mHC, a connection scheme Z.ai document in arXiv 2602.15763. It is how the model holds a 300K context at 18B active without the memory curve becoming impractical.

03

### Image input, which the flagship lacks

Flash is image-text-to-text where GLM-5.3 is text-only. The cheaper, more permissively licensed model in the family is also the more capable one on modality, which is an unusual and welcome inversion.

04

### 163,840 max output

An unusually large output ceiling, well beyond the 128K that has become standard. For document generation, long refactors, or anything that produces a large artefact in one pass, that ceiling is the constraint that usually bites.

05

### reasoning\_effort at low, high and max

Three explicit levels rather than an opaque default. Set it per call site: max for planning and hard debugging, low for mechanical turns. On a long agent run this moves cost more than most model choices do.

06

### Shipped in FP8

Already quantised in the published checkpoint, so the memory bill is lower than a BF16 release of the same size and the artefact you download is the artefact that was measured. That removes the most common open-weight evaluation trap before you hit it.

\[THE SPEC\]

## Everything an integration depends on.

The limits, the endpoints, and the lines that decide whether this model fits your transport and your budget. Kept here so nobody has to reconstruct it from three vendor pages.

<table class="w-full min-w-[560px] border-collapse text-left"><tbody><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Hugging Face repo</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">zai-org/GLM-5.3-Flash</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Licence</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">MIT. No field-of-use restriction, no user threshold, no bespoke terms</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Parameters</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">321,323,031,390 total, 320B nominal. Shipped in FP8 E4M3 with BF16 components</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Active parameters</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">18B per token</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Architecture</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">Glm5NextForConditionalGeneration. Hybrid sparse and linear attention with Manifold-Constrained Hyper-Connections</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Context window</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">300,000 tokens, as evaluated</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Max output</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">163,840 tokens</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Modalities</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">Image and text in, text out. English and Chinese</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Reasoning control</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">reasoning_effort at low, high, or max</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Serving</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">SGLang, vLLM, TokenSpeed, Transformers, KTransformers, Unsloth</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Hugging Face signal</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">784k downloads and 2,131 likes, ahead of the 753B flagship on both</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Paper</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">arXiv 2602.15763</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[34%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Flagship sibling</th><td class="p-4 align-top text-[14px] leading-relaxed text-ink-2">GLM-5.3 at 753B, bespoke licence, 88.2 on Terminal-Bench 2.1</td></tr></tbody></table>

\[PRICING\]

## What it costs to run.

The weights are MIT and free. These are the numbers a self-hosted deployment gets measured against.

<table class="w-full min-w-[560px] border-collapse text-left"><tbody><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[30%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Flash weights</th><td class="w-[22%] p-4 align-top text-[15px] font-bold text-ink">Free</td><td class="p-4 align-top text-[13px] leading-relaxed text-ink-4">MIT. No commercial restriction and no legal review needed.</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[30%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Resident memory</th><td class="w-[22%] p-4 align-top text-[15px] font-bold text-ink">320B params</td><td class="p-4 align-top text-[13px] leading-relaxed text-ink-4">FP8 in the published checkpoint, so lower than a BF16 release of the same size.</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[30%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">Active per token</th><td class="w-[22%] p-4 align-top text-[15px] font-bold text-ink">18B</td><td class="p-4 align-top text-[13px] leading-relaxed text-ink-4">Compute per token near an 18B dense model.</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[30%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">GLM-5.3 hosted API</th><td class="w-[22%] p-4 align-top text-[15px] font-bold text-ink">$1.40 / $4.40</td><td class="p-4 align-top text-[13px] leading-relaxed text-ink-4">The flagship through Z.ai. Cached input $0.26.</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[30%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">GLM Coding Plan</th><td class="w-[22%] p-4 align-top text-[15px] font-bold text-ink">$18 to $168 / mo</td><td class="p-4 align-top text-[13px] leading-relaxed text-ink-4">Flat rate. About $12.60 to $117.60 billed annually.</td></tr><tr class="border-b border-rule last:border-b-0"><th scope="row" class="w-[30%] bg-paper-3 p-4 align-top font-mono text-[10.5px] font-normal uppercase tracking-mono-wide text-ink-4">For comparison</th><td class="w-[22%] p-4 align-top text-[15px] font-bold text-ink">$10 / $50</td><td class="p-4 align-top text-[13px] leading-relaxed text-ink-4">GPT-6 Astra and Claude Fable 5.1 per million tokens.</td></tr></tbody></table>

The honest framing is that Z.ai's hosted pricing is low enough to make self-hosting a harder case than it would be against a frontier vendor. At $1.40 input you need real sustained volume, a privacy constraint, or a latency floor before your own GPUs win on a spreadsheet that includes engineering time. What MIT weights buy you is not cheaper inference, it is the right to deploy on your own terms, inside your own boundary, with no counterparty. That is worth a great deal on some engagements and nothing at all on others, and it is worth being clear about which one you are on.

\[RISK + GOVERNANCE\]

## What to watch.

### MIT means nobody can change the terms later.

That is the substantive benefit, and it is larger than it sounds. Community and bespoke licences carry acceptable-use policies that can be revised, thresholds that trigger renegotiation, and field-of-use limits that are read differently by different lawyers. MIT settles all of it once, and the settlement survives the vendor.

### You inherit the whole safety layer.

No vendor classifier between prompt and model, no refusal path, no abuse monitoring. Input filtering, output review, rate limiting and audit logging become your build. Teams consistently underestimate this and it is the usual reason a self-hosted pilot stops short of production.

### The family's security positioning travels with it.

Z.ai aim GLM at security research and the flagship tops CyberGym. Flash sits in the same family and behind the same absence of vendor gating. Agree scope, tool permissions, logging and a human gate on consequential actions before it goes into a client environment, rather than after.

\[OUR TAKE\]

## How we are using it.

01

### This is the open-weight model we would defend in a legal review.

MIT at 320B multimodal removes the conversation that stalls most open-weight plans. We have watched capable models lose to less capable ones purely on licence terms, and this is the first at this scale where that argument does not arise.

02

### The community picked it over the flagship, and they were right.

More downloads and more likes than the 753B, despite scoring four points lower. Practitioners are optimising for deployability and licence rather than for the top line, which is the same calculation we make on client work.

03

### Image input on the cheaper tier is the quiet win.

The flagship is text-only. A coding agent that can look at a screenshot, a document pipeline that reads a scan, a support flow that handles a photo, all of those need Flash and none of them can use GLM-5.3 at all.

04

### Set reasoning\_effort per call site on day one.

Three explicit levels is a gift and leaving it at one setting for the whole session wastes it. High on planning and debugging, low on mechanical turns. This single decision moves an agent's cost more than the choice between these two models does.

05

### Z.ai's own API is the number to beat, not the frontier.

At $1.40 / $4.40 the hosted path undercuts most self-hosting maths once engineering time is counted honestly. Self-host these weights for privacy, latency or genuine sustained volume. Do not self-host them to save money without doing the sums first.

\[METHODOLOGY · K-FRAMEWORK\]

## Integrated through the  
K-Framework.

Every model we integrate runs through the same operating system. Three pillars, sixteen layers, one Compound Growth Loop. The methodology that keeps AI work from rotting after the first ship.

[Read the K-Framework](https://www.kensink.com/k-framework)

01

### Foundations

Direct API integration with the model. No LangChain, no orchestration vendor, no agent framework built on quicksand. Typed contracts, the same way we wire up Postgres.

02

### Amplification

An eval suite built from your real tasks gates every prompt and model change. Quality is measured before it ships, not vibed in a demo.

03

### Judgment

Governance, audit, and oversight wired in from day one. Who called what, with which prompt version, at what cost. Your auditors get answers, not screenshots.

\[OBSERVABILITY\]

## Observability your team can read.

A model in production without observability is roulette. We instrument every integration so engineering and finance can see the same numbers, and so a regression at 3am surfaces before a customer opens a ticket.

Instrumented

### Cost per call

Tokens in, tokens out, dollars spent. Sliced by feature, tenant, and route. Budgets enforced where it matters.

Instrumented

### Latency p50 / p95 / p99

Real distributions, not averages. We know which routes are slow, and why.

Instrumented

### Eval pass rates

The same eval suite that gates a release runs continuously in production. A regression on real traffic surfaces fast.

Instrumented

### Prompt + completion logs

PII scrubbed at the proxy, shipped to your SIEM. Retention controls match your compliance window.

Dashboards your team owns, not ours. At handoff you get the queries, the alerts, and the runbook. We are not in the path to read your metrics.

\[COMMON QUESTIONS\]

## Questions we are getting asked.

Is MIT really unrestricted?

For the weights, yes. MIT grants use, modification, distribution and sale with attribution and no warranty, with no field-of-use restriction and no user threshold. What it does not do is settle your obligations under the EU AI Act, sectoral regulation, or your own customer contracts. Those attach to what you build rather than to what you downloaded.

Should we use Flash or the 753B GLM-5.3?

Flash, for most work. It is four points behind on Terminal-Bench 2.1, less than half the size, MIT rather than a bespoke licence, and it takes image input the flagship cannot. Move up to the 753B only when an eval on your own tasks shows the ceiling matters and legal has read the licence.

What hardware does it need?

320B resident, shipped in FP8, with 18B active per token. That is a multi-GPU deployment but a reachable one, unlike the 753B flagship. Size against your p95 context and concurrency rather than a single-request demo, since a 300K context window is where these deployments run out of memory in production having been fine in testing.

How does it compare to Qwen3.8-27B?

Different shapes. Qwen3.8-27B is 27B dense under Apache 2.0 and fits on one card. Flash is 320B sparse under MIT and needs several. Both are permissively licensed and multimodal. If you have one accelerator the decision is already made; if you have a cluster, run both against your own tasks, because the published benchmarks do not use a common harness.

Is 300K context real or extrapolated?

300,000 tokens is what the card documents as evaluated, alongside a 163,840 max output. We would still run a needle-style retrieval test at your actual context length before relying on it, as we would for any model. Published context and useful context are different numbers more often than vendors like to acknowledge.

Can we fine-tune it?

MIT permits it without restriction, and Unsloth is listed among the supported stacks, so the tooling path exists. The practical constraint is that fine-tuning a 320B sparse model is a serious undertaking. On most engagements retrieval and prompt work reach the same outcome for a fraction of the cost, and we would exhaust those first.

Share[](https://twitter.com/intent/tweet?url=https%3A%2F%2Fwww.kensink.com%2Fmodels%2Fglm%2Fglm-5-3-flash%2F&text=GLM-5.3-Flash%20brief)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Fwww.kensink.com%2Fmodels%2Fglm%2Fglm-5-3-flash%2F)

[View .md](https://www.kensink.com/models/glm/glm-5-3-flash.md)

\[RELATED\]

## Worth a look next.

[

MODEL

GLM-5.3 (the 753B flagship)

Read more

](https://www.kensink.com/models/glm/glm-5-3/)[

MODEL

Qwen3.8-27B (Apache 2.0)

Read more

](https://www.kensink.com/models/qwen/qwen3-8-27b/)[

COLLECTION

Open-weight models, 2026

Read more

](https://www.kensink.com/models/open-weight-2026/)

DIRECT INTEGRATION · NO FRAMEWORK

## Want GLM-5.3-Flash  
in your product?

Eval suite at handoff, full source ownership. We integrate against the model API the same way we integrate against Postgres, and route by task at runtime. Sized to your scope.

[Start a conversation →](https://www.kensink.com/contact) [All GLM models](https://www.kensink.com/models/glm)
