We integrate the model like we integrate Postgres.
Direct API integration against frontier and open-weight models. No LangChain, no orchestration vendor, no migration every six months. A thin abstraction so swapping models is a config change.
Claude
Direct Claude integration, eval-tested, no framework lock-in.
Kimi
Latest: Kimi K3, the first open 3T-class model. Open-weight frontier coding, 1M context and native vision, hosted or self-hosted.
Sakana Fugu
Fugu and Fugu Ultra: a multi-agent system delivered as one model. Frontier results, no single-vendor dependency.
Meta (Muse + Llama)
Muse Glimmer: an Apache 2.0 agent model on one GPU. Plus Muse Spark and the Llama line, routed by task and sensitivity.
OpenAI GPT
Production GPT integration with evals and cost control.
Google Gemini
Large context and multimodal understanding, integrated directly.
Llama
Open-weight models you host yourself, for control and residency.
Mistral
Compact, efficient models, hosted or self-run.
DeepSeek
Open-weight reasoning models, evaluated against your tasks.
Embeddings
Vector representations powering RAG, search, and recommendations.
Voice & Speech
Transcription, voice agents, and synthesis, routed by job across vendors.
Qwen
Apache 2.0 models you can self-host, from 0.6B to 180B.
GLM
Coding-first open weights, with an MIT-licensed 320B tier.
Vision
Document AI, image generation, screen grounding, and segmentation.
Or slice it the other way.
The families above are how we integrate. These are how a decision actually gets made: what is open, what is cheap, what codes, what holds a long context. One dataset, updated against the Hugging Face API rather than retyped, with every row dated.
The fifty models worth knowing about.
The fifty models worth knowing in September 2026, grouped by job: frontier APIs, open weights, vision, speech and embeddings, with licences and prices.
Open-weight models, sorted by what the licence lets you do.
Every downloadable model worth running in 2026, sorted by what the licence permits: MIT at 320B, Apache 2.0 at 27B, and the ones needing a legal read.
Models for coding agents, and the harness caveat.
Coding models ranked by published benchmark, with the harness named in every row because the scores are not comparable across them.
The million-token club, and what it actually costs.
Models with 250K to 1M token context windows, ordered by size, with the pricing cliffs and retrieval caveats that decide real usability.
What the models cost, cheapest first.
Every token-priced model ordered by input price, from $0.14 to $10 per million, with the cached-input rates that decide what an agent really costs.