Kensink Labs
★ Voice & SpeechLLM Models8-week engagement
VOICE · TRANSCRIPTION, AGENTS, SYNTHESIS

Production voice. Transcription, agents, translation, synthesis.

Whisper made transcription a commodity, and the category has moved on twice since. Streaming models replaced chunk-and-stitch, realtime models put reasoning inside the audio loop, and the vendor that leads on word error rate is not the one that leads on turn detection. We build the whole pipeline and route by job.

Realtime APILLM APIEval pipelines
Cycle
8 weeks · fixed price
Stack
Speech + realtime audio
Output
Production code + eval suite
Handoff
Full source ownership
[THE SHORT VERSION]

Transcription is step one, not the whole job.

Modern speech models transcribe accurately enough that accuracy is rarely what decides a project. The value is in the pipeline: diarisation, timestamps, formatting, redaction, and a clean handoff into summarisation, search, or extraction. The model choice is an afternoon. Building the path that makes the words useful, and proving it on your own recordings rather than on a public leaderboard, is the work.

When it fits
  • Meeting, call, and media transcription at volume
  • Voice agents that resolve rather than deflect
  • Live captions, dictation, and real-time note taking
  • Cross-language conversation and multilingual support desks
  • Audio search, quality review, and analysis pipelines
When it does not
  • Audio that cannot leave your infrastructure, unless you self-host open weights
  • Anything where you have not settled recording consent for your jurisdictions
[HOW WE BUILD IT]

How we build with Voice & Speech.

01

Scope and fit

We decide where Voice & Speech earns its place in your system, and where a simpler tool wins. No resume-driven architecture.

02

Build on a tested foundation

We integrate Voice & Speech against a foundation we trust: typed code, CI, and observability from the first commit. Boring infrastructure, modern surface.

03

Eval before launch

An eval suite proves the build behaves before it reaches a user. We measure, then ship.

04

Handoff with ownership

Your team gets the code, the tests, and a runbook. No lock-in to us or to a vendor framework.

[WHAT YOU GET]

What the engagement leaves behind.

Senior
Engineers who have shipped this before
100%
Source ownership at handoff
Eval-first
Tested before it ships
0
Framework lock-in
[MODELS + VENDORS]

The speech models worth wiring in.

Transcription, speech to speech, translation, and synthesis, across the vendors that lead on each. We integrate them directly behind one abstraction and route by job, because no single provider wins every axis: accuracy, latency, turn detection, language coverage, and cost per hour pull in different directions.

FlagshipOpenAI
May 2026

GPT-Realtime-2.1

gpt-realtime-2.1
Speech to speech, with reasoning and tool use
Audio in
$32 / 1M
Audio out
$64 / 1M
  • The voice-agent model. GPT-5-class reasoning with function calling in the loop
  • 128K context, audio and text and image in, audio and text out
  • Cached input at $0.40 per million, which is the lever that makes it affordable
Read the technical brief
CurrentOpenAI
May 2026

GPT-Realtime-2.1 Mini

gpt-realtime-2.1-mini
Speech to speech, compact tier
Audio in
$10 / 1M
Audio out
$20 / 1M
  • Roughly a third the audio price of the full 2.1, on the same Realtime API
  • Text at $0.60 in and $2.40 out, against $4 and $24 on the full tier
  • Our default for scripted or narrow voice flows that do not need deep reasoning
In the routing table · brief not yet published
FlagshipOpenAI
May 2026

GPT-Transcribe

gpt-transcribe
File transcription, batch
Audio
$0.0045 / minute
  • The cheapest path to accurate text. About $0.27 per hour of audio
  • Takes keyword hints and unstructured context for domain terms and code-switching
  • Serves both the file endpoint and a realtime transcription session
Read the technical brief
CurrentOpenAI
May 2026

GPT-Realtime-Whisper

gpt-realtime-whisper
Streaming speech to text
Audio
$0.017 / minute
  • Whisper's successor, rebuilt as a streaming model rather than a batch one
  • Live transcript deltas for captions, notes, and voice interfaces
  • Just under four times the price of GPT-Transcribe, for latency you can feel
Read the technical brief
CurrentOpenAI
May 2026

GPT-Live-Transcribe

gpt-live-transcribe
Streaming speech to text, tunable
Audio
$0.017 / minute
  • Same price as Realtime-Whisper, with an explicit latency and accuracy dial
  • Takes keyword hints, unstructured context, and multiple language hints
  • Realtime transcription sessions only, no file endpoint
In the routing table · brief not yet published
CurrentOpenAI
May 2026

GPT-Realtime-Translate

gpt-realtime-translate
Speech to speech translation
Audio
$0.034 / minute
  • Conversational translation that keeps pace with the speaker
  • More than 70 input languages, 13 output languages
  • Billed by the minute, so cost per call is predictable before you ship
In the routing table · brief not yet published
CurrentOpenAI
Mar 2025

GPT-4o Mini TTS

gpt-4o-mini-tts
Text to speech
Audio out
$12 / 1M
Text in
$0.60 / 1M
  • Synthesis for the paths where you generate the words yourself
  • Cheaper than a full realtime session when there is no conversation to hold
  • Right for notifications, IVR prompts, and read-aloud
In the routing table · brief not yet published
CurrentGoogle
2026

Gemini 3.5 Transcribe

gemini-3.5-transcribe
Speech to text, with diarisation
Audio in
$0.003 / min
Text out
$0.002 / min
  • Speaker diarisation and word timestamps in the model, not in your pipeline
  • Utterance-based language detection for multilingual and code-switching audio
  • A live variant, gemini-3.5-transcribe-live, for low-latency streams
In the routing table · brief not yet published
CurrentElevenLabs
Nov 2025

ElevenLabs Scribe v2 Realtime

scribe-v2-realtime
Streaming speech to text
Billing
Per minute, tiered
  • Lowest word error rate on the Artificial Analysis leaderboard as of June 2026, at 2.2%
  • Sub-150ms latency with more than 90 languages
  • Where we go when transcript accuracy is the product, not a step in it
In the routing table · brief not yet published
CurrentDeepgram
Apr 2026

Deepgram Flux Multilingual

flux-multilingual
Conversational speech to text
Billing
Per minute, tiered
  • End-of-turn detection built into the model, so you can drop the external VAD
  • Median end-of-turn detection under 300ms
  • The pick when interruption handling decides whether a voice agent feels human
In the routing table · brief not yet published
LegacyOpenAI
Mar 2023

Whisper v1

whisper-1
File transcription, original
Audio
$0.006 / minute
  • The model that made transcription a commodity, and still self-hostable open weights
  • Superseded on the API by GPT-Transcribe, which is cheaper and stronger
  • Keep it only where you need to run the weights on your own hardware
In the routing table · brief not yet published
[METHODOLOGY · K-FRAMEWORK]

Integrated through the
K-Framework.

Every model we integrate runs through the same operating system. Three pillars, sixteen layers, one Compound Growth Loop. The methodology that keeps AI work from rotting after the first ship.

Read the K-Framework
01

Foundations

Direct API integration with the model. No LangChain, no orchestration vendor, no agent framework built on quicksand. Typed contracts, the same way we wire up Postgres.

02

Amplification

An eval suite built from your real tasks gates every prompt and model change. Quality is measured before it ships, not vibed in a demo.

03

Judgment

Governance, audit, and oversight wired in from day one. Who called what, with which prompt version, at what cost. Your auditors get answers, not screenshots.

[OBSERVABILITY]

Observability your team can read.

A model in production without observability is roulette. We instrument every integration so engineering and finance can see the same numbers, and so a regression at 3am surfaces before a customer opens a ticket.

Instrumented

Cost per call

Tokens in, tokens out, dollars spent. Sliced by feature, tenant, and route. Budgets enforced where it matters.

Instrumented

Latency p50 / p95 / p99

Real distributions, not averages. We know which routes are slow, and why.

Instrumented

Eval pass rates

The same eval suite that gates a release runs continuously in production. A regression on real traffic surfaces fast.

Instrumented

Prompt + completion logs

PII scrubbed at the proxy, shipped to your SIEM. Retention controls match your compliance window.

Dashboards your team owns, not ours. At handoff you get the queries, the alerts, and the runbook. We are not in the path to read your metrics.

[COMMON QUESTIONS]

Questions we get asked.

What replaced Whisper?
Two things, depending on the job. For batch transcription it is GPT-Transcribe at $0.0045 per minute, which is cheaper than the legacy whisper-1 endpoint and stronger. For live audio it is GPT-Realtime-Whisper at $0.017, which streams natively and removes the chunk-and-stitch layer every live Whisper deployment had to write. Whisper's open weights remain the answer when audio cannot leave your infrastructure.
How much does transcription cost at scale?
About $0.27 per hour of audio on the batch model, so a thousand hours a month is roughly $270. The streaming tier is about $1.02 an hour. The mistake we see most often is a single streaming pipeline handling everything, which turns a $270 line item into a $1,020 one for a latency property only a fraction of the streams actually use.
Which provider is most accurate?
On the Artificial Analysis leaderboard as of June 2026, ElevenLabs Scribe v2 led at 2.2% word error rate against 5.2% for Deepgram Nova-3. That ranking is weaker evidence than it looks: clean English accuracy has plateaued and the differences that matter show up on accents, background noise, overlapping speakers and domain vocabulary. Test on your own recordings.
Do these models do speaker diarisation?
The OpenAI transcription models do not; that is pipeline work, usually pyannote. Google took the other approach: Gemini 3.5 Transcribe carries diarisation, word timestamps and utterance-based language detection in the model at $0.003 per minute of audio. If those are on your requirements list, price both routes before committing to building the layer.
Should we build a voice agent or stitch a pipeline?
Ask whether anyone is waiting. If a human is on the line and can interrupt, use a realtime session: no stitched pipeline matches it on turn latency or overlap handling. If the output is a transcript or a summary, stitch it, because a cheap transcription model plus a strong text model is cheaper and gives you a better text model.
APPLIED K-FRAMEWORK

Bring the problem.
We’ll bring the build.

Senior engineers, eval suite at handoff, full source ownership. Sprint, program, or ongoing. We shape the engagement to the work.