Kensink Labs
PRACTICE · 02 · FRONTIER FIRM

Reorganise your company around AI.
The Frontier Firm program.

A twelve-week organisational transformation, led by a lab that already operates this way. Diagnostic, roadmap, rollout. The org-chart change, not a deck.

▲ THE PRODUCTIVITY CLIFF

Productivity is no longer a slope. It's a cliff.

A new class of company is emerging: the Frontier Firm. They don't just "plug in" AI. They redesign their entire operational DNA around it. The gap between them and everyone else is widening exponentially.

−80%
Cost on human tasks
+60%
Throughput, day 90
2–5×
Productivity gains, frontier firms
8 wk
From kickoff to production
▲ WHERE YOU ACTUALLY ARE · 05 ZONES

One in five companies is
where it thinks it is.

Capability and organisational readiness are separate axes. A company can be strong on one and flat on the other, and the remedy for each failure is different. Buying more licences fixes exactly one of the five positions below.

19%

Frontier

Capability and organisational readiness are both high. The work has been redesigned, not just tooled. This is the group everyone else is benchmarking against.

~50%

Emergent

Real usage, real intent, neither axis settled. Progress is genuine but it lives in pockets and does not survive a reorg or a departure.

16%

Stalled

Low capability and thin support. Licences were bought, training was optional, and nobody owns the outcome. The tooling is present and the practice never formed.

10%

Blocked agency

Strong skills, obstructive systems. A fast engine bolted to a chassis that cannot take the torque. The people are ready; procurement, data access, and review gates are not.

5%

Unclaimed capacity

The organisation is ready and the people have not caught up. Platform, policy, and budget are in place, and demand for them has not materialised.

Find your zone

Twelve questions across leadership, tech, process, and talent. You see the score before we see your email.

Run the diagnostic →
67% / 32%
Share of AI impact from organisational vs individual factors
26%
Say leadership is clearly and consistently aligned on AI
13%
Are rewarded for redesigning work, regardless of outcome
15×
Year-on-year growth in active agents, 18× at large enterprises
▲ EVOLUTION · 03 STAGES

From automating tasks
to augmenting minds.

ERA 01 · 1960s to 1980s

Process automation

Mainframes automated siloed, backend tasks like accounting. Impact was specific, deterministic, and isolated.

ERA 02 · 1990s to 2010s

Workflow digitization

The PC, Internet, and SaaS brought digital efficiency to every department. Software helped humans work faster.

ERA 03 · Today

Cognitive augmentation

AI agents don't just do tasks; they reason, create, and collaborate. AI becomes a teammate, automating cognitive and creative work at scale.

▲ HOW THE WORK HAPPENS · 04 MODES

The skill is not using AI.
It is knowing which mode the task calls for.

Two dials sit behind every session: how much the agent does, and how much the person stays in it. Strong practitioners are not distinguished by which setting they favour. They are distinguished by choosing deliberately, before the work starts.

MODE 01

Asking

Low agent · low human

A question with a short answer. Cheap, fast, and no workflow attached. Most organisations never get past this and call it adoption.

MODE 02

Exploration

Low agent · high human

The person is doing the thinking and using the model to widen the search. Judgement stays entirely on the human side.

MODE 03

Collaboration

High agent · high human

Sustained back and forth on work that matters. The most expensive mode in attention, and the one that produces work neither party could produce alone.

MODE 04

Delegation

High agent · low human

A specified job handed over end to end. Only safe where the standard is written down and the output can be checked.

And the job title moves up the ladder.

The same progression a writer makes going from filing copy, to editing a staff writer, to briefing an agency, to running the newsroom. The output grows at every rung. So does the cost of a bad standard.

01

Author

You produce the work and reach for the model on individual components.

02

Editor

The model drafts. You revise, approve, and carry the byline.

03

Director

You write the specification and the quality bar. The agent executes the whole task.

04

Orchestrator

Several agents run in parallel across a workflow. You handle the exceptions they escalate.

STRUCTURAL SHIFT

The structure isn't just different. It's inverted.

THE TRADITIONAL FIRM

Human-led, manual tasks

  • Linear productivity growth
  • Sequential, manual workflows
  • AI tools sit beside the team
  • Departments operate in silos
  • Project timelines measured in quarters
THE FRONTIER FIRM

Hybrid human + AI collaboration

  • Exponential productivity growth
  • AI-augmented developers, +40% velocity
  • AI project managers that automate scheduling & reporting
  • Hybrid teams where humans and agents work as one unit
  • Project timelines measured in weeks
THE ORG-CHART CHANGE · 04 PARTS

An org chart is a seating plan.
A work chart is a production schedule.

This is the part that cannot be bought. Four structural pieces have to exist before agent capacity turns into company output, and none of them ship with a licence agreement.

01

The work chart

An org chart maps who reports to whom. A work chart maps outcomes to the mix of people and agents that produce them. One is a seating plan, the other is a production schedule. Companies that reorganise around jobs to be done stop asking which department owns a process.

02

The agent boss

Every manager now runs a mixed team. Direction, standards, and review move up the job description; step-by-step execution moves out of it. The skill is knowing which of the four modes a task calls for, and being answerable when the agent gets it wrong.

03

Intelligence resources

Someone has to run digital labour at the organisational level: provisioning, scope, retirement, performance review. It sits between IT and HR and belongs cleanly to neither, which is why in most companies it currently belongs to nobody.

04

Owned intelligence

The know-how a firm accumulates about its own work: the prompts, the evals, the failure cases, the escalation rules. Model access is a commodity your competitor can buy on the same terms. This is the part they cannot.

▲ WHAT CHANGES FOR YOUR PEOPLE

The question in the room
is never about the model.

Organisational factors account for roughly twice the AI impact of individual capability. The awkward implication for leadership is that when a rollout fails, the people were usually not the constraint.

Headcount is the wrong first question

Half of organisations report reduced need for entry-level roles, and the firms cutting first are largely the ones that automated a process nobody had redesigned. We recommend against cuts in year one, and it is written into our scope. The capacity you free is worth more redeployed than removed.

Managers decide whether this works

Strategy sets direction and managers operationalise it. Where managers use AI in the open, set an explicit quality bar, and make experimentation safe, every downstream measure moves. Where they do not, the licences go unused and the pilot dies quietly.

Skill atrophy is a real cost

The strongest practitioners deliberately do some work without AI to keep the underlying skill sharp. A pilot who only ever flies on autopilot still has to land the aircraft when the weather turns. Build the exception path into the operating model rather than discovering it during one.

Reward the redesign, not just the result

Only thirteen percent of AI users say they are rewarded for reinventing how work is done when the result falls short. Until the incentive covers the attempt, rational people will keep doing the work the old way and using AI at the margins.

Measured lift where managers model the behaviour themselves

+17 pts
Reported AI value when managers visibly use AI themselves
+22 pts
Critical thinking about AI output under the same managers
+30 pts
Trust in agentic AI when managers model the behaviour
+20 pts
Readiness where experimentation is psychologically safe
▲ THE PART VENDORS LEAVE OUT

Most of this fails.
It fails in four predictable places.

Bridges fail at the joints, not in the span. Agent programs are the same. The model is rarely the weak member. The connections between the model and the organisation are where the load goes unheld.

95%
Of GenAI deployments show no measurable P&L impact (MIT NANDA)
88%
Of agent pilots never reach production
40%+
Of agentic projects expected to be cancelled by 2027 (Gartner)
22%
Of agent deployments report negative ROI at twelve months
FAILURE 01

The pilot has no unit of work

“Modernise support” cannot be shipped, measured, or refused. It can only be discussed.

What we do instead

Weeks one and two are spent with the team timing the actual work. We leave with a named unit, a baseline number, and a written definition of done.

FAILURE 02

Nobody can say whether it is working

Without an eval suite, quality is a matter of opinion, and the loudest opinion in the room wins.

What we do instead

Every agent ships with a behaviour eval suite and a gate in CI. Prompt and model changes are diffed against it before merge.

FAILURE 03

The pilot cannot survive contact with the org

Bridges fail at the joints, not in the span. The model works. The handoffs to review, escalation, and audit do not exist.

What we do instead

The workflow is restructured with the team while the agent is being built, not after. Named owners, weekly review, written escalation path.

FAILURE 04

The bill arrives in month four

Agentic tasks consume five to thirty times the tokens of a chatbot turn. Budgets built on chatbot arithmetic break quietly.

What we do instead

Unit cost per workflow is modelled before the build and instrumented after it. Per-route spend and p95 latency are alerted on from day one.

THE PROGRAM · 04 PHASES

Twelve weeks,
with a gate on each one.

Each gate works the way a pressure test works on a pipeline. Nothing moves to the next section until this one holds. If a phase fails its gate, you find out in week four rather than week eleven.

PHASE 01 · Weeks 01 to 02

Find the unit of work

We sit with the team and time the work. Candidates are ranked by value and risk, not by how interesting they are to automate.

You leave the phase with
  • A named unit of work with a measured baseline
  • A ranked backlog with value and risk on each item
  • A written definition of done
GateA unit the team agrees is worth changing
PHASE 02 · Weeks 03 to 06

Build and pilot

One agent, in front of one team, on real volume. We measure deflection and read every correction the team makes.

You leave the phase with
  • One agent in production behind a feature flag
  • Behaviour eval suite wired into CI
  • Observability on cost, latency, and correction rate
GateEvals pass and the pilot team keeps using it
PHASE 03 · Weeks 07 to 10

Harden and restructure

Failure modes you did not imagine at kickoff surface here. We close them, then reshape the team's workflow around what the agent now handles.

You leave the phase with
  • Expanded eval coverage on discovered failure modes
  • Reshaped workflow with named owners and escalation path
  • Review cadence the team runs without us
GateThe workflow holds for two weeks unattended
PHASE 04 · Weeks 11 to 12

Generalize the pattern

The handover is the deliverable. What you get is the method, so the second and third units do not need us.

You leave the phase with
  • Written playbook for the next unit of work
  • Onboarding doc for the next team
  • Runbook, reviewed in person
GateYour team runs the next unit without us

Full written scope, deliverables, exclusions, and SLA are published rather than quoted on request.

Read the full engagement spec →
▲ WHAT IT COSTS TO RUN

The price per token keeps falling.
The bill keeps rising.

Agent spend behaves the way cloud spend behaved a decade ago. The unit price drops every year and the monthly invoice grows anyway, because the falling price is precisely what makes the new usage affordable enough to attempt.

5–30×
More tokens per task than a standard chatbot turn
73%
Of reviewed agentic projects came in over budget
−67%
Fall in price per token, early 2025 to early 2026
80–90%
Of lifetime compute spend goes to inference, not training

Four controls, fitted during the build rather than after it.

01

Unit cost before the build

Every candidate workflow gets a modelled cost per run before anyone writes a prompt. If the arithmetic does not clear the baseline it replaces, we say so in week two and you have not spent the build budget.

02

Routing, not one big model

Classification and extraction do not need the frontier tier. Work is routed by difficulty, with the expensive model reserved for the steps that actually need it.

03

Caching at the boundary

Stable context is cached rather than resent on every turn. On long-running agent loops this is the difference between a viable unit cost and an unviable one.

04

Alerts per route, not per month

Spend and p95 latency are alerted on at the route level. You find out on the day a loop starts retrying, not when the invoice arrives.

CONTROL PLANE · 04 REQUIREMENTS

The regulator will not ask
whether the agent acted.

They will ask why it was permitted to. Answering that after the fact is expensive and usually impossible. Four things have to be true at build time for the answer to exist at all.

01

Agents get identities

An agent running on a shared service account is a night shift signing in with the building's master key. The work gets done and there is no record of who did what. Each agent carries its own identity, scope, and expiry.

02

A principal hierarchy

Every action traces back to who authorised it, with what scope, for how long. When one agent calls another, accountability has to survive the handoff or it fragments on the first incident.

03

The audit trail is the product

Authentication events, tool invocations, delegation handoffs, and policy decisions are logged in a form that answers the regulator's question, which is not whether the agent acted but why it was permitted to.

04

Evals as the release gate

Autonomy scales output and it scales bad output at the same rate. The eval suite is what stands between a prompt change and a thousand wrong answers before lunch.

YOUR PATH · 02 ENGAGEMENTS

Two ways in.

Start with strategy. Or skip to integration if your team has already done the discovery work.

ENGAGEMENT 01

Frontier strategy consulting

A custom roadmap for AI integration that targets maximum ROI and minimizes disruption. A four-week sprint with two senior architects on-site, ending in a written brief naming the smallest valuable thing to ship.

  • Org-readiness audit across leadership, tech, process, talent
  • Prioritized AI use-case backlog with ROI estimates
  • 12-week transformation roadmap
  • Hand-off to your team, or to track 02 below
Fixed scope · fixed feeStart a conversation →
ENGAGEMENT 02

AI workforce integration

Deploying a hybrid workforce of custom AI agents and your existing talent into new, more efficient workflows. Twelve weeks of build + handoff, with the same senior team that scoped you.

  • Agent design across 3–5 priority workflows
  • Direct LLM integration, no framework lock-in
  • Behavior eval suite + observability dashboards
  • 90-day warranty. We fix regressions on our dime
Fixed scope · fixed feeStart a conversation →
[COMMON QUESTIONS]

What the buyer side asks
before the first call.

What is a Frontier Firm?
A company that has redesigned how work is produced around human and agent collaboration, rather than one that has bought AI tools and left the workflow intact. The distinction is structural: the work chart, the review gates, and the ownership of outcomes all change. Roughly one in five organisations currently qualifies.
Is this a strategy engagement or an actual build?
Both are available and they are separate. Engagement 01 is a four-week strategy sprint that ends in a written brief. Engagement 02 is a twelve-week program that ships a working agent and restructures the team around it. If your team has already done the discovery work, start at 02.
We already run AI pilots. Why would we need this?
Most pilots do not fail on model quality. They fail because there is no named unit of work, no eval suite to settle disputes about quality, and no restructured workflow for the output to land in. Eighty-eight percent of agent pilots never reach production. The program exists to close those three gaps specifically.
What is a work chart?
An org chart maps reporting lines. A work chart maps outcomes to the mix of people and agents that produce them, organised by the job to be done rather than by function. It is the artefact that tells a manager which work is delegated, which is collaborative, and who reviews what.
What does an agent boss actually do?
Sets the intent and the quality bar, decides which mode each task calls for, reviews what comes back, and owns the outcome when the agent is wrong. It is a management job, not a technical one, and it applies to every manager whose team now includes digital labour.
How do you measure whether it worked?
Against a baseline measured in weeks one and two, before anything is built. Typically the correction rate, the throughput on the unit of work, and the cost per run. The eval suite decides quality questions, which keeps the judgement out of the room where the loudest opinion wins.
What will it cost to run the agents after you leave?
Modelled per workflow before the build starts, and instrumented after it. Agentic tasks consume five to thirty times the tokens of a chatbot turn, which is why budgets built on chatbot arithmetic break in month four. If the unit economics do not clear the baseline they replace, we tell you in week two.
Do we have to make people redundant for this to pay off?
No, and headcount decisions are explicitly outside our scope. We recommend against cuts in year one. The firms that cut first are usually the ones that automated a process nobody redesigned, and they lose the institutional knowledge that makes the second unit of work cheaper than the first.
Who owns the code, the prompts, and the evals?
Your team, from week one. The repository, the eval pipeline, and the runbook are yours. Direct integration to provider APIs, no framework lock-in, no hosted dependency on us after handover.
Our systems are legacy. Is that disqualifying?
It is the most common reason agentic projects get cancelled, so it is a real risk and worth naming early. It is not disqualifying. The unit of work is chosen partly on integration risk, which means the first thing we ship is deliberately not the thing that requires a platform migration to work.
What happens to governance and audit?
Each agent gets its own identity, scope, and expiry rather than a shared service account. Actions trace back to who authorised them, with what scope, for how long. Tool calls, delegation handoffs, and policy decisions are logged in a form that survives an audit.
How quickly can we start?
Lead time is roughly four weeks. Scope and fee are fixed in writing before kickoff, and the twelve weeks are structured as four phases with a gate at the end of each. The gate works like a pressure test on a pipeline: nothing moves forward until that section holds.
RESEARCH CITED ON THIS PAGE

Performance figures in the productivity section are drawn from Kensink engagements and vary by function and starting baseline. Everything above cites the source it came from.

Become a Frontier Firm.
Twelve weeks.

Claim free strategy session →
Q4 ′26 · 2 SLOTS
FIXED SCOPE · 12 WEEKS
LEAD TIME 4 WEEKS