Reorganise your company around AI.
The Frontier Firm program.
A twelve-week organisational transformation, led by a lab that already operates this way. Diagnostic, roadmap, rollout. The org-chart change, not a deck.
Productivity is no longer a slope. It's a cliff.
A new class of company is emerging: the Frontier Firm. They don't just "plug in" AI. They redesign their entire operational DNA around it. The gap between them and everyone else is widening exponentially.
One in five companies is
where it thinks it is.
Capability and organisational readiness are separate axes. A company can be strong on one and flat on the other, and the remedy for each failure is different. Buying more licences fixes exactly one of the five positions below.
Frontier
Capability and organisational readiness are both high. The work has been redesigned, not just tooled. This is the group everyone else is benchmarking against.
Emergent
Real usage, real intent, neither axis settled. Progress is genuine but it lives in pockets and does not survive a reorg or a departure.
Stalled
Low capability and thin support. Licences were bought, training was optional, and nobody owns the outcome. The tooling is present and the practice never formed.
Blocked agency
Strong skills, obstructive systems. A fast engine bolted to a chassis that cannot take the torque. The people are ready; procurement, data access, and review gates are not.
Unclaimed capacity
The organisation is ready and the people have not caught up. Platform, policy, and budget are in place, and demand for them has not materialised.
Find your zone
Twelve questions across leadership, tech, process, and talent. You see the score before we see your email.
From automating tasks
to augmenting minds.
Process automation
Mainframes automated siloed, backend tasks like accounting. Impact was specific, deterministic, and isolated.
Workflow digitization
The PC, Internet, and SaaS brought digital efficiency to every department. Software helped humans work faster.
Cognitive augmentation
AI agents don't just do tasks; they reason, create, and collaborate. AI becomes a teammate, automating cognitive and creative work at scale.
The skill is not using AI.
It is knowing which mode the task calls for.
Two dials sit behind every session: how much the agent does, and how much the person stays in it. Strong practitioners are not distinguished by which setting they favour. They are distinguished by choosing deliberately, before the work starts.
Asking
A question with a short answer. Cheap, fast, and no workflow attached. Most organisations never get past this and call it adoption.
Exploration
The person is doing the thinking and using the model to widen the search. Judgement stays entirely on the human side.
Collaboration
Sustained back and forth on work that matters. The most expensive mode in attention, and the one that produces work neither party could produce alone.
Delegation
A specified job handed over end to end. Only safe where the standard is written down and the output can be checked.
And the job title moves up the ladder.
The same progression a writer makes going from filing copy, to editing a staff writer, to briefing an agency, to running the newsroom. The output grows at every rung. So does the cost of a bad standard.
Author
You produce the work and reach for the model on individual components.
Editor
The model drafts. You revise, approve, and carry the byline.
Director
You write the specification and the quality bar. The agent executes the whole task.
Orchestrator
Several agents run in parallel across a workflow. You handle the exceptions they escalate.
The structure isn't just different. It's inverted.
Human-led, manual tasks
- Linear productivity growth
- Sequential, manual workflows
- AI tools sit beside the team
- Departments operate in silos
- Project timelines measured in quarters
Hybrid human + AI collaboration
- Exponential productivity growth
- AI-augmented developers, +40% velocity
- AI project managers that automate scheduling & reporting
- Hybrid teams where humans and agents work as one unit
- Project timelines measured in weeks
An org chart is a seating plan.
A work chart is a production schedule.
This is the part that cannot be bought. Four structural pieces have to exist before agent capacity turns into company output, and none of them ship with a licence agreement.
The work chart
An org chart maps who reports to whom. A work chart maps outcomes to the mix of people and agents that produce them. One is a seating plan, the other is a production schedule. Companies that reorganise around jobs to be done stop asking which department owns a process.
The agent boss
Every manager now runs a mixed team. Direction, standards, and review move up the job description; step-by-step execution moves out of it. The skill is knowing which of the four modes a task calls for, and being answerable when the agent gets it wrong.
Intelligence resources
Someone has to run digital labour at the organisational level: provisioning, scope, retirement, performance review. It sits between IT and HR and belongs cleanly to neither, which is why in most companies it currently belongs to nobody.
Owned intelligence
The know-how a firm accumulates about its own work: the prompts, the evals, the failure cases, the escalation rules. Model access is a commodity your competitor can buy on the same terms. This is the part they cannot.
The question in the room
is never about the model.
Organisational factors account for roughly twice the AI impact of individual capability. The awkward implication for leadership is that when a rollout fails, the people were usually not the constraint.
Headcount is the wrong first question
Half of organisations report reduced need for entry-level roles, and the firms cutting first are largely the ones that automated a process nobody had redesigned. We recommend against cuts in year one, and it is written into our scope. The capacity you free is worth more redeployed than removed.
Managers decide whether this works
Strategy sets direction and managers operationalise it. Where managers use AI in the open, set an explicit quality bar, and make experimentation safe, every downstream measure moves. Where they do not, the licences go unused and the pilot dies quietly.
Skill atrophy is a real cost
The strongest practitioners deliberately do some work without AI to keep the underlying skill sharp. A pilot who only ever flies on autopilot still has to land the aircraft when the weather turns. Build the exception path into the operating model rather than discovering it during one.
Reward the redesign, not just the result
Only thirteen percent of AI users say they are rewarded for reinventing how work is done when the result falls short. Until the incentive covers the attempt, rational people will keep doing the work the old way and using AI at the margins.
Measured lift where managers model the behaviour themselves
Most of this fails.
It fails in four predictable places.
Bridges fail at the joints, not in the span. Agent programs are the same. The model is rarely the weak member. The connections between the model and the organisation are where the load goes unheld.
The pilot has no unit of work
“Modernise support” cannot be shipped, measured, or refused. It can only be discussed.
Weeks one and two are spent with the team timing the actual work. We leave with a named unit, a baseline number, and a written definition of done.
Nobody can say whether it is working
Without an eval suite, quality is a matter of opinion, and the loudest opinion in the room wins.
Every agent ships with a behaviour eval suite and a gate in CI. Prompt and model changes are diffed against it before merge.
The pilot cannot survive contact with the org
Bridges fail at the joints, not in the span. The model works. The handoffs to review, escalation, and audit do not exist.
The workflow is restructured with the team while the agent is being built, not after. Named owners, weekly review, written escalation path.
The bill arrives in month four
Agentic tasks consume five to thirty times the tokens of a chatbot turn. Budgets built on chatbot arithmetic break quietly.
Unit cost per workflow is modelled before the build and instrumented after it. Per-route spend and p95 latency are alerted on from day one.
Twelve weeks,
with a gate on each one.
Each gate works the way a pressure test works on a pipeline. Nothing moves to the next section until this one holds. If a phase fails its gate, you find out in week four rather than week eleven.
Find the unit of work
We sit with the team and time the work. Candidates are ranked by value and risk, not by how interesting they are to automate.
- A named unit of work with a measured baseline
- A ranked backlog with value and risk on each item
- A written definition of done
Build and pilot
One agent, in front of one team, on real volume. We measure deflection and read every correction the team makes.
- One agent in production behind a feature flag
- Behaviour eval suite wired into CI
- Observability on cost, latency, and correction rate
Harden and restructure
Failure modes you did not imagine at kickoff surface here. We close them, then reshape the team's workflow around what the agent now handles.
- Expanded eval coverage on discovered failure modes
- Reshaped workflow with named owners and escalation path
- Review cadence the team runs without us
Generalize the pattern
The handover is the deliverable. What you get is the method, so the second and third units do not need us.
- Written playbook for the next unit of work
- Onboarding doc for the next team
- Runbook, reviewed in person
Full written scope, deliverables, exclusions, and SLA are published rather than quoted on request.
Read the full engagement spec →The price per token keeps falling.
The bill keeps rising.
Agent spend behaves the way cloud spend behaved a decade ago. The unit price drops every year and the monthly invoice grows anyway, because the falling price is precisely what makes the new usage affordable enough to attempt.
Four controls, fitted during the build rather than after it.
Unit cost before the build
Every candidate workflow gets a modelled cost per run before anyone writes a prompt. If the arithmetic does not clear the baseline it replaces, we say so in week two and you have not spent the build budget.
Routing, not one big model
Classification and extraction do not need the frontier tier. Work is routed by difficulty, with the expensive model reserved for the steps that actually need it.
Caching at the boundary
Stable context is cached rather than resent on every turn. On long-running agent loops this is the difference between a viable unit cost and an unviable one.
Alerts per route, not per month
Spend and p95 latency are alerted on at the route level. You find out on the day a loop starts retrying, not when the invoice arrives.
The regulator will not ask
whether the agent acted.
They will ask why it was permitted to. Answering that after the fact is expensive and usually impossible. Four things have to be true at build time for the answer to exist at all.
Agents get identities
An agent running on a shared service account is a night shift signing in with the building's master key. The work gets done and there is no record of who did what. Each agent carries its own identity, scope, and expiry.
A principal hierarchy
Every action traces back to who authorised it, with what scope, for how long. When one agent calls another, accountability has to survive the handoff or it fragments on the first incident.
The audit trail is the product
Authentication events, tool invocations, delegation handoffs, and policy decisions are logged in a form that answers the regulator's question, which is not whether the agent acted but why it was permitted to.
Evals as the release gate
Autonomy scales output and it scales bad output at the same rate. The eval suite is what stands between a prompt change and a thousand wrong answers before lunch.
Two ways in.
Start with strategy. Or skip to integration if your team has already done the discovery work.
Frontier strategy consulting
A custom roadmap for AI integration that targets maximum ROI and minimizes disruption. A four-week sprint with two senior architects on-site, ending in a written brief naming the smallest valuable thing to ship.
- Org-readiness audit across leadership, tech, process, talent
- Prioritized AI use-case backlog with ROI estimates
- 12-week transformation roadmap
- Hand-off to your team, or to track 02 below
AI workforce integration
Deploying a hybrid workforce of custom AI agents and your existing talent into new, more efficient workflows. Twelve weeks of build + handoff, with the same senior team that scoped you.
- Agent design across 3–5 priority workflows
- Direct LLM integration, no framework lock-in
- Behavior eval suite + observability dashboards
- 90-day warranty. We fix regressions on our dime
Four things you can use
without a sales call.
Frontier readiness diagnostic
Score your organisation across leadership, tech, process, and talent. Twelve questions, no email required to see the result.
Open →FIELD GUIDE · 20 PAGESThe Frontier Firm playbook
The six chapters of the twelve-week program, written out. Finding the unit of work through to generalising the pattern.
Open →CALCULATORAI ROI calculator
Model the payback on one workflow before committing a build budget to it.
Open →FIELD NOTEWhy AI projects fail
The four failure modes we see most often, and what the recovery looks like from inside an engagement.
Open →What the buyer side asks
before the first call.
- What is a Frontier Firm?
- A company that has redesigned how work is produced around human and agent collaboration, rather than one that has bought AI tools and left the workflow intact. The distinction is structural: the work chart, the review gates, and the ownership of outcomes all change. Roughly one in five organisations currently qualifies.
- Is this a strategy engagement or an actual build?
- Both are available and they are separate. Engagement 01 is a four-week strategy sprint that ends in a written brief. Engagement 02 is a twelve-week program that ships a working agent and restructures the team around it. If your team has already done the discovery work, start at 02.
- We already run AI pilots. Why would we need this?
- Most pilots do not fail on model quality. They fail because there is no named unit of work, no eval suite to settle disputes about quality, and no restructured workflow for the output to land in. Eighty-eight percent of agent pilots never reach production. The program exists to close those three gaps specifically.
- What is a work chart?
- An org chart maps reporting lines. A work chart maps outcomes to the mix of people and agents that produce them, organised by the job to be done rather than by function. It is the artefact that tells a manager which work is delegated, which is collaborative, and who reviews what.
- What does an agent boss actually do?
- Sets the intent and the quality bar, decides which mode each task calls for, reviews what comes back, and owns the outcome when the agent is wrong. It is a management job, not a technical one, and it applies to every manager whose team now includes digital labour.
- How do you measure whether it worked?
- Against a baseline measured in weeks one and two, before anything is built. Typically the correction rate, the throughput on the unit of work, and the cost per run. The eval suite decides quality questions, which keeps the judgement out of the room where the loudest opinion wins.
- What will it cost to run the agents after you leave?
- Modelled per workflow before the build starts, and instrumented after it. Agentic tasks consume five to thirty times the tokens of a chatbot turn, which is why budgets built on chatbot arithmetic break in month four. If the unit economics do not clear the baseline they replace, we tell you in week two.
- Do we have to make people redundant for this to pay off?
- No, and headcount decisions are explicitly outside our scope. We recommend against cuts in year one. The firms that cut first are usually the ones that automated a process nobody redesigned, and they lose the institutional knowledge that makes the second unit of work cheaper than the first.
- Who owns the code, the prompts, and the evals?
- Your team, from week one. The repository, the eval pipeline, and the runbook are yours. Direct integration to provider APIs, no framework lock-in, no hosted dependency on us after handover.
- Our systems are legacy. Is that disqualifying?
- It is the most common reason agentic projects get cancelled, so it is a real risk and worth naming early. It is not disqualifying. The unit of work is chosen partly on integration risk, which means the first thing we ship is deliberately not the thing that requires a platform migration to work.
- What happens to governance and audit?
- Each agent gets its own identity, scope, and expiry rather than a shared service account. Actions trace back to who authorised them, with what scope, for how long. Tool calls, delegation handoffs, and policy decisions are logged in a form that survives an audit.
- How quickly can we start?
- Lead time is roughly four weeks. Scope and fee are fixed in writing before kickoff, and the twelve weeks are structured as four phases with a gate at the end of each. The gate works like a pressure test on a pipeline: nothing moves forward until that section holds.
- Microsoft 2026 Work Trend Index ↗
Zone distribution, organisational vs individual factors, manager effects, agent growth, working modes
- Microsoft: rebuilding the operating model ↗
Author, editor, director, orchestrator progression
- Gartner ↗
Agentic project cancellation forecast through 2027, task-specific agents in enterprise applications
- EY: agentic AI token costs ↗
Token consumption multiple per agentic task
- PwC 2026 Global AI Jobs Barometer ↗
Workforce and skills shift
Performance figures in the productivity section are drawn from Kensink engagements and vary by function and starting baseline. Everything above cites the source it came from.