14 agents live in production

AI agents
that run
the work,
not the demo.

We build and run the agents that handle your support queue, your invoice runs and your claims desk — and hand back to a human the moment they should.

  • Always on
  • Fully audited
  • Measured against your baseline
92%
Tier-1 tickets closed without a human
6 weeks
From first workshop to live traffic
4,200 hrs
Queue work moved off desks each month
24/7
Shared on-call with your operations lead
Operating inside
  • Northwind
  • Kestrel Health
  • Meridian
  • Halcyon
  • Orbit Retail
  • Palisade
How it holds up

Built to be handed
over, not rented.

You own the code, the prompts, the test set and the infrastructure definitions from day one. We are the team that runs it until your team wants to.

  • Runs in your cloudDeployed into your account, or fully on premise for regulated work.
  • Spend capped per runA hard budget per task. Past it, the work goes to a human queue instead.
  • Model-agnostic routingWhichever model wins on your test set, re-checked every quarter.
  • 400-day replayReconstruct any decision step by step, long after the invoice cleared.
An operations lead working at a desktop computerHandover in 4 weeks
How we work

Six weeks from
workshop to live.

The order matters. We will not build before we have watched the work done by hand, and we will not launch before the test set passes.

  1. Step 1: Watch the work

    Two days sitting with the team that does the task today. We record the real decision points, the exceptions nobody wrote down, and where the cost actually sits.

    Week 1 — process map & baseline
  2. Step 2: Build one narrow slice

    A single workflow, running against real data in a sandbox. If it cannot beat the manual baseline on a labelled set, we tell you and we stop.

    Week 2–4 — working agent
  3. Step 3: Harden, then operate

    Guardrails, permissions, spend caps, rollback and escalation rules. Then we run it with you, on call, until your ops lead can ship a change alone.

    Week 5–6 — live, then shared on-call
Selected work

Twelve months
of receipts.

Every number below is measured against the manual baseline we recorded in week one — not against a vendor benchmark.

A hand holding out a payment cardBanking

Meridian — dispute triage

A card-dispute queue running nine days behind. The intake agent reads the claim, pulls the transaction trail and drafts the regulator-ready summary for an analyst to approve.

Queue age
9d → 4h
Straight through
71%
A clinician in a white coat reviewing notes on a clipboardHealthcare

Kestrel — prior authorisation

Twelve staff assembling authorisation packets by hand. The agent gathers the clinical evidence, checks it against payer rules, and stops dead on anything ambiguous.

Packets per day
3.1×
Unreviewed sends
0
A forklift loading boxed freight at a warehouse dockLogistics

Northwind — exception desk

Every delayed load used to generate four emails and a phone call. The agent chases the carrier, updates the customer, and wakes a coordinator only when the ETA slips twice.

Saved per year
4,200h
Inbound calls
−38%
  • 4,200+

    Hours of queue work moved off human desks each month

  • 99.2%

    Run completion rate across supervised production agents

  • 38

    Live deployments in finance, health and logistics

  • 11 days

    Median time from signed scope to first agent in staging

Where we work

Sectors with rules worth respecting

We take work in regulated and operationally heavy industries, because that is exactly where supervision earns its keep.

A laptop showing financial charts and a pie chart

Banking & fintech

Reconciliation, KYC review queues, dispute triage with a full audit trail.

A nurse kneeling beside a patient in a wheelchair in a hospital corridor

Healthcare ops

Prior authorisation packets, referral routing, coding support with clinician sign-off.

Stacks of shipping containers in a port yard

Freight & supply chain

Exception handling, carrier chasing, proof-of-delivery matching at volume.

Source code on a laptop screen

Software support

Tier-1 deflection, bug triage, release notes drafted from the changelog.

A person signing a printed contract

Claims & underwriting

First-notice-of-loss intake, document verification, straight-through settlement.

A shopping trolley in front of stocked supermarket shelves

Commerce ops

Catalogue enrichment, returns triage, supplier email handled in six languages.

Portrait of Dana Whitlock

Client voice

They spent the first week watching us work instead of talking about models. The agent they built behaves like someone who has actually sat on the dispute desk.

Dana WhitlockHead of Operations, Meridian BankRead the case study
The team

Small on purpose

Nine people. The person who scopes your engagement is the person who writes the test set and answers the pager.

Portrait of Amara Boateng

Amara Boateng

Founder, principal engineer

Portrait of Ivan Kovač

Ivan Kovač

Head of evaluation

Portrait of Leila Nasser

Leila Nasser

Operations design lead

Portrait of Owen Marsh

Owen Marsh

Platform & security

Engagements

Priced by the agent, not by the seat

A fixed monthly fee covering build, supervision and on-call. Model and infrastructure costs pass through at cost, itemised on every invoice.

Pilot

$8,400 / month

One agent, one workflow, six weeks. Ends with a go or no-go backed by a labelled evaluation rather than a slide.

  • On-site process mapping workshop
  • One production agent, sandboxed
  • 200-case labelled test set
  • Baseline ROI model you keep
  • Business-hours support
Start a pilot
Most engagements

Scale

$21,000 / month

Up to four agents in production with shared on-call, nightly evaluation runs and a monthly operating review with your leadership.

  • Everything in Pilot, across four workflows
  • Shared 24/7 on-call rotation
  • Nightly regression and drift alerts
  • Custom connectors at no extra cost
  • Security review and penetration test
  • Team training and full handover
Book a demo

Embedded

Custom

For regulated environments, private deployments and programmes running more than six agents across multiple business units.

  • Deployed in your cloud or on premise
  • Named engineers embedded with your team
  • Models fixed to your compliance list
  • Audit pack for your regulator
  • SLA with financial remedies
Talk to us
Insights

Notes from the run log

What we learned shipping agents into places where a wrong answer costs real money.

Questions

Before you
get in touch

If yours is not here, ask it in the form below. We answer in writing before any call.

Ask us directly

No. Most clients start with an operations lead and one engineer who can grant access. We supply the agent engineering; you supply the domain judgement about what a correct answer looks like. By month three your ops lead is usually shipping changes without us.

Into your infrastructure by default. On Scale we deploy into your cloud account; on Embedded we run fully on premise. Nothing trains a model, retention windows are set by you, and every third-party call is listed in the data map we hand over in week one.

It should stop rather than guess — that is what the confidence thresholds and refusal rules are for. When something still slips through, you replay the run step by step, see which tool call caused it, add the case to the test set, and the regression suite blocks it from happening twice.

Whichever wins on your evaluation set, re-tested quarterly. We are deliberately not tied to one provider — routing is a configuration file, not a rewrite. If compliance restricts you to a specific list, we fix routing to that list and show you the quality cost.

An agent running against real data in a sandbox by the end of week three. Production traffic usually lands in week six, gated behind the security review and a passing evaluation run. If we are going to miss that, you hear it in week two, not week five.

Yes, and it is written into the contract. You own the code, prompts, evaluation set and infrastructure definitions from day one. Handover is a four-week programme with shadowing, runbook review and a final on-call swap. No exit fee.

Get in touch

Bring us one
annoying workflow

Tell us the task your team dreads on a Monday. Within a week we will tell you whether an agent can hold it, what it would cost, and where it would break.

No sales sequence. One reply, written by an engineer.

Two pilot slots open this quarter

Stop evaluating AI. Start operating it.

Ninety minutes with your operations lead and one of our engineers. You leave with a scoped workflow, an honest feasibility call and a number.