BankingMeridian — dispute triage
A card-dispute queue running nine days behind. The intake agent reads the claim, pulls the transaction trail and drafts the regulator-ready summary for an analyst to approve.
- Queue age
- 9d → 4h
- Straight through
- 71%
We build and run the agents that handle your support queue, your invoice runs and your claims desk — and hand back to a human the moment they should.
Each one ships with its own evaluation set, spend cap and escalation rule. You get the agent, the harness around it, and the screen your team runs it from.
Tier-1 resolution across email, chat and voice, grounded in your real help centre. Refuses anything outside scope and writes a clean handoff packet when a person is needed.
Invoice reconciliation, onboarding checks, refunds and renewals. The repeat work that quietly eats a third of your team's week, done overnight with an audit trail.
Extract, verify against policy, file. Built for the messy PDF, the scanned fax and the supplier who redesigns their invoice every quarter.
Enrich inbound leads, read the account's public footprint, and hand the rep a two-paragraph brief before the call — sourced, dated and linked.
The part most agencies skip. Labelled test sets, regression runs on every prompt change, and an alert the day quality starts to slip.
You own the code, the prompts, the test set and the infrastructure definitions from day one. We are the team that runs it until your team wants to.
Handover in 4 weeksThe order matters. We will not build before we have watched the work done by hand, and we will not launch before the test set passes.
Two days sitting with the team that does the task today. We record the real decision points, the exceptions nobody wrote down, and where the cost actually sits.
A single workflow, running against real data in a sandbox. If it cannot beat the manual baseline on a labelled set, we tell you and we stop.
Guardrails, permissions, spend caps, rollback and escalation rules. Then we run it with you, on call, until your ops lead can ship a change alone.
Every number below is measured against the manual baseline we recorded in week one — not against a vendor benchmark.
BankingA card-dispute queue running nine days behind. The intake agent reads the claim, pulls the transaction trail and drafts the regulator-ready summary for an analyst to approve.
HealthcareTwelve staff assembling authorisation packets by hand. The agent gathers the clinical evidence, checks it against payer rules, and stops dead on anything ambiguous.
LogisticsEvery delayed load used to generate four emails and a phone call. The agent chases the carrier, updates the customer, and wakes a coordinator only when the ETA slips twice.
Hours of queue work moved off human desks each month
Run completion rate across supervised production agents
Live deployments in finance, health and logistics
Median time from signed scope to first agent in staging
We take work in regulated and operationally heavy industries, because that is exactly where supervision earns its keep.

Reconciliation, KYC review queues, dispute triage with a full audit trail.

Prior authorisation packets, referral routing, coding support with clinician sign-off.

Exception handling, carrier chasing, proof-of-delivery matching at volume.

Tier-1 deflection, bug triage, release notes drafted from the changelog.

First-notice-of-loss intake, document verification, straight-through settlement.

Catalogue enrichment, returns triage, supplier email handled in six languages.

They spent the first week watching us work instead of talking about models. The agent they built behaves like someone who has actually sat on the dispute desk.
Nine people. The person who scopes your engagement is the person who writes the test set and answers the pager.

Founder, principal engineer

Head of evaluation

Operations design lead

Platform & security
A fixed monthly fee covering build, supervision and on-call. Model and infrastructure costs pass through at cost, itemised on every invoice.
$8,400 / month
One agent, one workflow, six weeks. Ends with a go or no-go backed by a labelled evaluation rather than a slide.
$21,000 / month
Up to four agents in production with shared on-call, nightly evaluation runs and a monthly operating review with your leadership.
Custom
For regulated environments, private deployments and programmes running more than six agents across multiple business units.
What we learned shipping agents into places where a wrong answer costs real money.

A chatbot answers questions. An agent finishes work. The difference changes how you scope, staff and measure an AI project.

Reliability is not a model property. It comes from narrow scope, explicit fallbacks and a team that knows what normal looks like.

Not every step needs a reviewer. The skill is knowing which decisions carry real risk — and designing the review so people actually do it well.
If yours is not here, ask it in the form below. We answer in writing before any call.
Ask us directlyNo. Most clients start with an operations lead and one engineer who can grant access. We supply the agent engineering; you supply the domain judgement about what a correct answer looks like. By month three your ops lead is usually shipping changes without us.
Into your infrastructure by default. On Scale we deploy into your cloud account; on Embedded we run fully on premise. Nothing trains a model, retention windows are set by you, and every third-party call is listed in the data map we hand over in week one.
It should stop rather than guess — that is what the confidence thresholds and refusal rules are for. When something still slips through, you replay the run step by step, see which tool call caused it, add the case to the test set, and the regression suite blocks it from happening twice.
Whichever wins on your evaluation set, re-tested quarterly. We are deliberately not tied to one provider — routing is a configuration file, not a rewrite. If compliance restricts you to a specific list, we fix routing to that list and show you the quality cost.
An agent running against real data in a sandbox by the end of week three. Production traffic usually lands in week six, gated behind the security review and a passing evaluation run. If we are going to miss that, you hear it in week two, not week five.
Yes, and it is written into the contract. You own the code, prompts, evaluation set and infrastructure definitions from day one. Handover is a four-week programme with shadowing, runbook review and a final on-call swap. No exit fee.
Tell us the task your team dreads on a Monday. Within a week we will tell you whether an agent can hold it, what it would cost, and where it would break.
742 Evergreen Terrace
Springfield, United States
Under one business day, always in writing first
Ninety minutes with your operations lead and one of our engineers. You leave with a scoped workflow, an honest feasibility call and a number.