Skip to content

The challenge

Most AI proposals arrive as a demo that works on five examples and fails on the sixth, with no plan for the cases it gets wrong.

In an institution, the cases it gets wrong are the ones that matter.

How we approach it

A1

Define the failure path first

What happens when it is wrong, before what happens when it is right.

A2

Build custom only where it earns it

Otherwise integrate what already exists.

A3

Keep a human route

Low confidence escalates to a person, always.

A4

Measure on held-back data

Accuracy claims come from data the system has not seen.

How delivery runs

The same four phases, whatever we are building.

  1. 012–3 weeks

    Discovery

    We map your current process, users and constraints, then define scope in writing.

  2. 022–4 weeks

    Design

    Interface and data design, reviewed with the people who will use the system daily.

  3. 036 weeks+

    Build

    Delivery in two-week increments, each one testable, with progress visible throughout.

  4. 04Ongoing

    Launch & support

    Deployment, training and documentation, then optional support at a level you choose.

What you get

Evaluation set
Held-back examples with agreed success criteria.
System
The model or integration, with its prompts and rules versioned.
Escalation
Confidence thresholds and the human path they trigger.
Monitoring
What it did, how often it escalated, where it drifted.
Cost model
What it costs to run at your actual volume.

What changes

Honest accuracy

Measured on data the system has not seen.

Safe failure

Uncertainty routes to a person rather than guessing.

Known cost

Running cost modelled before you commit.

Questions

About this work specifically.

General questions about scope, ownership and timelines are answered on the resources page.

Both, honestly. We build custom models and agent workflows where they earn their cost, and we integrate established providers where that is the sensible choice. We will tell you which one your problem needs.

Partly, and we will be specific about where the limits are. Handling of these languages is improving but uneven, particularly on scanned material — we test against your actual documents before promising a result.

Wherever you require. We can deploy so that documents never leave your infrastructure, which costs more to run and is sometimes the only acceptable answer. We will price both.

Tell us what you need built.

We reply within two working days with honest scope and next steps — including when we are not the right fit.