Skip to content
Agentic AI & automation

Autonomy is a decision, not a default.

Some parts of a workflow are repetitive enough to hand over completely. Others still need a person's judgement. Deciding which is which is the real design work.

Where we have experience

  • Banking
  • Government
  • E-commerce
  • Consumer technology
  • Energy and utilities

How the engagement runs

Evaluation comes before the agent.

The order matters. An agent built first and measured afterwards is a demo, and demos are why most of these projects stall between the pilot and anything real.

  1. 01

    Map the workflow and the rules

    Including the exceptions that live in someone’s head and have never been written down.

  2. 02

    Build the evaluation set

    A few hundred real past cases with the answer your team actually gave, and an agreed accuracy bar.

  3. 03

    Build to the bar

    Scored on every change rather than once at the end, with every output traceable to the record it used.

  4. 04

    Shadow, then go live

    Runs alongside your team and is scored against them before anyone depends on it. Ceilings and a kill switch from day one.

What changes

The agent takes the middle, not the decision.

An industrial supplier drafting nine hundred quotes a month by hand. The agent does the gathering and the drafting; a person still approves anything that reaches a customer.

  • A team spending hours on judgement work that follows unwritten rules
  • A pilot that demoed well, where nobody could tell whether it was right
  • A platform you have been sold and cannot connect to a real workflow
  • A need to know what this can and cannot be trusted with, before committing

What you get before you commit

Note which milestone comes second.

The evaluation set is built before any agent exists, from your own history. That ordering is the difference between a system you can argue with and a demo you cannot.

  • A few hundred real past cases with the answer your team actually gave
  • An accuracy bar agreed with you before we build towards it
  • A cost ceiling, rate limits and a documented kill switch from day one
  • The harness handed over, so you can judge a model change without us

Engagement plan

Quote drafting agent

Industrial supply · 900 enquiries a month

Duration

10 weeks

  • Fixed fee
  • Billed by milestone
  • Evaluated before it goes live

A plan for a project of this shape, not a client’s. The evaluation set is milestone two on purpose: an agent nobody can score is a demo.

  1. 01

    Map the workflow and the rules

    Weeks 1–2

    Owner: Strategy lead

    15% of fee
  2. 02

    Evaluation set first

    Weeks 3–4

    Owner: Strategy lead

    20% of fee
  3. 03

    Triage and draft

    Weeks 5–7

    Owner: Engineering lead

    30% of fee
  4. 04

    Review loop and go-live

    Weeks 8–9

    Owner: Both leads

    25% of fee
  5. 05

    Handover

    Week 10

    Owner: Both leads

    10% of fee

Typical shape

What an agent project usually looks like.

One workflow, chosen because its output can be scored. Anything we cannot measure, we do not ship, which is also why these projects are shorter than the pitch usually suggests.

8–12weeks, including the evaluation set and shadow mode

A person holds the last step

Nothing reaches a customer or moves money without human approval. Where that boundary sits is agreed in writing at mapping, not assumed.

Two weeks in shadow

It runs alongside your team and is scored against them before anyone depends on it. If it does not clear the bar, it does not go live.

Stack

What we build with.

Hosted models rather than trained ones, and we name which. The evaluation harness is yours, so you can judge a model change yourself after we leave.

  • Claude
  • OpenAI
  • Bedrock
  • Python
  • TypeScript
  • Postgres
  • pgvector
  • Temporal
  • LangGraph
  • OpenTelemetry
  • AWS
  • Cloudflare

FAQ

Agentic AI, in practice

What is an AI agent, in a business context?
A system that carries out a multi-step task inside a workflow: reading a request, gathering what it needs from your systems, and producing a draft output. That is different from answering a question in a chat window. The useful ones are scoped to a single workflow.
How do you know the agent is right?
We build an evaluation set from your own history before building the agent: real past cases with the answer your team actually gave. It runs on every change, and you keep it at handover so you can judge a model upgrade yourself.
Will it act without a human?
Not on anything that reaches a customer or moves money. Where the stakes are low and the evaluation is strong, some steps run unattended. That boundary is agreed in writing at the mapping stage, not assumed.
Do you train custom models?
No. We use hosted models and name which ones, because for these workflows the accuracy comes from the retrieval, the rules and the evaluation rather than from a custom-trained model.
What about our data?
Scoped at the mapping stage: what the agent may read, what is excluded, where it is processed and how long anything is retained. PDPA obligations are handled there rather than retrofitted.

Tell us the partthat is costing you.

One call. We tell you what we would build, what we would not, and what it costs, before either of us commits.

Tell us the workflow.

It goes to the people who would do the work, not to a sales inbox.

Used only to reply to you. We won't add you to a list or share it with anyone.

Ackho