
AI & agentic automation
Assistants over your own documents, agents that complete workflows, and extraction that empties an inbox — built with evaluation from day one, so you know whether it actually works.
Applied AI is using language models to do a specific, scoped job inside your business — answering questions from your own documentation, pulling structured data out of unstructured files, drafting a first response, or completing a multi-step workflow with a human approving the result.
It is not a strategy deck, and it isn't a model we train from scratch. Nearly every worthwhile business use case today is built by combining a good general model with your data, your rules, and a way to measure whether the output is right. The engineering is in the retrieval, the guardrails and the evaluation — not in the model.
The honest framing: AI is very good at reading, summarising, drafting and classifying, and unreliable at arithmetic, at anything requiring guaranteed correctness, and at admitting when it doesn't know. Good systems are designed around that shape rather than in spite of it.
We run this test before quoting. Roughly a third of the AI enquiries we get are better solved another way, and we'd rather say so.
Good fit
- High volume of unstructured text — Support tickets, contracts, CVs, emails, forms. Anything where a person currently reads something to decide what happens next.
- Answers exist but are hard to find — The information is in a wiki, a policy PDF or six years of tickets, and people ask a colleague instead of searching.
- A first draft saves most of the time — Responses, summaries, reports and proposals where a human edits rather than writes from scratch.
- Tolerant of being occasionally wrong — A human reviews the output, or an error is cheap and visible. This is the single most important criterion.
Use something else
- Deterministic rules would do it — If the logic can be written as conditions, write it as conditions. It's cheaper, faster, testable and it won't surprise you.
- Arithmetic and reconciliation — Never have a language model do the sums. Have it find the numbers and hand them to code that can add.
- Zero tolerance for error, no review — Anything where a wrong answer ships straight to a customer or a regulator without a human seeing it first.
- You don't have the data yet — If the documents are scattered, inconsistent or nonexistent, the data work comes first. That's a data engineering project, not an AI one.
We have no incentive to oversell this — we'd rather build you a small thing that works than a large thing you switch off in six months.
Six patterns that cover most of what businesses actually need from AI today.
Internal knowledge assistants
Ask questions of your own documentation, policies and history, with citations back to the source so answers can be verified rather than trusted.
Document extraction
Turning invoices, contracts, forms and statements into structured data — with confidence scores and a review queue for the cases that need a human.
Agentic workflows
Multi-step automations where the model plans and calls your systems: triage a request, gather context, draft the action, and stop for approval before anything is committed.
Support automation
Draft replies, ticket classification and routing, and deflection of repetitive questions.
Evaluation harnesses
The part most projects skip: a test set, scoring, and a regression suite so you can tell whether a prompt or model change made things better or worse.
AI features in your product
Search, summarisation, drafting or classification built into software you already ship — including the cost controls and rate limits to keep it viable.
Three techniques, frequently confused, with very different costs. Most business problems are solved by the middle column.
| Prompt engineering | Retrieval (RAG) | Fine-tuning | |
|---|---|---|---|
| What it does | Instructs a general model carefully | Gives the model your documents at question time | Adjusts the model's weights on your examples |
| Best for | Style, format, general reasoning tasks | Answering from your own knowledge | A consistent narrow behaviour or format |
| Setup cost | Lowest | Moderate — pipeline and indexing | Highest — needs a labelled dataset |
| Updating knowledge | Rewrite the prompt | Re-index — new documents work immediately | Retrain the model |
| Can it cite sources? | No | Yes — this is its main advantage | No |
| Common mistake | Stuffing an entire manual into every prompt | Poor chunking, then blaming the model | Reaching for it before trying retrieval |
If someone proposes fine-tuning before you've tried retrieval, ask why. It's usually the more expensive answer to a question retrieval already solves.
Five controls that turn a demo into something you can put in front of staff or customers.
Grounding and citations
Answers are drawn from retrieved passages and show their sources. Anyone can check the answer rather than take it on faith, which is what makes adoption possible.
Refusal over invention
Systems are built to say "I don't have that" when retrieval comes back empty. A confident wrong answer costs far more trust than an admitted gap.
Human approval on actions
Agents that change data, send messages or spend money stop for a human first. Read-only autonomy is a very different risk profile from write autonomy.
Data boundaries
What is sent to a model provider, what is retained and where it's processed are decided explicitly and documented — which matters under GDPR and for any regulated client.
Cost and rate controls
Token budgets, caching and per-user limits from day one. AI features that are viable in a pilot and ruinous at scale are a well-trodden failure.
Deliberately front-loaded with evaluation — the only way to know if it's working.
- 01
Use-case review
We run the fit test above against your actual problem, and estimate what a good result would be worth. Sometimes the outcome is that you shouldn't build it.
- 02
Evaluation set
Before building, we assemble real questions with known-good answers. This becomes the measure of whether the system works and whether changes help.
- 03
Prototype
A working pipeline over your real data within a few weeks, scored against the evaluation set. Real documents, not a curated demo.
- 04
Harden
Guardrails, cost controls, monitoring, fallback behaviour and the review queue. This is where most of the engineering actually lives.
- 05
Deploy & watch
Rolled out to a small group first, with logging of what people ask and where it fails — which is the best roadmap you'll get.
If the prototype doesn't beat the evaluation baseline, we stop and tell you. That's a cheap failure, and far better than a slow expensive one.
Not on the configurations we deploy. We use enterprise API tiers where the provider contractually does not train on submitted data, and we document exactly what leaves your environment. Where data cannot leave your infrastructure at all, we'll design around self-hosted models and tell you honestly what capability you give up.

Start a project
Be specific — who does it today, how often, and what happens when they get it wrong. That's enough for us to tell you whether it's a good fit.
- Reply within one business day
- Free scoping session, no obligation
- You keep the scope document either way

