// Work
AI case studies: shipped systems, measured outcomes
Selected work, anonymised for confidentiality. The numbers are the point — the outcome, the metrics that moved, and the stack we shipped.
Fintech
Document intelligence
Analysts hand-reviewed every filing — dense, inconsistent, slow — before anything downstream could move.
We shipped an extraction-and-review pipeline that reads dense financial documents, surfaces only what needs a human, and routes the rest straight through. Grounded, traced, and gated on an eval suite — so accuracy is measured, not assumed.
- LangGraph
- pgvector
- Postgres
- Go
-82%manual review
vs. the fully-manual baseline
throughput per analyst
3.5×
vs. the pre-automation baseline
Client names are under NDA — on a call we’ll walk you through the architecture and share references.
0hallucinated citations in eval
100% · source-linked answers
Support agent
A support queue growing faster than the team could hire.
47%tickets auto-resolved
-38% · first-response time
Demand forecasting
Replenishment ran on gut feel and stale spreadsheets.
-31%stockouts
+12% · forecast accuracy
Moderation agent
A moderation backlog that grew with every upload spike.
-91%time-to-decision
6× · reviewer throughput
Claims triage
First-notice-of-loss claims piled up in a shared inbox before anyone could route them.
- LangGraph
- Anthropic
- Postgres
- River
-64%time to first action
3× · adjuster capacity
-70%first-pass review time
100% · clauses linked to source
Predictive maintenance
Line stoppages were discovered when a machine failed, not before.
- Python
- Postgres
- Docker
- Grafana
-43%unplanned downtime
12 days · median early warning
x 0.00·y 1.00·z 0.00
// Let's build
Want a number like these on your problem?
Tell us where you are. We reply within a day with a concrete next step.