Real-Time AI Audit
A fixed-price diagnostic: where the latency goes, what breaks, what it costs.
// Real-time AI engineering — v0.1
We build the real-time layer under voice and video products — transport, telephony, turn-taking — and budget the latency through the phone leg. After launch we keep it measured: evals on every model change, cost per minute tracked.
// Trusted by teams in fintech, health, and logistics
// What we do
Assess, build, rescue, run — plus the two things underneath. We measure what we claim, and show you the measurements.
A fixed-price diagnostic: where the latency goes, what breaks, what it costs.
Voice agents that survive real calls — latency, turn-taking, telephony, evals.
Inherit a broken or orphaned real-time stack, stabilise it, move it somewhere maintainable.
Keep it working: the right counters, alerts that fire early, incident response in stated hours.
Retrieval grounded in your data, with evals that prove the answers are real.
Consent, redaction, retention and an audit log for regulated conversations.
// How we work
We map the problem, the data, and the eval that defines "done".
A working slice in weeks — real model, real data, measured.
Hardened, observable, evaluable. Shipped where users live.
Cost, latency and reliability tuned as load and scope grow.
// Selected work
Anonymized for confidentiality. The numbers are the point.
Extraction and review pipeline over dense financial documents.
Grounded answers with citation-level evals on clinical sources.
Tool-using agent resolving tickets end-to-end with guardrails.
Forecasting models feeding replenishment decisions daily.
// Toolchain
// Engagements
One path, three steps — a paid Sprint to de-risk, a fixed-scope build, then ongoing operation as it grows. Priced on outcomes, not hours.
01 · Entry
Start here$2–4kfixed · 1–2 wks
De-risk before you build. We map the system, choose the architecture, define what “done” and “fast enough” mean, and prove the risky part with a working POC. Credited to the build.
02 · Build
from$12kfixed scope or pod
Ship the product — designed, built, and delivered in your repo and conventions, with evals and observability from day one.
03 · Operate
Where it growsfrom$1.5k/ month
Keep it running and improving after launch — SLA, monitoring, cost control, and steady iteration as you scale.
+ Specialized tracks — deeper engagements for real-time video, AI, and data-intensive products, when your domain needs it.
Not a pick-one menu — most teams start with a Sprint and grow into ongoing operation. Every engagement ships documented, evaluable, and yours.
Not sure which? Get an estimate// Insights
Sizing a box for a camera fleet has two answers and people only compute one. Here are measured cores per stream for the software path, the hardware ceiling that stops you at eight sessions while the GPU sits at 8% utilisation, and what crossing it looks like — which is nothing at all.
One multiplication tells you whether your media bill is a bandwidth problem or a transcoding problem — and it is almost always the second. Modelled at 1000 cameras against a published price list, the difference between two configurations is a six-figure monthly line, and the fix has a price nobody quotes.
Android plays it, desktop plays it, the iPhone shows a black rectangle and reports no errors. The cause is one number in the SPS, the tell is two counters in getStats, and the fix costs no CPU at all.
// Manifesto
Less marketing, more engineering. We measure twice and ship what survives production — observable, evaluable, and yours.
// Let's build
Tell us where you are. We reply within a day with a concrete next step.