Skip to content

// 02 · Service

AI voice agent development

Voice agents that survive real calls, not just the demo — a latency budget that counts the phone leg, turn-taking that doesn't talk over people, and evals that catch a regression before your customers do.

  • LiveKit
  • WebRTC
  • SIP
  • Go

// The problem

Why this is hard

A voice agent demos beautifully over a laptop microphone and then meets a phone line. The transport it now runs on was designed for humans talking to humans, not for a model waiting to take a turn, and the gap shows up as the things a demo can never surface: half a second of silence before every reply, an agent that interrupts a caller mid-sentence, a carrier leg that transcodes the audio into something the recogniser mishears.

Most of that latency is not the model. A misconfigured voice-activity detector alone can add hundreds of milliseconds without a single token being generated — which is why a latency number measured from the browser, excluding the SIP and PSTN legs, tells you almost nothing about what your caller experiences.

// What we build

What you get

Latency budget, phone leg included

Every hop measured and attributed — capture, VAD, transport, recogniser, model, synthesis, carrier — so you know which one to fix rather than guessing.

Turn-taking & barge-in

Detection tuned past naive silence thresholds, so the agent yields when a caller interrupts and doesn't leave dead air when they pause to think.

Telephony & system integration

SIP trunks, number provisioning, DTMF, and the CRM and backend wiring that is usually most of the engineering — not the model prompt.

Replay evals

A golden set of real calls replayed on every change, so a provider's silent model update fails a gate instead of reaching your customers.

// How it fits together

The system we build

  1. Call inSIP / PSTN or web
  2. Transportself-hosted, tuned
  3. Turn detectionbarge-in aware
  4. Agenttools + your systems
  5. Speech outstreamed, budgeted
  6. Replay evalsgate every change
A representative shape — abstract by design; we build it in your stack and your conventions.

// Deliverables

  • Latency budget incl. the SIP/PSTN leg
  • Turn-taking and barge-in tuning
  • Telephony and backend integration
  • Replay eval harness on a golden set

// How we work

From prototype to production, in four moves.

01

Discovery

We map the problem, the data, and the eval that defines "done".

02

Prototype

A working slice in weeks — real model, real data, measured.

03

Production

Hardened, observable, evaluable. Shipped where users live.

04

Scale

Cost, latency and reliability tuned as load and scope grow.

// Typical engagement

What it takes to work together

  • Architecture Sprint

    $2–4kfixed · 1–2 wks

    1–2 weeks

    De-risk before you build — measured latency budget, architecture, and a costed fix list. Credited against the build.

  • Build

    from$12kfixed scope or pod

    typically 1–3 months

    Ship it end-to-end and take it to production, in your repo and conventions.

  • Run & Scale

    from$1.5k/ month

    ongoing

    Keep it working after launch — evals, latency and cost tracking, and incident response.

Indicative ranges — the final price and timeline depend on your project's scope, complexity, and integrations. A paid Architecture Sprint pins them down.

// FAQ

Common questions

Under roughly 600ms end to end is the current production bar — but only if that number includes the SIP and PSTN legs. A figure measured browser-to-browser hides the part of the path your callers are actually on, so ask any vendor where their measurement starts and stops.

A hosted platform is the right default to start, and we will say so. Self-hosting earns its keep when you need the transport tuned below what the platform exposes, when data residency or on-prem is contractual, or when per-minute cost at your volume outgrows the convenience.

We build inbound and consented outbound. Unsolicited outbound AI calling is a legal product as much as an engineering one — in the US the FCC treats AI-generated voices as artificial under the TCPA, with statutory damages per call — so we scope consent and disclosure with your counsel before writing any of it.

// Related

Related services

All services

// Industries

Where teams put this to work:

15+systems shipped
6+ yrsin production
~4 wksto a first slice
99.9%uptime SLA

// Let's build

Ready to build with voice agents in production?

Tell us where you are. We reply within a day with a concrete next step.