The problem
Hosted voice-agent platforms require API keys and per-minute fees from several AI vendors. Teams that want control over data and cost need a version they can run themselves.
What I built
An API for calls, campaigns, contacts, queues, call flows and numbers; a voice agent worker; an outbound dialer; a browser softphone; metrics and alerting.
My role
Founder and sole builder.
Architecture
- LiveKit for real-time media and SIP, connected to a carrier trunk for phone calls.
- Agent worker on LiveKit Agents with local voice activity detection and barge-in handling.
- Speech to text, language model and text to speech each pluggable, with a default profile that runs open models locally.
- SQLite or PostgreSQL, with a Valkey-backed job queue for the dialer.
- Prometheus metrics and alert rules.
Technologies
How it works
When a call starts, the agent worker joins the LiveKit room, listens with voice activity detection, transcribes speech, generates a reply and speaks it back, stopping when the caller interrupts.
Key engineering decisions
- Five deployment profiles, with a default that needs no cloud AI keys.
- The worker checks text-to-speech provider compatibility when a job starts, instead of failing in the middle of a call.
- Model reasoning is turned off for voice turns to keep latency low.
- A pinned model and licence manifest keeps code licences and model-weight licences separate.
Challenges
- The CPU-only voice profile needs about 16 GB of memory, which limits cheap hosting.
- No production load test has been run yet, so capacity is not claimed.
What I learned
Self-hosting is a product feature only if the defaults work without keys. Designing the zero-key profile first shaped every other choice.
Current status
Code ready for deployment. Marketing site live; the API has no production host yet.
Links
More case studies: Capital Intelligence OS · GridResolve AI · Rankelo · Vaani · Bizia