The registry, image references and configuration values are specific to the release you will run and are provided with the runbook for your deployment.
Choose the deployment form
- Docker
- Kubernetes
- Single container
One GPU server, Docker with NVIDIA GPU support. The usual starting point: a single node is a production configuration, not a trial one. Choose this unless you have a reason not to.
The path
Scoping
One call covering capacity, environment, obligations and operations. Bring the numbers from Requirements, peak concurrent calls above all.Out of it: a sizing for your traffic, a deployment shape, and the list of what each side provides.
Environment pre-flight
Your team provisions hosts, GPUs, storage and network paths, and creates the service accounts the agent will act through. Voho reviews the environment against the sizing before anything is installed.Out of it: an environment that passes review, so install day is install day.
Install
Voho and your team deploy together, in your environment, against your runbook: pull the images, place the signed licence file, start the deployment, and confirm the GPU is addressed. Nothing is installed that your team has not seen installed.Out of it: a running deployment, health-checked, with the local dashboard showing usage and system health.
Connect telephony
Point a test number at the deployment over your existing trunk. Audio is 8 kHz mulaw, so there is no transcoding tier to build.Out of it: a real call, answered, in your building.
Build the first workflow
Configure the agent’s instructions and wire the actions it needs into your systems. Start with one workflow that matters and finishes (a status enquiry, a ticket being raised) not a broad assistant.See Instructions and Actions.Out of it: an agent that completes a task end to end, including its handover path.
Test before traffic
Run it against real recordings and awkward cases: heavy dialect, code-switching mid-sentence, callers who interrupt, reference numbers read back digit by digit. See Testing.Out of it: a pass criterion agreed before the first customer hears it.
Pilot
Route a slice of live traffic. Watch time to first audio, handover rate and action failures. Fix what the transcripts show rather than what the demo suggested.Out of it: evidence from your own callers.
Cutover and handover
Widen traffic, and move first-line operations to your team with the runbooks, dashboards and escalation path in place.Out of it: a deployment your team runs. See Operations.
Prepare before day one
Every item here is something only your organisation can produce, and each one blocks the step it belongs to.- Peak concurrent calls, from the busiest hour of the last year.
- A GPU host provisioned to the sizing (16–24 GB VRAM or better, 8 cores, 32 GB RAM, 50–100 GB disk) on Ubuntu or Debian, with NVIDIA drivers and Docker GPU support in place.
- Network paths opened: trunk to call handling, action layer to each target system.
- Service accounts on each system the agent will act in, scoped to only what its actions need.
- Identity provider ready for SSO, with your role definitions decided.
- Retention policy per data type (audio, transcripts, summaries) agreed with your legal and data protection people.
- A test number you can route freely without touching production traffic.
- A firewall decision on the periodic usage sync, or an agreed offline path if you are deploying with zero egress.
- The first workflow chosen, with the person who owns that process available.
Who needs to be in the room
Next
Requirements
What to size and provision.
Data residency
What your security review will ask.
Start a scoping conversation
Bring your peak concurrency, your obligations, and the first workflow you want live.

