An on-premise deployment is sized against your traffic, not against a fixed shopping list. This page covers what your environment must be able to provide, and the numbers to bring to scoping so the sizing is done once rather than three times. A deployment starts at a single GPU server. Scale is a matter of adding nodes, not of moving to a different class of machine.

Hardware

One node needs: A single GPU is genuinely the floor, not a stripped-down trial configuration. Organisations run production lines on one node and add more as concurrency grows.
Size for headroom above peak, not at it. A voice deployment that is exactly sized fails during the busiest hour of the year, which is the hour it is judged on.

Software

  • OS: Linux. Ubuntu or Debian are the tested distributions.
  • Container runtime: Docker with NVIDIA GPU support, or Kubernetes for a multi-node cluster.
  • GPU drivers: NVIDIA drivers and the container toolkit, so containers can address the GPU.

Deployment forms

Three, so the deployment fits the platform you already run rather than the other way round.

Docker

A single GPU server. The usual starting point, and enough for a production line.

Kubernetes

A cluster, for scaling across nodes and for estates that standardise on Kubernetes.

Single container

One self-contained image, for platforms that require a single-container deployment.

Storage

The 50–100 GB above covers the deployment itself. Records are separate and are sized by your retention policy:
  • Transcripts, summaries and action logs for the retention period you set.
  • Audio, if you choose to retain it: this dominates the volume, and it does not need the same retention as text.
  • Backup and restore, on your existing regime.
Decide retention per data type before go-live rather than after. See Security and data handling.

Network

Low-latency paths between call handling, the speech components and the agent runtime. Every round trip between them is added to the silence the caller hears.
SIP connectivity from your trunk or PBX to the call-handling component. Audio is carried as 8 kHz mulaw, so no transcoding tier is needed.
Reachability from the action layer to each system the agent will act in, plus a service account on each with only the permissions those actions require.
No live connection is needed to keep the deployment running, licensing is checked locally against a signed file. Usage is reported back to Voho periodically for billing, which is the one outbound flow to account for in your firewall policy. See Licensing and egress.

Identity

  • A SAML or OIDC identity provider for operator sign-in.
  • SCIM, if you want provisioning and de-provisioning to follow your directory automatically.
  • Your role definitions: who may read transcripts, who may change agent instructions, who may export.

What to measure before scoping

Bring these and the sizing conversation takes one call instead of four.
1

Peak concurrent calls

Not calls per day, the highest number of calls live at the same moment. Pull it from your PBX or contact-centre reporting for the busiest hour of the last twelve months.
2

Call profile

Average call length, and the split between languages and dialects. A three-minute Arabic service call and a twenty-second balance enquiry size very differently.
3

Traffic shape

When your peaks land: time of day, day of week, and the seasonal spikes your business knows about. Ramadan, salary week and results day are real capacity events.
4

Systems in scope

Which systems the agent must act in for the first workflow, and whether an API and a service account already exist for each.
5

Retention and residency obligations

What you must keep, for how long, and where it must stay. This one can change the architecture, so it belongs at the start.
6

Who operates it

The team that will run this after handover, and what they already run. Monitoring that does not fit their existing tooling does not get watched.

Send us these numbers

Sizing starts from your peak concurrency and call profile.