Get an API token, synthesise a sentence, and hear it back.

Prerequisites

  • A Voho account at app.voho.ai
  • curl, or Python/Node if you prefer

Get a token

1

Open the console

Sign in at app.voho.ai and go to API tokens.
2

Create a token

Give it a name that says where it will be used: billing-service, ivr-staging. When a token needs revoking, you will want to know what breaks.
3

Copy it now

The token is shown once and starts with voho_sk_live_.
Only the hash is stored, so a lost token cannot be recovered. You would create a new one and revoke the old. Put it in your secret manager before closing the page.

Your first request

Play hello.mp3. That is layla in Najdi Arabic on the default model.

Find a different voice

The catalog is browsable by dialect, which is usually the first thing you want to filter on.
Any voice ID drops straight into the voice field. See Voices and languages for the full catalog.

For telephony

If the audio is going onto a phone line, ask for mulaw at 8 kHz: that is what SIP trunks carry, so you skip a transcoding step:

For live conversation

Waiting for a complete file adds its whole synthesis time to the silence a caller hears. Stream instead, first audio arrives in roughly 200 ms rather than 1.7 s:
If your text comes from an LLM and arrives incrementally, use the WebSocket. It accepts text while audio is already playing.

Check before you ship

Run your templates through /v1/text/normalize once to see how numbers and currency will actually be spoken. Expansion is where most pronunciation surprises live.

Next

API reference

Every endpoint, parameter, and error code.

Voices

All 14 voices with descriptions.

Models

Sada versus Nabra, and what each costs.

Streaming

Audio that starts before synthesis finishes.