WS
wss://app.voho.ai/v1/transcribe/wsPOST /v1/transcribe takes a finished recording. A live call needs the opposite: audio arriving in small frames and text coming back while the person is still speaking, so whatever is listening can decide when the caller has finished and when to stop talking over them.
Transcription follows Saudi Arabic as it is spoken on the phone (Najdi, Hijazi and Gulf) and keeps the English words that arrive mid-sentence.
Protocol
Authenticate with your token in theAuthorization header. A browser, which cannot set headers on a WebSocket, may pass ?token= instead. Control messages are JSON text frames; audio is sent as binary frames.
Client to server
frame
Opens transcription.
{ "type": "start", "language": "ar-SA", "sample_rate": 16000, "encoding": "pcm" }. All fields are optional; these are the defaults. encoding is pcm (16-bit little-endian, mono) or mulaw. sample_rate is between 8000 and 48000.frame
Audio in the encoding and sample rate you declared. Send frames as they are captured; 20 to 100 ms each works well. Keep sending during silence, since pauses are how a final transcript is recognised.
frame
No more audio. The final transcript for the last thing said is sent, then
done. { "type": "stop" }Server to client
frame
Authenticated, waiting for
start.frame
Transcription is open. Echoes
language, sample_rate, encoding.frame
{ "type": "transcript", "text": "...", "final": false }. Interim transcripts arrive while the caller is speaking and may be revised. A frame with "final": true settles a stretch of speech and carries confidence, and language when it was detected.frame
Session complete. Carries
seconds of audio received and cost_cents.frame
Carries
error.code and error.message.Languages
Any Arabic code (ar-SA, ar-AE, ar-OM, ar-EG, ar) also transcribes English spoken in the same sentence. Use en-US or en-GB for an English-only line.
Billing
3 cents per started minute of audio received, the same rate asPOST /v1/transcribe. A session is billed once, when it ends.

