Realtime Voice API
Build spoken conversations with authenticated WebSockets, PCM audio, transcripts, and reliable session renewal.
Build a two-way spoken conversation with Kendr Voice over a WebSocket. The
protocol is kendr.voice.v1, backed by amazon.nova-2-sonic-v1:0. Input audio,
spoken output, transcripts, status, and control messages share one connection.
For generated MP3/WAV files or uploaded recording transcription, use the
file audio API. Realtime Voice is not the OpenAI Realtime protocol
and cannot be called through Chat Completions or an OpenAI audio SDK.
See the published integration guide,
OpenAPI description, and runnable Python client.
OpenAPI describes HTTP admission; its x-websocket extension describes messages
after the 101 upgrade. Ordinary generated HTTP clients do not implement that stream.
Credentials and endpoints
| Operation | URL | Required API-key scope |
|---|---|---|
| Discover voices and audio settings | GET https://api.kendr.org/v1/voice/voices |
models:read |
| Open a conversation | GET wss://api.kendr.org/v1/voice/stream (WebSocket upgrade) |
models:invoke |
Send Authorization: Bearer <KENDR_API_KEY> or X-API-Key: <KENDR_API_KEY> in
the HTTP request/upgrade headers. These public routes require a Kendr API key;
browser cookies, app-session tokens, and OAuth access tokens do not authenticate
them. Never put credentials in query strings or a WebSocket subprotocol.
Server and native WebSocket clients may omit Origin. If supplied, it must
match an origin trusted by the deployment. The browser's native WebSocket
constructor cannot set these authentication headers. Authenticate browser users
to your own backend, enforce their permissions and spending limits there, and
relay the WebSocket through that backend. Keep the Kendr API key on the backend;
do not embed a long-lived key in browser JavaScript.
curl https://api.kendr.org/v1/voice/voices \
-H "Authorization: Bearer $KENDR_API_KEY"
Example response (data is abbreviated):
{
"object": "list",
"data": [{"id": "tiffany", "object": "voice"}],
"protocol": "kendr.voice.v1",
"model": "amazon.nova-2-sonic-v1:0",
"input_audio": {"encoding": "pcm_s16le", "sample_rate_hz": 16000, "channels": 1, "frame_bytes": 1024},
"output_audio": {"encoding": "pcm_s16le", "sample_rate_hz": 24000, "channels": 1},
"max_duration_seconds": 450
}
Use an ID from this catalog. File-audio voices such as nova are a separate
catalog. This describes configuration, not available capacity or provider
health. Duration is deployment dependent, up to 450 seconds. Honor each
session.ready.max_duration_seconds.
Start a session
Send this JSON as the first text message, promptly after the upgrade:
{
"type": "start", "mode": "full", "voice_id": "tiffany",
"context": [{"role": "user", "content": "We are practicing a support conversation."}],
"tools_enabled": false,
"client_context": {"timezone": "Asia/Kolkata", "utc_offset_minutes": 330, "locale": "en-IN", "temperature_unit": "celsius"}
}
| Field | Contract |
|---|---|
type |
Required, exactly start |
mode |
Required, exactly full; public sessions are billed; preview is unavailable |
voice_id |
Required catalog ID on every connection; does not read or change the saved browser voice |
context |
Optional ordered history, default []; only user and assistant roles; maximum 40 messages, each nonblank and at most 4,000 Unicode characters; total default limit 24,000 characters (deployment configurable) |
tools_enabled |
Optional boolean, default false; enables Kendr built-in tools executed by the service, not client-defined function calling |
client_context |
Optional object; timezone up to 100 characters, locale up to 40, temperature_unit empty/celsius/fahrenheit, utc_offset_minutes integer from -840 to 840; strings cannot contain control characters |
continuation_token |
Optional opaque token from the previous session.ready; bound to the same account and voice; still requires API-key authentication and matching voice_id |
Unknown request fields are rejected. Do not send a custom model, system prompt,
preview_text, or custom tool definitions. The first-message timeout defaults
to 15 seconds (configurable from 3 to 30). Provider setup has its own bounded
timeout. Client messages must each fit within 32 KiB; this can constrain context
before its character limit.
Wait for session.ready before sending PCM. While connecting you may send
{"type":"end"} to cancel. A typical ready event is:
{
"type": "session.ready", "protocol": "kendr.voice.v1", "session_id": "session-id",
"mode": "full", "model": "amazon.nova-2-sonic-v1:0", "voice_id": "tiffany",
"continuation_token": null,
"input_audio": {"encoding": "pcm_s16le", "sample_rate_hz": 16000, "channels": 1},
"output_audio": {"encoding": "pcm_s16le", "sample_rate_hz": 24000, "channels": 1},
"max_duration_seconds": 450
}
The token is an opaque string when available, otherwise JSON null. Store it
privately if you plan to renew. Do not log it.
Send and receive audio concurrently
Send raw binary WebSocket messages of exactly 1,024 bytes, containing signed 16-bit little-endian PCM, mono, at 16,000 Hz. Each represents 512 samples or 32 ms. Resample microphone data, remove any WAV header, buffer exact frames, and pad the final partial frame with silence. Pace at real time; do not upload an entire recording in a burst.
Binary server messages contain 24,000 Hz mono signed 16-bit little-endian PCM. Sizes vary (currently up to 16 KiB); concatenate them in arrival order or queue them into a 24 kHz player. These are not MP3, WAV containers, JSON, or base64. Keep receiving while sending, including during assistant speech. Use your WebSocket library's ping/pong support.
| Direction / message | Behavior |
|---|---|
Client {"type":"barge_in"} |
Requests local playback clearing; keep forwarding microphone audio for provider speech detection |
Client {"type":"end"} |
Ends the entire session, not just an utterance; await session.ended and close |
Server status |
state: connecting, listening, thinking, speaking; UI hints, not reliable utterance-end markers |
Server transcript |
role, text, final, delta; append delta:true to that role's current partial; final:true,delta:false replaces it with the complete utterance |
Server playback.clear |
Stop current audio and discard queued playback immediately; preserve committed history |
Server tool |
Built-in tool progress: id, name, state (started, completed, failed); no client tool-result response is supported |
Server error |
code, message, retryable, fatal; see recovery below |
Server session.ended |
Terminal reason, canonical renewable boolean, legacy reconnect alias, cumulative usage, billing_units |
The transcript has no stable turn ID. Keep separate partial buffers by role, commit each final utterance once, and reset partial buffers at connection boundaries. Preserve finalized history across retries. Ignore unknown response fields/event types for forward compatibility.
{
"type": "session.ended", "reason": "client_end", "renewable": false, "reconnect": false,
"usage": {"input_speech_tokens": 100, "input_text_tokens": 20, "output_speech_tokens": 80, "output_text_tokens": 15, "total_input_tokens": 120, "total_output_tokens": 95},
"billing_units": {"audio_input_million_tokens": "0.000100", "text_input_million_tokens": "0.000020", "audio_output_million_tokens": "0.000080", "text_output_million_tokens": "0.000015"}
}
Usage totals are cumulative per connection; do not sum repeated snapshots. Zero-valued billing-unit classes may be omitted.
Runnable Python example
Download voice_stream.py, use Python 3.11 or newer:
python -m pip install 'websockets>=14,<16'
export KENDR_API_KEY='your-kendr-api-key'
ffmpeg -i question.webm -ar 16000 -ac 1 -c:a pcm_s16le question.wav
python voice_stream.py question.wav --voice tiffany --output reply.wav --listen-seconds 10
On PowerShell use $env:KENDR_API_KEY = 'your-kendr-api-key' instead of export.
The example validates the WAV, waits for readiness, sends 32 ms frames while
receiving concurrently, prints final transcripts, sends silence during a
bounded reply window, gracefully ends, and writes a 24 kHz WAV. Increase
--listen-seconds (maximum 60) for longer answers. This captures a single
session; it does not implement microphone capture or live playback, and never
automatically retries or renews a paid session.
Errors and renewal
Before upgrade, HTTP errors use the platform envelope:
{"ok":false,"error":{"code":"invalid_authentication","message":"A Kendr API key is required."}}
Optional request_id and details may be present. Typical handshake/catalog
statuses are 400 (invalid request), 401 (missing/invalid key), 403 (scope/origin
denied), 503 (service unavailable). After upgrade, HTTP success means only
that the socket opened. Validation, credit admission, capacity, and provider
failures arrive as WebSocket error events:
{"type":"error","code":"voice_invalid_audio_frame","message":"Audio frames must contain exactly 1024 bytes of 16-bit PCM audio.","retryable":false,"fatal":true}
Treat fatal:true as terminal; the server may close with code 1008 without
session.ended. Correct input, credentials, scopes, or credit balance before
retrying. Useful categories include insufficient_credits, rate_limit_exceeded,
voice_invalid_start, voice_audio_before_ready, voice_start_timeout, and
provider failure codes; handle unknown codes too. A nonfatal error may precede
a renewable session.ended: keep reading. A code-1000 close alone does not prove
an error-free session; inspect its events.
For longer calls or transient failures:
- Stop the old audio sender and discard its playback queue. Explicit user hang-up must cancel reconnects already in progress.
- If
session.ended.renewableis true (such asmax_durationorserver_shutdown), wait for the old socket to close. For retryable errors or unexpected disconnects, use bounded exponential backoff with jitter and a finite attempt/time budget. Avoid tight reconnect loops. - Open a new authenticated socket and send
startwith the samevoice_id, bounded recent finalizedcontext, and the last non-nullcontinuation_token. Tokens expire about ten minutes after issue; an expired token is an error, not authorization to silently discard conversation history. - Wait for readiness, save the newest token, and begin fresh microphone audio. Do not replay old PCM or send two streams at once. Without a token, a fresh session can use explicit voice and bounded context under normal billing.
Continuation preserves voice identity, not a provider socket, transcript store, audio replay buffer, or billing reservation. Your client supplies history. Each accepted connection is a separate billed session. There is no voice idempotency key or response-replay endpoint. Renewals must respect an overall application time/spend budget. Capacity is shared per account (default one live session, configurable); browser calls and previews share that capacity.
Credits and operational checks
Full sessions reserve credits before provider connection and settle actual reported speech-input, speech-output, text-input, and text-output tokens at termination. The default hold is 500 credits (configurable), not the price of a call. Settlement is bounded by the hold, uses the approved voice rate card when available and configured interim rates otherwise, and releases unused credit. Unused sessions release their hold. Interrupted sessions can still have billable usage.
Settlement happens after the terminal event/socket close. usage and
billing_units are usage reports, not a final credit receipt. Check
wallet/usage records after settlement. An immediate renewal can temporarily
require credit for the old pending hold plus the new hold. Do not apply the
file-audio API's response headers, 5% markup, or replay rules to live voice.
Deploy voice-api and gateway WebSocket/catalog routes together. Mock tests
do not establish provider availability, audio quality, or production billing
correctness. Before production acceptance, send real spoken input through an
API key, listen to the reply, verify transcripts and wallet settlement,
interrupt playback, cancel during connection, renew at the duration limit,
and reject invalid credentials/scopes and exhausted capacity.
Audio and context are sent to the configured Amazon Bedrock voice provider. Tell users they are interacting with AI and apply your own consent, transcript retention, and relay access policies. Do not log API keys, continuation tokens, or raw audio.