# File audio API

Kendr exposes speech generation and recording transcription through the same
Kendr API key and wallet used for text. These endpoints require `models:invoke`.
Use `GET /v1/models` with `models:read` to discover enabled audio models and their
credit prices. Audio models have `audio_speech` or `audio_transcription` in
`capabilities` and are separate from chat and live voice sessions. For two-way
spoken conversations over WebSockets, use the [Realtime Voice API](voice-api.md)
and its [browser-readable guide](api-voice.html).

## Speech

`POST /v1/audio/speech` accepts JSON and returns a complete audio file:

```sh
curl https://api.kendr.org/v1/audio/speech \
  -H "Authorization: Bearer $KENDR_API_KEY" \
  -H 'Content-Type: application/json' \
  -H 'Idempotency-Key: exam-listening-unique-request' \
  -d '{"model":"kendr-tts","input":"The fee is £3.50, payable on 14th March.","voice":"nova","instructions":"Speak in a clear British accent, at a measured exam-listening pace.","response_format":"mp3"}' \
  --output listening.mp3
```

| Field | Supported values |
| --- | --- |
| `model` | `kendr-tts` or `gpt-4o-mini-tts` |
| `input` | Nonblank text, up to 4,096 Unicode characters |
| `voice` | `alloy`, `ash`, `ballad`, `coral`, `echo`, `fable`, `nova`, `onyx`, `sage`, `shimmer`, `verse`, `marin`, `cedar` |
| `instructions` | Optional delivery/accent instructions, up to 4,096 characters |
| `response_format` | `mp3` (default) or `wav` |
| `speed` | Optional number from 0.25 to 4.0 |

The response is binary `audio/mpeg` or `audio/wav`, normalized to 24 kHz mono.
MP3 uses 64 kbps; WAV uses 16-bit PCM. Output is limited to ten minutes. Split
long passages when using very slow speech. Requests for accents and pacing are
forwarded to the speech model; they are instructions rather than a guarantee
of a particular speaker identity. Applications should tell listeners that the
voice is AI generated.

## Transcription

`POST /v1/audio/transcriptions` accepts multipart form data:

```sh
curl https://api.kendr.org/v1/audio/transcriptions \
  -H "Authorization: Bearer $KENDR_API_KEY" \
  -H 'Idempotency-Key: exam-speaking-unique-request' \
  -F model=kendr-transcribe \
  -F file=@answer.webm \
  -F response_format=json \
  -F language=en
```

| Field | Supported values |
| --- | --- |
| `model` | `kendr-transcribe`, `gpt-4o-mini-transcribe`, or `whisper-1` |
| `file` | Up to 3 MiB and 120 seconds; WebM/Opus, Ogg/Opus, M4A or MP4/AAC, MP3, or PCM WAV |
| `response_format` | `json` (default) |
| `language` | Optional lowercase two-letter language code, such as `en` |

```json
{"text":"Um, I think the answer is seventeen.","usage":{"type":"duration","seconds":8.4}}
```

The adapter asks the model to preserve fillers and false starts, and returns
the transcript without editorial cleanup. Exact transcription remains model
dependent. Empty files and decoded digital silence return an empty `text`
without a provider transcription charge. Invalid media, unsupported codecs,
oversized uploads, and overlong recordings return 400.

Both endpoints also accept the `/api/v1/audio/...` path prefix. Public streaming
and audio through `/v1/chat/completions` are outside this file-audio API.

## Credits, retries, and errors

Requests reserve a bounded credit hold before contacting the provider, then
settle from the selected approved rate card and actual provider usage, with
the audio service's 5% markup. Any unused hold is released; the charge cannot
exceed the hold. Speech is billed by text-input and audio-output tokens,
mini-transcribe by input/output tokens, and Whisper by audio duration. A hold
may exceed the final cost, so available credits must cover admission. Current
unit prices are published in the model listing.

The seed prices below include the 5% markup and Kendr's $0.002 per-credit
conversion. An existing approved operator rate card takes precedence; always
use the live model listing for the deployed price.

| Model | Seed credit price |
| --- | --- |
| `kendr-tts` / `gpt-4o-mini-tts` | 315 per million text-input tokens plus 6,300 per million audio-output tokens |
| `kendr-transcribe` / `gpt-4o-mini-transcribe` | 656.25 per million input tokens plus 2,625 per million output tokens; approximately 1.575 credits/minute |
| `whisper-1` | 3.15 credits/minute |

The mini-transcribe per-minute figure is a planning estimate derived from
[OpenAI's estimated minute cost](https://developers.openai.com/api/docs/pricing).
Speech has no fixed character or minute price: pace and pronunciation affect
audio tokens. Measure a representative passage using `X-Kendr-Credits-Charged`
before budgeting a multi-hour exam library.

Successful responses include `X-Kendr-Request-Id`,
`X-Kendr-Credits-Charged`, and `Cache-Control: no-store`. Recordings, generated
audio, and transcripts are not saved in billing replay storage. An optional
`Idempotency-Key` prevents duplicate generation: concurrent, completed, or
released requests using the same key return 409 instead of another file.
Reusing a key with different input also returns 409. Because results are not
retained, a lost successful response cannot be downloaded again. A new key is
a new paid generation. Omit the header only when automatic retries are disabled.

Errors use this envelope:

```json
{"error":{"message":"Not enough available Kendr credits for this audio request.","type":"insufficient_quota","code":"credit_balance_exhausted","param":null}}
```

| Status | Meaning |
| --- | --- |
| 400 | Invalid input, fields, voice, format, or recording |
| 401 / 403 | Invalid credentials or denied scope/wallet access |
| 402 | Insufficient available credits |
| 409 | Existing idempotency key or changed route/rate card |
| 429 | Rate/concurrency limit; includes `Retry-After` |
| 5xx | Provider, codec, billing, or service failure |

Provider calls are never automatically retried. A failure before delivery
releases the customer hold unless settlement may already have committed.
Kendr may incur an upstream cost for a failed or interrupted generation; that
cost is not charged as a successful customer response. An ambiguous settlement
is retried idempotently and its hold is not blindly released.

## Privacy and deployment

Kendr processes recordings and generated files in memory, including codec
conversion through ffmpeg pipes. It does not write them to temporary files,
retain learner filenames, or store input text/transcripts in billing records.
Recordings and speech text are sent to the configured OpenAI API to fulfill
the request. Kendr does not use these inputs for training. Upstream retention
and any zero-data-retention controls are governed by the OpenAI account's
configuration; this implementation alone does not establish zero retention
at the upstream provider. See [OpenAI's data controls](https://developers.openai.com/api/docs/guides/your-data).

Deploy the updated `model-api`, `connector-api`, and `billing-api` images and
apply `202609250001_batch_audio_routes.sql` through the usual migration runner.
The connector image includes ffmpeg. The OpenAI connector must have a valid
credential, be enabled and healthy, and have an enabled audio route with an
approved rate card. Existing administrator disables remain authoritative.
These endpoints do not require the realtime `voice-api` service.

Local automated tests cover API contracts, authentication rejection, multipart
parsing, actual codec conversion, mocked provider responses, and credit
lifecycle failures. Before production acceptance, run a paid smoke test with
the deployed API key: synthesize MP3, transcribe it, check credit debits, listen
to all requested voices and accents, and test real Chrome/Firefox/Safari
recordings. Mocked transport and codec tests cannot certify accent quality,
pronunciation, live provider availability, or production deployment.

Provider references: [speech endpoint](https://developers.openai.com/api/reference/resources/audio/subresources/speech/methods/create),
[transcription endpoint](https://developers.openai.com/api/reference/resources/audio/subresources/transcriptions/methods/create).
