Overview
The Nia API turns text into natural speech, transcribes audio back into text, translates and dubs between languages, hosts live two-way voice conversations, and lets you talk to Nia's brain — all over HTTP. It's the same engine that powers the Playground and the studios — exposed behind an API key so your own apps can call it.
Speech spans many languagesacross two local engines — Kokoro (English & European voices, multiple speakers each) and Meta MMS (Indian languages like Tamil, Malayalam, Telugu, Bengali & more, plus Arabic). You never pick the engine: choose a language code and Nia routes it. The full, always-current list is below and at GET /v1/voices.
Base URL
https://platform.naslabs.aiEvery endpoint runs on Nia's own engine — the models are ours and self-hosted, not a third-party API behind a wrapper.
Create an API key
- Get access. Sign-up is invite-only — ask your Nia contact for an invite code, then create your account at /signup.
- Open the API Keys tab in the dashboard.
- Click + New key, give it a name, and choose scopes (
speech,voices,chat,dub,transcribe,translate,realtime). - Copy the key immediately— it's shown once and never again. Store it as an environment variable.
Authentication
Send your key in the Authorization header as a Bearer token on every request:
curl https://platform.naslabs.ai/v1/voices \
-H "Authorization: Bearer $NIA_KEY"A missing, unknown, or revoked key returns 401. A valid key used on an endpoint outside its scopes returns 403.
Authorization header, so the credential has to travel in the signaling URL — and a long-lived key must never go there (URLs leak into proxy logs and browser history). Instead you exchange your key at POST /v1/realtime/session for a token that expires in about a minute and opens exactly one call. Mint it on your server, hand it to your client.For a live call to actually carry audio, two things must hold: your page is served over HTTPS (or
http://localhost) so the browser grants the microphone, and the WebRTC client is configured with the ice_servers from the session — a caller behind a symmetric NAT reaches the agent only through the TURN relay in that list, or the call connects silently.Quickstart
From zero to a spoken WAV in two commands:
# 1. Export the key you copied from the API Keys tab
export NIA_KEY="nia_sk_..."
# 2. Turn text into speech — saves a WAV you can play
curl -s https://platform.naslabs.ai/v1/speech \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"Hello from the Nia API.","voice":"af_heart"}' \
--output hello.wavEndpoints
/v1/speechscope: speechSynthesize speech from text. Returns audio bytes (WAV by default). Send the text in the target language's own script — the Indic and Arabic voices are script-locked single-speaker models, so Latin text is not transliterated and comes back as a near-empty clip or an error.
| Field | Type | Required | Description |
|---|---|---|---|
| text | string | yes | The text to speak (up to 3000 chars). Must be in the script of the chosen language — Malayalam text for a Malayalam voice, Arabic for an Arabic voice. |
| language | string | — | Language code from GET /v1/voices, e.g. "en-us", "hi", "ta", "ar". Picks the voice engine. Default "en-us". |
| voice | string | — | Voice id, e.g. "af_heart". Multi-voice on English/European (Kokoro) languages; Indian/Arabic (MMS) languages have one voice per language, so language alone is enough. |
| speed | number | — | 0.5–2.0. Playback speed. Default 1.0. |
| format | string | — | "wav" (default) or "mp3". |
curl -s https://platform.naslabs.ai/v1/speech \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"Good morning.","voice":"af_heart","speed":1.0}' \
--output speech.wav200 OK
Content-Type: audio/wav
X-Nia-Duration-Ms: 1180
X-Nia-Sample-Rate: 24000 # 24000 for Kokoro voices, 16000 for MMS (Indian/Arabic)
<binary WAV data>/v1/voicesscope: voicesList every available language and its voices — the same catalog the Playground uses. Poll this to discover language codes and voice ids; new languages appear here automatically.
curl -s https://platform.naslabs.ai/v1/voices \
-H "Authorization: Bearer $NIA_KEY"{
"languages": [
{
"code": "en-us",
"label": "English (US)",
"engine": "kokoro",
"voices": [
{ "id": "af_heart", "label": "Heart (F)" },
{ "id": "am_puck", "label": "Puck (M)" }
]
},
{
"code": "ta",
"label": "Tamil",
"engine": "mms",
"voices": [
{ "id": "mms_ta", "label": "Tamil (Neural)" }
]
}
]
}/v1/chatscope: chatSend a message to Nia's brain and get a text reply. The response carries real token spend in "usage": cached_tokens is the part of input_tokens served from the model's prefix cache — a subset, never an addition, so your total is input_tokens + output_tokens. It is null (not 0) when the model doesn't report caching. The same accounting applies to /v1/translate and /v1/dub and is summarised in the Analytics tab.
| Field | Type | Required | Description |
|---|---|---|---|
| message | string | — | A single user message. (Or use "messages".) |
| messages | array | — | OpenAI-style [{role, content}] history. Overrides "message". |
| temperature | number | — | Sampling temperature. Optional. |
| model | string | — | Model id. Default "nia-chat-1". |
| stream | boolean | — | Stream the reply as Server-Sent Events instead of one blocking JSON response. Default false. Use it when a cold model would otherwise leave your user staring at nothing (or push you past a serverless timeout) — the first words render as they're generated. |
curl -s https://platform.naslabs.ai/v1/chat \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"message":"Say hello in four words."}'{
"reply": "Hello, nice to meet you.",
"model": "nia-chat-1",
"usage": {
"input_tokens": 12,
"output_tokens": 15,
"cached_tokens": null
}
}/v1/dubscope: dubDub an audio clip into another language: transcribe → translate → re-speak, in one call. Send audio as multipart/form-data. Optionally keep the original speaker's voice (clone) or match the source timing. Returns JSON with base64 audio + the transcript and translation.
| Field | Type | Required | Description |
|---|---|---|---|
| file | file | yes | The audio to dub (multipart form field). mp3/m4a/wav/webm/mp4, up to 25 MB. |
| target_language | string | yes | Language code to dub into, from GET /v1/voices (e.g. "es", "hi", "ar"). |
| source_language | string | — | Source language code, or "auto" (default) to detect it. |
| voice | string | — | Target voice id for the dub. Ignored when preserve_voice=true. |
| speed | number | — | 0.5–2.0. Dub playback speed. Default 1.0. |
| preserve_voice | boolean | — | Clone the original speaker and re-speak in their voice (ignores voice). Default false. |
| timing | boolean | — | Align each dubbed segment to the source's timestamps. Default false. |
| format | string | — | "wav" (default) or "mp3" — the encoding of the returned base64 audio. |
# Dub an English clip into Spanish, keeping the original voice.
# The base64 audio is returned in JSON — decode "audio" to a file to play it.
curl -s https://platform.naslabs.ai/v1/dub \
-H "Authorization: Bearer $NIA_KEY" \
-F "file=@clip.mp3" \
-F "target_language=es" \
-F "preserve_voice=true"{
"audio": "<base64 WAV>",
"format": "wav",
"sampleRate": 24000,
"durationMs": 5470,
"sourceText": "Hello, I'm doing quite well, thanks for asking.",
"translation": "Hola, estoy bastante bien, gracias por preguntar.",
"sourceLanguage": "English",
"targetLanguage": "Spanish",
"targetLanguageCode": "es",
"detectedLanguage": "en",
"truncated": false,
"timing": false,
"preserveVoice": true
}/v1/transcribescope: transcribeTranscribe an audio clip to text (speech-to-text). Send audio as multipart/form-data. Returns the transcript, the detected language, and the clip duration. Leaving the language out runs a dedicated identification stage and confirms it with a script-aware second model, so Indic languages that sound alike — Malayalam and Tamil especially — come back in the right language and the right script. That cascade costs time: measured on ~9s clips, auto-detect returns English in ~2.2s, Hindi ~3.5s, Tamil ~3.9s and Malayalam in 10–16s, against 2.1–2.9s for the same clips with language passed. On a very short utterance identification can still land on a phonetic neighbour. Pass language explicitly whenever you know it — none of that runs, and it is the most accurate path.
| Field | Type | Required | Description |
|---|---|---|---|
| file | file | yes | The audio to transcribe (multipart form field). mp3/m4a/wav/webm/mp4, up to 25 MB. |
| language | string | — | Language code hint from GET /v1/voices (e.g. "en-us", "hi", "ta"), or "auto" (default) to detect it. Naming the language skips detection entirely — faster, and more accurate than auto-detect for languages that sound alike. |
curl -s https://platform.naslabs.ai/v1/transcribe \
-H "Authorization: Bearer $NIA_KEY" \
-F "file=@clip.mp3" \
-F "language=auto"{
"text": "Hello, I'm doing quite well, thanks for asking.",
"language": "en",
"durationMs": 3200,
"model": "nia-scribe-1"
}/v1/realtime/sessionscope: realtimeStart a live, two-way voice conversation — the same realtime agent as the Playground (you talk, Nia listens and talks back, interrupting and turn-taking naturally). Unlike the other endpoints, this one carries no audio: audio streams peer-to-peer over WebRTC. It returns a short-lived session token plus the signaling URL and ICE servers your client needs to open that connection.
| Field | Type | Required | Description |
|---|---|---|---|
| voice | string | — | Voice id for Nia's replies, e.g. "af_heart". Default from the server config. |
| language | string | — | Language code from GET /v1/voices, e.g. "en-us", "hi". Sets both what Nia hears and what she speaks. |
| tone_preset | string | — | Persona tone, e.g. "warm", "concise". Optional. |
| style_text | string | — | Free-text instructions for the persona — who you are, hours, escalation rules, what not to promise. Up to 4,000 characters (it was 400; the cap is policy, not a model limit). Optional. |
| temperature | number | — | Sampling temperature for Nia's brain. Optional. |
| persona_id | string | — | A saved persona ("p_...") — its voice, tone, and language. Create one with POST /v1/personas. Optional. |
| mode | string | — | "assistant" (default) runs Nia's own brain. "external" runs the call with NO built-in LLM: you receive every user turn and push the reply text back for Nia to speak — that's how you put your own RAG behind the voice. See "Live voice with your own brain" below. |
| ttl_seconds | number | — | How long the session token stays valid, 1–600. Default 60. This is a deadline to *connect*, not a call length limit — once connected, the call runs as long as you like. |
# Exchange your API key for a short-lived session token.
# No audio here — this is the handshake that authorizes the live connection.
curl -s https://platform.naslabs.ai/v1/realtime/session \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"voice":"af_heart","language":"en-us"}'{
"session_id": "rts_9f2c1a4b7e8d3c5a6b0f2e1d",
"token": "nia_rt_eyJzaWQiOiJydHNfOWYyYzFhNGI...",
"expires_at": "2026-07-17T09:15:00.000Z",
"expires_in": 60,
"url": "https://api.naslabs.ai/api/offer",
"ice_servers": [
{ "urls": "stun:stun.l.google.com:19302" },
{ "urls": ["turn:turn.example.com:3478"], "username": "nia", "credential": "..." }
],
"settings": { "voice": "af_heart", "language": "en-us" },
"model": "nia-realtime-1"
}/v1/realtime/ttsscope: realtimeOpen a realtime text-to-speech stream: send text over a WebSocket, get Kokoro audio back chunk-by-chunk as it synthesizes (PCM16 mono @ 24 kHz). Just the voice — no LLM, no mic. This POST mints a short-lived token and returns the wss:// URL to open; the audio streams over that socket. Unlike the live agent, a WebSocket is plain TCP, so this works anywhere HTTP does (including a RunPod pod) with no TURN relay.
| Field | Type | Required | Description |
|---|---|---|---|
| voice | string | — | Voice id for the stream, e.g. "af_heart". Kokoro voices only for now. Optional. |
| language | string | — | Kokoro language code from GET /v1/voices, e.g. "en-us", "hi". Optional. |
| speed | number | — | Speaking rate 0.5–2.0. Optional. |
| ttl_seconds | number | — | How long the token stays valid, 1–600. Default 60. A deadline to *connect*, not a stream-length limit. |
# 1. Mint a token + the wss:// URL (no audio here — this is the handshake).
curl -s https://platform.naslabs.ai/v1/realtime/tts \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"voice":"af_heart","language":"en-us"}'
# → { "url": "wss://<host>/api/realtime/speak?token=nia_rt_...", "sample_rate": 24000, ... }
# 2. Open that wss:// URL, send {"type":"speak","text":"..."} and read PCM16 frames.
# (curl can't stream WebSocket audio — see the Python tab.){
"session_id": "rts_9f2c1a4b7e8d3c5a6b0f2e1d",
"token": "nia_rt_eyJzaWQiOiJydHNf...",
"expires_in": 60,
"url": "wss://api.naslabs.ai/api/realtime/speak?token=nia_rt_...",
"stream_url": "wss://api.naslabs.ai/api/realtime/speak",
"sample_rate": 24000,
"format": "pcm_s16le",
"channels": 1,
"model": "nia-realtime-tts-1"
}/v1/realtime/sttscope: realtimeOpen a realtime speech-to-text stream: push PCM16 mono @ 16 kHz audio up a WebSocket and get live transcripts back. The server segments your audio with voice-activity detection and emits a "final" message per finished utterance, plus "interim" messages while the speaker is still talking so you can show text as it forms. Every utterance is language-identified on its own, so a speaker who switches language between utterances is transcribed correctly in each; Indic languages are settled by a dedicated acoustic model rather than generic Whisper, which is what stops Malayalam, Tamil and their neighbours collapsing into whichever the general model guesses. A large improvement, not a guarantee — a short opening utterance can still be labelled a neighbouring language, and a speaker who switches without pausing stays in the language the utterance started in. Pass a language to pin the decoder and skip detection entirely — faster, and more accurate whenever you already know what will be spoken. Just the ears — no LLM, no voice. This POST mints a short-lived token and returns the wss:// URL to open. Plain TCP, so it works anywhere HTTP does with no TURN relay.
| Field | Type | Required | Description |
|---|---|---|---|
| language | string | — | Pin the decoder to one language, from GET /v1/voices (e.g. "en", "hi"). Omit for automatic per-utterance detection. Pinning skips language identification, so it is both faster and more accurate than letting auto-detect choose between languages that sound alike. Optional. |
| ttl_seconds | number | — | How long the token stays valid, 1–600. Default 60. A deadline to *connect*, not a stream-length limit. |
# 1. Mint a token + the wss:// URL.
curl -s https://platform.naslabs.ai/v1/realtime/stt \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"language":"en"}'
# → { "url": "wss://<host>/api/realtime/transcribe?token=nia_rt_...", "sample_rate": 16000, ... }
# 2. Open that wss:// URL, send raw PCM16 @ 16 kHz binary frames, read {"type":"final"} JSON.
# Send {"type":"flush"} to force-finalize. (curl can't stream — see the Python tab.)# The POST returns the token + URL:
{
"session_id": "rts_1b2c3d4e5f6a7b8c9d0e1f2a",
"token": "nia_rt_eyJzaWQiOiJydHNf...",
"expires_in": 60,
"url": "wss://api.naslabs.ai/api/realtime/transcribe?token=nia_rt_...",
"stream_url": "wss://api.naslabs.ai/api/realtime/transcribe",
"sample_rate": 16000,
"format": "pcm_s16le",
"channels": 1,
"model": "nia-realtime-stt-1"
}
# Then, over the WebSocket — provisional text while the speaker talks, and one
# settled message per finished utterance. Switch on "type"; treat any type you
# do not recognise as ignorable, so new ones can be added without breaking you.
# { "type": "interim", "text": "hi there how", "language": "en" }
# { "type": "final", "text": "hi there, how are you", "language": "en" }/v1/realtime/stt/autoscope: realtimeLive transcription with the language identified per utterance. A sibling of /v1/realtime/stt, not a replacement — that route still points at the socket you built against, so nothing changes for existing integrations. Use this one when you do not know what will be spoken: each utterance is identified on its own (acoustic language ID, then a dedicated Indic recognizer to settle Malayalam / Tamil / Telugu / Kannada, then a cross-check), and a speaker who changes script family mid-sentence is split into two utterances. Same wire protocol as /v1/realtime/stt: PCM16 mono @ 16 kHz up, "ready" / "interim" / "final" down, 4401 on a bad token. Identification needs roughly 1.5–3 seconds of speech to be reliable — a very short opening utterance can still come back as a neighbouring language, so pin with /v1/realtime/stt/lang whenever you know the language.
| Field | Type | Required | Description |
|---|---|---|---|
| ttl_seconds | number | — | How long the token stays valid, 1–600. Default 60. A deadline to *connect*, not a stream-length limit. |
curl -s https://platform.naslabs.ai/v1/realtime/stt/auto \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" -d '{}'
# → { "url": "wss://<host>/api/realtime/transcribe/auto?token=nia_rt_...", "auto_detect": true, ... }
# Then open that URL and stream PCM16 @ 16 kHz exactly as for /v1/realtime/stt.{
"session_id": "rts_1b2c3d4e5f6a7b8c9d0e1f2a",
"token": "nia_rt_eyJzaWQiOiJydHNf...",
"expires_in": 60,
"url": "wss://api.naslabs.ai/api/realtime/transcribe/auto?token=nia_rt_...",
"stream_url": "wss://api.naslabs.ai/api/realtime/transcribe/auto",
"sample_rate": 16000,
"format": "pcm_s16le",
"channels": 1,
"interim": true,
"auto_detect": true,
"model": "nia-live-stt-auto-1"
}/v1/realtime/stt/langscope: realtimeLive transcription pinned to one language. The counterpart of /v1/realtime/stt/auto, and the one to prefer whenever you already know what will be spoken: pinning skips language identification entirely, which is both faster per utterance and more accurate than letting a detector choose between languages that sound alike. "language" is required here — an omitted one returns 400 rather than silently falling back to detection. The code is appended to the returned url as &language=, because a browser WebSocket can send neither headers nor a body on the upgrade; server-side clients may instead send {"type":"config","language":"..."} once connected. Same wire protocol as /v1/realtime/stt.
| Field | Type | Required | Description |
|---|---|---|---|
| language | string | yes | The language to pin, from GET /v1/voices (e.g. "ml", "ta", "en"). Required — use /v1/realtime/stt/auto for detection. |
| ttl_seconds | number | — | How long the token stays valid, 1–600. Default 60. A deadline to *connect*, not a stream-length limit. |
curl -s https://platform.naslabs.ai/v1/realtime/stt/lang \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"language":"ml"}'
# → { "url": "wss://<host>/api/realtime/transcribe/lang?token=nia_rt_...&language=ml", ... }
# Then open that URL and stream PCM16 @ 16 kHz exactly as for /v1/realtime/stt.{
"session_id": "rts_1b2c3d4e5f6a7b8c9d0e1f2a",
"token": "nia_rt_eyJzaWQiOiJydHNf...",
"expires_in": 60,
"url": "wss://api.naslabs.ai/api/realtime/transcribe/lang?token=nia_rt_...&language=ml",
"stream_url": "wss://api.naslabs.ai/api/realtime/transcribe/lang",
"sample_rate": 16000,
"format": "pcm_s16le",
"channels": 1,
"interim": true,
"auto_detect": false,
"settings": { "language": "ml" },
"model": "nia-live-stt-lang-1"
}/v1/translatescope: translateTranslate text from one language into another — the text engine behind the live Interpreter. Send JSON, get the translation back.
| Field | Type | Required | Description |
|---|---|---|---|
| text | string | yes | The text to translate (up to 3000 chars). |
| target_language | string | yes | Language code to translate into, from GET /v1/voices (e.g. "es", "hi", "ar"). |
| source_language | string | — | Source language code, or "auto" (default) to detect it. |
curl -s https://platform.naslabs.ai/v1/translate \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"Good morning, how are you?","target_language":"es"}'{
"source": "Good morning, how are you?",
"translation": "Buenos días, ¿cómo estás?",
"sourceLanguage": "English",
"targetLanguage": "Spanish",
"targetLanguageCode": "es",
"model": "nia-translate-1"
}/v1/embeddingsscope: chatTurn text into vectors, so retrieval can live on your side. We don't store your documents, chunk them, or search them — you keep your own store (pgvector, or anything else) and call this for the one part that needs a model on a GPU. It's what lets "can I send it back" match a policy that says "returns", which keyword search never will. Uses the "chat" scope: it runs a text model on the same host as the brain.
| Field | Type | Required | Description |
|---|---|---|---|
| input | string | string[] | yes | The text to embed. Send an array to embed a batch in one call — at most 64 items and 32,000 characters total, because each one runs a forward pass on the same GPU your live calls use. |
curl -s https://platform.naslabs.ai/v1/embeddings \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"input":["Standard delivery is next-day within Dubai.","Returns are accepted within 14 days."]}'{
"embeddings": [[0.0123, -0.0456, "..."], [0.0219, -0.0031, "..."]],
"dimensions": 768,
"model": "nia-embed-1",
"usage": { "input_tokens": 24 }
}/v1/personasscope: voicesCreate and manage saved personas — a voice (cloned from a reference clip, or a base voice) bundled with a tone, style and language. The id you get back is the "persona_id" accepted by /v1/realtime/session and /v1/speech. GET lists them, GET/PATCH/DELETE /v1/personas/{id} fetch, retune and remove one. Uses the "voices" scope: a persona is a saved voice. Personas belong to the workspace that owns the key — another key can never see or touch them.
| Field | Type | Required | Description |
|---|---|---|---|
| kind | string | — | "kokoro" (default for a JSON body) bundles an existing base voice. "clone" builds a new voice from your clip and must be sent as multipart/form-data with a "file" part. |
| name | string | yes | Display name, up to 80 chars. |
| base_voice | string | — | For kind="kokoro": the voice id to bundle, e.g. "af_heart". |
| file | file | — | For kind="clone": 3–40 seconds of clean speech (aim for ~12s). Any common audio format. |
| tone_preset | string | — | Persona tone, e.g. "warm". Optional. |
| style_text | string | — | Free-text instructions, up to 4,000 chars. Optional. |
| language | string | — | Language code from GET /v1/voices. Default "en-us". |
# Create a persona from an existing voice (JSON — no clip needed)
curl -s https://platform.naslabs.ai/v1/personas \
-H "Authorization: Bearer $NIA_KEY" \
-H "Content-Type: application/json" \
-d '{"name":"Freshpack support","base_voice":"af_heart","language":"en-us",
"tone_preset":"warm","style_text":"You are the Freshpack store assistant..."}'
# Clone a voice from a reference clip (multipart)
curl -s https://platform.naslabs.ai/v1/personas \
-H "Authorization: Bearer $NIA_KEY" \
-F kind=clone -F name="Store manager" -F language=en-us \
-F file=@reference.wav
# List / fetch / retune / delete
curl -s https://platform.naslabs.ai/v1/personas -H "Authorization: Bearer $NIA_KEY"
curl -s https://platform.naslabs.ai/v1/personas/p_ab12 -H "Authorization: Bearer $NIA_KEY"
curl -s -X PATCH https://platform.naslabs.ai/v1/personas/p_ab12 \
-H "Authorization: Bearer $NIA_KEY" -H "Content-Type: application/json" \
-d '{"style_text":"Updated brief..."}'
curl -s -X DELETE https://platform.naslabs.ai/v1/personas/p_ab12 -H "Authorization: Bearer $NIA_KEY"{
"persona": {
"id": "p_ab12cd34ef56",
"name": "Freshpack support",
"engine": "kokoro",
"kind": "kokoro",
"baseVoice": "af_heart",
"refText": null,
"durationMs": null,
"tonePreset": "warm",
"styleText": "You are the Freshpack store assistant...",
"language": "en-us",
"status": "ready",
"createdAt": "2026-08-02T10:12:00.000Z",
"lastUsedAt": null
}
}Live voice with your own brain
A live call runs on our infrastructure and its audio is peer-to-peer, so the obvious question is how your own knowledge gets into it mid-conversation. The answer is Mode B: mint the session with "mode": "external" and the built-in LLM is left out of the pipeline. You get every finished user turn as it happens, run whatever retrieval and model you like on your side, and push the reply back over the data channel for Nia to speak.
It needs no webhook, no inbound endpoint and no timeout tuning on our side — the connection is already open in both directions. And you see every turn, not only the ones a model decided to call a tool on. Turn-taking, barge-in and transcription are unchanged; only the LLM is removed.
The trade-off: in Mode B you own every word Nia says, so reply latency is your round trip and our persona work is out of the loop. If you want our brain to answer and merely consult your data, that's tool calling — not shipped yet; talk to us.
// Mode B: Nia is the ears and the voice, your server is the brain.
// Mint the session with mode: "external" — the built-in LLM is left out of the
// pipeline entirely (so you're not paying its latency), while VAD, turn-taking,
// transcription and barge-in all stay exactly as they are.
const session = await fetch("/my-server/nia-session").then((r) => r.json());
// POST /v1/realtime/session { mode: "external", voice: "af_heart", language: "en-us" }
const client = new PipecatClient({
transport: new SmallWebRTCTransport({ iceServers: session.ice_servers }),
enableMic: true,
callbacks: {
// Every finished user turn arrives here. This is the hook: run your own
// retrieval + LLM, then push the answer back for Nia to say.
onUserTranscript: async (t) => {
if (!t.final) return;
const answer = await fetch("/my-server/rag", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ question: t.text }),
}).then((r) => r.text());
client.sendClientMessage("speak", { t: "speak", d: { text: answer } });
},
},
});
await client.connect({
webrtcRequestParams: {
endpoint: `${session.url}?token=${encodeURIComponent(session.token)}`,
requestData: { settings: session.settings },
},
});
// Retrieval takes a second or two? Say something first — it's just another
// speak, and the interrupt flag lets the real answer cut it off when it lands.
// client.sendClientMessage("speak", { t: "speak", d: { text: "Let me check that." } });
// client.sendClientMessage("speak", { t: "speak", d: { text: answer, interrupt: true } });Languages & voices
Every language below works on /v1/speech — pass its code as language. This table is pulled live from the running engine, so it always matches GET /v1/voices.
Couldn't reach the engine to list languages. Start it (python -m nia.server) or call GET /v1/voices for the current catalog.
Errors
Errors are JSON with a consistent shape:
{
"error": {
"type": "unauthorized",
"message": "Invalid or revoked API key."
}
}| Status | Type | When |
|---|---|---|
| 400 | invalid_request | Missing/invalid body (e.g. no "text"). |
| 401 | unauthorized | Missing, unknown, or revoked API key. |
| 403 | forbidden | The key is valid but lacks the endpoint's scope. |
| 429 | rate_limited | Too many requests — see the Retry-After header. |
| 502 | upstream_error | The voice engine or model backend was unreachable. |
Rate limits
Each key is limited per minute (default 60 requests/minute). Exceeding it returns 429 with a Retry-After header telling you how many seconds to wait.
The limit counts requests you make to us. A live call is one request — the mint — however long it lasts, and its audio never touches the API. There is no fixed ceiling on concurrent live sessions per key today; the real limit is engine capacity, so tell us the number of simultaneous calls you expect and we'll confirm it against the deployment rather than quote an unmeasured figure.
The limit is enforced in-memory per node — fine for local/single-process use. A multi-process deployment would move this to a shared store.
Data & retention
What a call leaves behind, precisely:
- Audio: never stored. Live audio streams peer-to-peer over WebRTC and is never written to disk.
- Conversation context: memory only. A running call holds a trimmed window of recent turns to keep the thread; it dies with the call.
- Usage records: no content. Route, key, duration, voice, language, status and token counts — never what was said.
- Process logs: transcripts, by default. The engine logs each transcribed turn and reply, because that is what makes a stalled call diagnosable. A deployment that must not retain content anywhere can set
NIA_LOG_TRANSCRIPTS=0, which keeps the diagnostic lines but replaces the words with a character count. - Batch endpointsprocess in memory and return the result; inputs and outputs aren't persisted, only the usage row.
- Personas are the one thing stored on purpose: a cloned voice keeps its reference clip until you
DELETE /v1/personas/{id}.
Status
Last verified 22 August 2026 against production. Every figure below was measured, not estimated.
Operational: /v1/voices, /v1/speech, /v1/transcribe, /v1/chat, /v1/translate, /v1/dub, /v1/embeddings, /v1/personas, /v1/realtime/session, /v1/realtime/tts, /v1/realtime/stt, /v1/realtime/stt/auto, /v1/realtime/stt/lang.
Transcription latency, measured on ~9 second clips:
| Language | Pinned | Auto-detect |
|---|---|---|
| English | 2.1s | 2.2s |
| Hindi | 2.6s | 3.5s |
| Tamil | 2.9s | 3.9s |
| Malayalam | 2.9s | 10–16s |
| Arabic | 2.3s | — |
Auto-detect returned the correct language on every one of those clips. Send language whenever you know it — it skips identification entirely, and it is both faster and more accurate.
- Speech covers all 33 languages. Kokoro voices (English and European) return in 0.6–0.8s, the cloud voices used for Tamil, Malayalam, Kannada and the 12 Arabic dialects in 1.3–2.4s, and the MMS voices (Telugu, Bengali, Marathi, Gujarati, Assamese, Urdu, Odia, Maithili, Punjabi) in 0.8–1.6s.
- Send text in the language's own script. The Indic and Arabic voices are script-locked single-speaker models. Latin text sent to a Malayalam or Arabic voice is not transliterated — it returns a near-empty clip or an error. This is the most common cause of “the API returned 200 but there is no sound”.
- Language identification needs a moment of speech. Past roughly 1.5–3 seconds it is reliable. Below that it can be confidently wrong, and languages sharing a script or a phonetic neighbourhood are the usual confusion. If your callers speak in short bursts of a known language, pin it.
- Live voice. Every language in the catalog now speaks on a live call, including Tamil, Malayalam, Kannada and the Arabic dialects, which previously produced a turn with text but no audio. Two things to design around: Indic auto-detect costs a few seconds per turn (pin the language for the tightest latency), and a speaker who switches language mid-sentence stays in the language the utterance started in.
Recently resolved
- 22 Aug 2026 — Silent turns on Tamil, Arabic and Kannada live calls. Picking one of these languages produced a reply you could read on the data channel but could not hear: the cloud voice used for them had no local fallback, so a cold or failed synthesis left the turn silent with no error. Fixed by adding a local fallback voice and warming the cloud connection at startup. No client change was needed.
- 22 Aug 2026 — /v1/transcribe returning 422 for every upload. For about an hour every clip — including short English ones — came back 422 “Couldn’t read that audio”. The audio was fine: the transcription host had run out of GPU memory and was reporting it as an unreadable-file error. Resolved by restarting the host and moving a synthesis model off that GPU. Files that failed in that window work now.
- 4 Aug 2026 — Live calls connected but stayed silent. Every live session failed to start its pipeline — the transcription model was loaded onto the GPU once per call instead of once per process, so after roughly ten calls the card was full. Because signaling and the data channel come up independently, callers saw a connected session with no events and no error. Fixed by sharing one model across connections. The batch and streaming routes were never affected.
- 2 Aug 2026 — Text generation outage. /v1/chat, /v1/translate, /v1/dub and the live agent’s replies were failing because the engine host’s language model had lost part of its install. Resolved; no client change was needed.
Known issues
- Uploads over ~4.5 MB are rejected. The platform edge returns a plain-text 413 before the request reaches us. Compress, split, or use the streaming transcription socket, which has no body limit. This goes away when /v1 moves onto the engine host directly.
- One GPU, no queue. Long jobs block each other. Serialize load tests, and tell us your expected concurrency rather than assuming a ceiling.