Inline vs Queued + Polling
How a request behaves depends on length: short text comes back immediately; long text becomes a job you poll.
At a glance
- Short text → inline audio in the response.
- Long text → a job id + poll URL.
- Poll until complete, then fetch the audio.
Inline
Below the inline threshold, synthesis returns audio directly. One request, one response, simplest for short snippets.
Queued + polling
Above the threshold, you get a job id and a poll URL. Poll GET /api/v1/tts/jobs/{id} until it's done, then download from /jobs/{id}/audio.
POST /api/v1/tts/synthesize -> { "job_id": "...", "poll": "/api/v1/tts/jobs/..." }
GET /api/v1/tts/jobs/{id} -> { "status": "running" | "done" | "failed" }
GET /api/v1/tts/jobs/{id}/audio -> the audio file