Audiobook Studio 2.0 docs

Studio 2.0 at a Glance

Studio 2.0 is a ground-up rearchitecture. The big idea: move synthesis into a managed, crash-isolated TTS Server, and make engines and voices first-class, installable pieces.

At a glance
  • Managed TTS Server runs engines in a separate, supervised process.
  • A task orchestrator replaces the old single worker loop.
  • An engine registry and voice bridge keep voices independent of engines.
  • Engines install from GitHub; voices come from a Hugging Face library.

A managed TTS Server

An engine can crash without taking down the app or losing your work. Studio watches the TTS Server and restarts it if it stops responding.

A task orchestrator

Work is queued and tracked in one place, and picks up where it left off after a restart.

Engine registry & voice bridge

The same voice can move between engines without being rebuilt. For the machinery behind each change, see Architectural Shifts.

Tip: You can add, update, or swap engines without rewrites rippling into the queue, the editor, or your projects.

Installable engines & a voice library

Engines and voices are bundles you can install and share. Engines are GitHub repos you install like Stable Diffusion extensions (clone to install, pull to update). Voices come from a Hugging Face library with icons, playable samples, and rich tags.

Coming soon: the engine browser and the Hugging Face voice library.

What stayed the same

The workflow. Projects, chapters, voices, generate, assemble.