> For the complete documentation index, see [llms.txt](https://docs.qolaba.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.qolaba.ai/qolaba-2.0/whats-new/speech-generation.md).

# Speech Generation

Text To Speech was a standalone page. Text To Music was another standalone page. Multi-speaker dialogue (which required a paid plan) was a tab inside Text To Speech. Three separate destinations for what is really one continuous surface: *turning text into sound*.

The new **AI Audio Studio** unifies all of this. When you arrive, you choose a *mode card*: Single speaker, Multi-speaker, or Music. Everything else — the prompt box, the right panel controls, the model row — reshapes itself around your choice. One page, three modes, zero context-switching.

### Before

![Old Text To Speech page showing Generated Audios list on the left with past entries and voice labels, and Single-Speaker/Multi-Speaker tabs. Right area shows Voice + Accent selectors, Style Prompt field, Your Script textarea, and Get Started With chips.](/files/731d49bdc4352df63096908bd3a23835e768465a)

The old Text To Speech page had your generation history on the left and controls on the right — a fine layout, but it shared no visual connection with Text To Music, which was an entirely separate page.

### After

![New AI Audio Studio showing three mode cards at top: Single speaker (blue mic icon), Multi-speaker (green speech bubble icon), Music (orange note icon). Below the cards is a model row with Flash TTS, Pro TTS, Flash v2.5 ElevenLabs. Right panel shows Model dropdown, Voice, Accent, Style Prompt fields. Prompt box at bottom reads 'Write what you want the voice to say'.](/files/4f564bb5591b5cae0dc5e7efb892b9b1342cc5a8)

The new Audio Studio presents the three modes as equal choices. Pick a mode and everything — the prompt, the model row, the right panel — adapts to it. Single speaker and multi-speaker share the same voice controls; Music replaces them with track settings.

### Multi-speaker

Multi-speaker mode lets you script a dialogue with multiple voices — ideal for podcast-style content, interview simulations, explainer videos, or any content that needs more than one narrator. You assign a distinct voice and gender to each speaker role, then write the script with speaker labels.

#### Multi-speaker mode selected

![Audio Studio with Multi-speaker mode card highlighted. The prompt box now shows a sample multi-speaker dialogue with Speaker 1 and Speaker 2 lines. Generate button at bottom right.](/files/606d03bc8f33ba5175f2d1f759e7792b81a978a8)

Selecting Multi-speaker updates the prompt box with a dialogue template, so you can see the expected format immediately. The right panel switches to show per-speaker voice assignment controls.

#### Generated output

![Multi-speaker output showing waveform audio player with speaker labels (Speaker 1: Zephyr, Speaker 2: Puck) visible above the waveform. Play, download, and share controls below.](/files/36f8aad01696abeabfa8a4eb7be99dadfe1dded3)

The generated audio file shows each speaker labelled with their voice name above the waveform. You can listen back, download the full clip, or regenerate individual segments if one voice needs adjusting.

| Was                                                  | Now — and why it changed                                                                                      |
| ---------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| Audio & Video → **Text To Speech** (separate page)   | Left rail → **Audio** → **Single speaker** mode (default on arrival)                                          |
| Single-Speaker tab / Multi-Speaker (Pro) tab         | Three mode cards: **Single speaker · Multi-speaker · Music** — all visible, all equal                         |
| Model dropdown: `Flash TTS (faster)`                 | Model row: **Flash TTS (Gemini), Pro TTS (Gemini), Flash v2.5 (ElevenLabs)** — same models, clearer labelling |
| Voice + Accent selectors, Style Prompt, Your Script  | Same controls in the right panel and prompt area — nothing removed                                            |
| Get Started With chips (Narration, Social media ad…) | Same starter chips below the prompt box                                                                       |
| Generated Audios list in left sidebar                | Sidebar **Recent** list + **History** button top-right of the Audio Studio                                    |
