For the complete documentation index, see llms.txt. This page is also available as Markdown.

Speech Generation

Text To Speech was a standalone page. Text To Music was another standalone page. Multi-speaker dialogue (which required a paid plan) was a tab inside Text To Speech. Three separate destinations for what is really one continuous surface: turning text into sound.

The new AI Audio Studio unifies all of this. When you arrive, you choose a mode card: Single speaker, Multi-speaker, or Music. Everything else — the prompt box, the right panel controls, the model row — reshapes itself around your choice. One page, three modes, zero context-switching.

Before

Old Text To Speech page showing Generated Audios list on the left with past entries and voice labels, and Single-Speaker/Multi-Speaker tabs. Right area shows Voice + Accent selectors, Style Prompt field, Your Script textarea, and Get Started With chips.

The old Text To Speech page had your generation history on the left and controls on the right — a fine layout, but it shared no visual connection with Text To Music, which was an entirely separate page.

After

New AI Audio Studio showing three mode cards at top: Single speaker (blue mic icon), Multi-speaker (green speech bubble icon), Music (orange note icon). Below the cards is a model row with Flash TTS, Pro TTS, Flash v2.5 ElevenLabs. Right panel shows Model dropdown, Voice, Accent, Style Prompt fields. Prompt box at bottom reads 'Write what you want the voice to say'.

The new Audio Studio presents the three modes as equal choices. Pick a mode and everything — the prompt, the model row, the right panel — adapts to it. Single speaker and multi-speaker share the same voice controls; Music replaces them with track settings.

Multi-speaker

Multi-speaker mode lets you script a dialogue with multiple voices — ideal for podcast-style content, interview simulations, explainer videos, or any content that needs more than one narrator. You assign a distinct voice and gender to each speaker role, then write the script with speaker labels.

Multi-speaker mode selected

Audio Studio with Multi-speaker mode card highlighted. The prompt box now shows a sample multi-speaker dialogue with Speaker 1 and Speaker 2 lines. Generate button at bottom right.

Selecting Multi-speaker updates the prompt box with a dialogue template, so you can see the expected format immediately. The right panel switches to show per-speaker voice assignment controls.

Generated output

Multi-speaker output showing waveform audio player with speaker labels (Speaker 1: Zephyr, Speaker 2: Puck) visible above the waveform. Play, download, and share controls below.

The generated audio file shows each speaker labelled with their voice name above the waveform. You can listen back, download the full clip, or regenerate individual segments if one voice needs adjusting.

Was
Now — and why it changed

Audio & Video → Text To Speech (separate page)

Left rail → AudioSingle speaker mode (default on arrival)

Single-Speaker tab / Multi-Speaker (Pro) tab

Three mode cards: Single speaker · Multi-speaker · Music — all visible, all equal

Model dropdown: Flash TTS (faster)

Model row: Flash TTS (Gemini), Pro TTS (Gemini), Flash v2.5 (ElevenLabs) — same models, clearer labelling

Voice + Accent selectors, Style Prompt, Your Script

Same controls in the right panel and prompt area — nothing removed

Get Started With chips (Narration, Social media ad…)

Same starter chips below the prompt box

Generated Audios list in left sidebar

Sidebar Recent list + History button top-right of the Audio Studio

Last updated