Speech Generation
Before
After
Multi-speaker
Multi-speaker mode selected
Generated output
Was
Now — and why it changed
Last updated
Text To Speech was a standalone page. Text To Music was another standalone page. Multi-speaker dialogue (which required a paid plan) was a tab inside Text To Speech. Three separate destinations for what is really one continuous surface: turning text into sound.
The new AI Audio Studio unifies all of this. When you arrive, you choose a mode card: Single speaker, Multi-speaker, or Music. Everything else — the prompt box, the right panel controls, the model row — reshapes itself around your choice. One page, three modes, zero context-switching.
The old Text To Speech page had your generation history on the left and controls on the right — a fine layout, but it shared no visual connection with Text To Music, which was an entirely separate page.
The new Audio Studio presents the three modes as equal choices. Pick a mode and everything — the prompt, the model row, the right panel — adapts to it. Single speaker and multi-speaker share the same voice controls; Music replaces them with track settings.
Multi-speaker mode lets you script a dialogue with multiple voices — ideal for podcast-style content, interview simulations, explainer videos, or any content that needs more than one narrator. You assign a distinct voice and gender to each speaker role, then write the script with speaker labels.
Selecting Multi-speaker updates the prompt box with a dialogue template, so you can see the expected format immediately. The right panel switches to show per-speaker voice assignment controls.
The generated audio file shows each speaker labelled with their voice name above the waveform. You can listen back, download the full clip, or regenerate individual segments if one voice needs adjusting.
Audio & Video → Text To Speech (separate page)
Left rail → Audio → Single speaker mode (default on arrival)
Single-Speaker tab / Multi-Speaker (Pro) tab
Three mode cards: Single speaker · Multi-speaker · Music — all visible, all equal
Model dropdown: Flash TTS (faster)
Model row: Flash TTS (Gemini), Pro TTS (Gemini), Flash v2.5 (ElevenLabs) — same models, clearer labelling
Voice + Accent selectors, Style Prompt, Your Script
Same controls in the right panel and prompt area — nothing removed
Get Started With chips (Narration, Social media ad…)
Same starter chips below the prompt box
Generated Audios list in left sidebar
Sidebar Recent list + History button top-right of the Audio Studio
Last updated