Scenema Audio Pro: Direct Voice Performance Through Text
Scenema Audio Pro is a professional speech generation engine designed for creators who need precise vocal control over long-form content. It supports audiobooks, podcasts, documentaries, corporate narration, character-driven storytelling, and any project where voice performance matters.
The demo above was generated from nothing more than a voice description and a short script with inline delivery directions. No voice recording was provided. No preset was selected. The engine constructed the vocal identity, accent, pacing, and emotional arc entirely from the text.
Voice Design
Scenema Audio Pro supports voice design through natural language descriptions. The voice description defines the speaker’s identity, and the engine generates speech that matches it. There is no need to browse preset libraries or settle for a voice that approximates the intended character. The description is the voice.
A description like “Female, mid-40s with a slight rasp, Southern Louisiana drawl, unhurried, the kind of voice that sounds like it has been soaked in bourbon and campfire smoke” produces a voice with those characteristics on the first generation. Changing the description changes the voice. The relationship between description and output is direct and responsive.
We also ship a library of curated system voices, each with a distinct character and vocal identity. These voices cover a range of accents, ages, and delivery styles. Select a voice and start generating immediately, or write a custom description for a fully original result.
Inline Delivery Control
Action tags provide moment-level control over vocal performance. They are placed inline with the script text, immediately before the passage they should affect. The engine interprets these directions and adjusts delivery accordingly.
<action>low, steady, almost confessional</action>I never told anybody what happened that summer.<action>pause, exhale through nose</action>Not because I was afraid. Because some things...once you say them out loud, they become real.<action>voice drops even lower, barely above a whisper</action>And I was not ready for that.Action tags are not spoken aloud. They function as stage directions for the voice engine, controlling emotional shifts, pacing changes, breath, and physical vocal effects at specific moments in the script. This approach mirrors how directors work with voice actors: specific direction at specific moments, rather than global parameters applied uniformly across the entire performance.
Long-Form Generation
Scenema Audio Pro supports up to 30 minutes of verified single-output generation. The system handles internal chunking, merging, and post-processing automatically. A full script is submitted as a single request, and the output is a single continuous audio file with consistent voice identity across the entire duration.
The underlying architecture supports longer durations, although we recommend keeping individual generations at or below the 30-minute mark for the best experience. For projects exceeding that length, splitting into chapter-length segments produces the most reliable results.
Sound Effects
Scenema Audio, our base model, supports inline sound effect generation through prompt tags. In practice, producing high-quality sound effects this way requires precise prompting. Small changes to wording produce different results, and the output can be unpredictable for users who are not familiar with the model’s sensitivity to prompt structure.
We made the decision to remove inline sound effects from Pro and instead direct users to our dedicated Foley service. The Foley model is trained specifically for sound effect generation and produces more precise and predictable results. If a project needs environmental audio, atmospheric textures, or specific sound effects, the generated speech can be brought into the studio and layered with Foley on a per-shot basis. The result is better separation, higher quality, and full control over sound design as a distinct production stage.
Multiple Takes
Every generation produces a unique take. This is by design and consistent with how professional voice production works. Voice actors record multiple takes, and the director selects the best performance. Scenema Audio Pro operates the same way. The version history preserves every generation, and the best take can be promoted as the active version at any time.
Multilingual Support
Scenema Audio Pro supports multilingual generation across English, Spanish, French, German, Japanese, Portuguese, Korean, Chinese, Hindi, Arabic, and more. The full list of supported languages is available in the studio. The same voice description and action tags work across all of them. A voice designed in English maintains its character when generating in any supported language.
Scenema Audio Pro is available now in the Audio Studio and the Studio voiceover inspector. Select “Scenema Audio Pro” from the model dropdown to get started.
Read about Scenema Audio, our expressive audio diffusion model that powers the base tier with zero-shot voice cloning, emotional acting, and scene-aware sound effects.