Prompt Guide: Scenema Audio Pro
Scenema Audio Pro generates speech from text. Whatever you write is what the model says. Your prompt is the script, and the model performs it exactly as written.
You control the output with two mechanisms. The voice description defines who is speaking. Action tags define how they speak at any given moment. Everything else in the prompt is spoken aloud, verbatim. Understanding this relationship is the single most important thing you can learn about writing effective prompts.
Voice Description
The voice description is the most impactful control you have. It tells the model who the speaker is, and every aspect of the generated audio follows from it. A detailed voice description produces a specific, consistent character. A vague one produces generic output.
A good voice description includes gender, approximate age, accent, and vocal quality. These four attributes anchor the voice. Beyond that, you can add personality, energy level, register, and any other characteristic that helps the model understand the speaker.
Vague: A nice female voiceSpecific: A woman, early 30s. American accent, warm alto. Conversational, unhurried. Slight vocal fry.The difference between these two descriptions is not subtle. The vague version gives the model almost nothing to work with, and the result sounds like a default text-to-speech voice. The specific version produces a recognizable character with consistent traits across the entire generation.
Here is a documentary narrator generated entirely from a voice description with no action tags at all. The voice description alone carries the performance.
Documentary Narrator
No action tags. The voice description alone produces an authoritative, measured BBC narrator.
Show prompt
Notice that there is no direction in the script itself. The voice description does all the work. For narration, explainers, and any content where the delivery is consistent throughout, a strong voice description may be all you need.
Action Tags
Action tags are inline delivery directions. They are placed in the script immediately before the passage they should affect, and they are not spoken aloud. Think of them as stage directions for a voice actor: they tell the model how to deliver the next line.
<action>Casual, warm, settling in</action>Okay so, I have been thinking about this all weekand I just have to talk about it.
<action>Quieter, conspiratorial</action>You know that feeling when you read somethingand it just... rewires your brain?Action tags describe vocal quality and emotional state. They should be short, specific descriptors of how the voice sounds in that moment. Good action tags read like adjectives: “quieter, conspiratorial” or “breathless, grinning” or “serious, thoughtful.”
Here is a podcast host whose delivery shifts five times across a single script. Each shift is driven by an action tag.
Conversational Podcast
Five distinct delivery shifts across a single script, each driven by an action tag.
Show prompt
Writing Emotions
Action tags set the emotional tone, but the model still needs actual content to perform. This is the most common source of flat output: the action tag says the right thing, but the script text gives the model nothing to work with.
The distinction is between directing and performing. The action tag directs. The script text is the performance. If you want laughter, write out the laughter. If you want hesitation, write the pauses and false starts. If you want sobbing, write the broken sentences and trailing words. The text IS the audio.
Weak: <action>She laughs</action> That was so funny.Strong: <action>Laughing, barely holding it together</action> Hahahaha... heh heh... oh man, that was so funny.
Weak: <action>She is upset</action> I miss her so much.Strong: <action>Voice breaking</action> I just... I... miss her sooo much.In the weak versions, the action tag describes what should happen but the script text is flat. There is nothing for the model to perform emotionally. The strong versions give the model actual material: written-out laughter, trailing words, hesitation, elongated vowels. The action tag and the script text work together.
Emotional Monologue
Action tags set the emotional arc. Written-out pauses and broken sentences give the model material to perform.
Show prompt
Villain Monologue
Written-out laughter in the script text, not just an action tag saying 'laughs.' The model performs what it reads.
Show prompt
In the villain example, the laughter is written out as “Hhhaaahahaha… hahahaha!” and “heh heh heh.” The action tags describe the vocal quality around the laughter (“low, sinister chuckle building,” “breathless, grinning”), but the actual laughing sounds come from the script text. If you replaced the written laughter with an action tag like <action>He laughs maniacally</action>, the model would read “He laughs maniacally” aloud.
Common Mistakes
Stage directions in the script text. Everything outside an action tag is spoken aloud. If you write narrative descriptions, physical actions, or scene directions in the body of the script, the model reads them as dialogue. This is the single most common prompting error.
Here is a real example. A user wanted an animated YouTube reviewer and wrote the prompt like a screenplay:
Wrong:
<action>Sudden full eruption, maximum volume,genuine exasperation barely disguised as a bit</action>OH MY GOD, THERE'S TOO MANY BLOODY SHOWS.
<action>Dropping back into casual reviewer mode,almost like he didn't just yell</action>Drops of God, an anime about expert sommelier wine lovers.The action tags here contain narrative stage directions (“genuine exasperation barely disguised as a bit,” “almost like he didn’t just yell”) instead of vocal descriptors. The fix is straightforward: action tags describe how the voice sounds, not what the character is doing or thinking.
Corrected:
<action>Loud, exasperated, explosive</action>OH MY GOD, THERE'S TOO MANY BLOODY SHOWS.
<action>Casual, amused, deadpan</action>Drops of God, an anime about expert sommelier wine lovers.Here is the corrected version, generated with proper action tags:
YouTube Reviewer (Corrected)
The same script with action tags rewritten as vocal descriptors instead of stage directions.
Show prompt
Vague voice descriptions. “A narrator” or “a friendly voice” gives the model almost no information. Always include gender, age, accent, and at least one vocal quality descriptor. The more specific the description, the more distinct and consistent the output.
Over-directing, under-writing. If every other line is an action tag but the script text between them is short and flat, the model has more direction than material. Action tags work best when the script gives the model enough text to actually perform the shift. A single sentence after an action tag is often too little for the delivery change to register fully.
Describing emotions instead of performing them. “He sighs heavily” as script text will be spoken as the words “he sighs heavily.” If you want a sigh, write something like “Hhhhhh…” before the next line and use an action tag like <action>Exhausted, deflated</action> to set the tone. The same applies to crying, gasping, stuttering, and any other vocal effect. Write the sound, not the description.
Multilingual
The same voice description and action tag system works in any supported language. Write the script in the target language, set the language field, and the model generates speech in that language with the same voice identity.
Spanish Storyteller
Native Castilian Spanish with poetic cadence and action tag transitions.
Show prompt
Japanese Narrator
Formal, contemplative Japanese narration with measured pacing.
Show prompt
Supported languages include English, Spanish, French, German, Japanese, Portuguese, Korean, Chinese, Hindi, Arabic, and more. The full list is available in the Audio Studio language selector.
Voice Cloning
Scenema Audio Pro supports zero-shot expressive voice cloning. Upload a short reference clip of any voice, and the model generates new speech in that voice performing your script. The reference only needs to be a few seconds of clean speech.
What makes this different from other voice cloning systems is that emotional range is not locked to the reference recording. If your reference clip is calm, the output can still shout, whisper, laugh, or cry. The emotion comes from the prompt and action tags, not from the reference. The reference provides the voice identity. The script provides the performance.
This means a single calm reference clip is enough to produce a full range of emotional performances in that voice. You do not need multiple reference recordings at different emotional registers.
Advanced Techniques
Pacing through the prompt. Scenema Audio Pro has no pace slider. All pacing control comes from the prompt itself. Write the way you want the voice to sound. Ellipses (“I just… I do not know…”) slow the delivery. Short punchy sentences speed it up. Action tags like “speaking faster, more urgent” or “slow, deliberate” give the model explicit pacing direction. The text and the action tags together determine the rhythm of the output.
Long-form consistency. For scripts longer than a few paragraphs, the voice description anchors identity across the entire duration. The model maintains the same speaker characteristics from beginning to end. Action tags handle moment-level variation within that consistent identity. For projects exceeding 30 minutes, splitting into chapter-length segments produces the most reliable results.
Shot type. The shot type selector (closeup, medium, wide) affects the perceived proximity of the voice. Closeup produces an intimate, present sound as if the speaker is right next to the microphone. Wide produces a more distant, environmental quality. For most speech generation, closeup is the right choice.
Scene description. The scene field adds environmental context. While Scenema Audio Pro does not generate sound effects (use the dedicated Foley service for that), the scene description can subtly influence the vocal performance. A scene like “quiet library” may produce softer, more restrained delivery, while “packed stadium” may produce something louder and more projected.
Scenema Audio Pro is available now in the Audio Studio. Select “Scenema Audio Pro” from the model dropdown to get started.
For technical details on the engine, read the Scenema Audio Pro announcement. For the base audio diffusion model with scene-aware sound effects, read about Scenema Audio.