Scenema Team

Prompt Guide: Scenema Audio Pro

Scenema Audio Pro generates speech from text. Whatever you write is what the model says. Your prompt is the script, and the model performs it exactly as written.

You control the output with two mechanisms. The voice description defines who is speaking. Action tags define how they speak at any given moment. Everything else in the prompt is spoken aloud, verbatim. Understanding this relationship is the single most important thing you can learn about writing effective prompts.

Voice Description

The voice description is the most impactful control you have. It tells the model who the speaker is, and every aspect of the generated audio follows from it. A detailed voice description produces a specific, consistent character. A vague one produces generic output.

A good voice description includes gender, approximate age, accent, and vocal quality. These four attributes anchor the voice. Beyond that, you can add personality, energy level, register, and any other characteristic that helps the model understand the speaker.

Vague: A nice female voice
Specific: A woman, early 30s. American accent, warm alto.
Conversational, unhurried. Slight vocal fry.

The difference between these two descriptions is not subtle. The vague version gives the model almost nothing to work with, and the result sounds like a default text-to-speech voice. The specific version produces a recognizable character with consistent traits across the entire generation.

Here is a documentary narrator generated entirely from a voice description with no action tags at all. The voice description alone carries the performance.

Documentary Narrator

No action tags. The voice description alone produces an authoritative, measured BBC narrator.

Show prompt
Voice A man, late 40s. Deep, authoritative baritone. British RP accent, measured and unhurried. BBC documentary narrator quality. The kind of voice that commands attention without raising volume.
The Amazon rainforest stretches across nine countries and covers an area larger than the entire European Union. It is home to more than ten percent of all species known to science, many of which have never been studied in detail. Every square kilometer of this vast green expanse contains thousands of species of plants, insects, birds, and mammals, each playing a role in an ecosystem so complex that researchers have spent decades attempting to understand even a fraction of its inner workings. At dawn, the canopy comes alive. Howler monkeys begin their territorial calls, a sound so deep and resonant that it carries for miles through the dense foliage. Below them, columns of leafcutter ants march along trails they have carved into the forest floor over generations.

Notice that there is no direction in the script itself. The voice description does all the work. For narration, explainers, and any content where the delivery is consistent throughout, a strong voice description may be all you need.

Action Tags

Action tags are inline delivery directions. They are placed in the script immediately before the passage they should affect, and they are not spoken aloud. Think of them as stage directions for a voice actor: they tell the model how to deliver the next line.

<action>Casual, warm, settling in</action>
Okay so, I have been thinking about this all week
and I just have to talk about it.
<action>Quieter, conspiratorial</action>
You know that feeling when you read something
and it just... rewires your brain?

Action tags describe vocal quality and emotional state. They should be short, specific descriptors of how the voice sounds in that moment. Good action tags read like adjectives: “quieter, conspiratorial” or “breathless, grinning” or “serious, thoughtful.”

Here is a podcast host whose delivery shifts five times across a single script. Each shift is driven by an action tag.

Conversational Podcast

Five distinct delivery shifts across a single script, each driven by an action tag.

Show prompt
Voice A woman, early 30s. American accent, slight California vocal fry. Warm, conversational, natural. Sounds like she is talking to a friend over coffee. Authentic, not performative.
Casual, warm, settling in
Okay so, I have been thinking about this all week and I just have to talk about it.
Quieter, conspiratorial
You know that feeling when you read something and it just... rewires your brain? Like you physically feel different after?
Amused
Heh, yeah, that happened to me on Tuesday.
Serious, thoughtful
So here is the thing. We spend so much time optimizing for productivity, for output, for metrics. And none of that is wrong. But I read this one line that said, the goal is not to be busy, the goal is to be changed. And I just sat there for like ten minutes staring at the wall.
Brighter, upbeat
Anyway, I want to hear what you all think. Drop me a message. I read every single one.

Writing Emotions

Action tags set the emotional tone, but the model still needs actual content to perform. This is the most common source of flat output: the action tag says the right thing, but the script text gives the model nothing to work with.

The distinction is between directing and performing. The action tag directs. The script text is the performance. If you want laughter, write out the laughter. If you want hesitation, write the pauses and false starts. If you want sobbing, write the broken sentences and trailing words. The text IS the audio.

Weak: <action>She laughs</action> That was so funny.
Strong: <action>Laughing, barely holding it together</action>
Hahahaha... heh heh... oh man, that was so funny.
Weak: <action>She is upset</action> I miss her so much.
Strong: <action>Voice breaking</action>
I just... I... miss her sooo much.

In the weak versions, the action tag describes what should happen but the script text is flat. There is nothing for the model to perform emotionally. The strong versions give the model actual material: written-out laughter, trailing words, hesitation, elongated vowels. The action tag and the script text work together.

Emotional Monologue

Action tags set the emotional arc. Written-out pauses and broken sentences give the model material to perform.

Show prompt
Voice A woman, late 20s. American accent. Thin, shaking voice. Trying not to cry but failing. Raw and unguarded. Each sentence more fragile than the last.
Trembling, barely audible
I keep thinking she is going to call.
Voice cracks
Every time the phone rings I think maybe... maybe this time.
Long pause
The worst part is the mornings. I still reach for the other side of the bed.
Breaking down, sobbing
I just... I just miss her so much. I do not know how to do this without her.

Villain Monologue

Written-out laughter in the script text, not just an action tag saying 'laughs.' The model performs what it reads.

Show prompt
Voice A man, late 50s. Deep theatrical baritone. British accent, stage actor energy. Savoring every word. Deranged but articulate. The kind of villain who monologues.
Low, sinister chuckle building
Hhhaaahahaha... hahahaha!
Breathless, grinning
Hhh... hhh... oh, you have no idea how long I have waited for this moment. Years. Every single day, watching, planning.
Whispered, menacing
And now... heh heh heh... now it is finally mine.
Louder, unhinged
AHAHAHA! All of it! MINE!

In the villain example, the laughter is written out as “Hhhaaahahaha… hahahaha!” and “heh heh heh.” The action tags describe the vocal quality around the laughter (“low, sinister chuckle building,” “breathless, grinning”), but the actual laughing sounds come from the script text. If you replaced the written laughter with an action tag like <action>He laughs maniacally</action>, the model would read “He laughs maniacally” aloud.

Common Mistakes

Stage directions in the script text. Everything outside an action tag is spoken aloud. If you write narrative descriptions, physical actions, or scene directions in the body of the script, the model reads them as dialogue. This is the single most common prompting error.

Here is a real example. A user wanted an animated YouTube reviewer and wrote the prompt like a screenplay:

Wrong:
<action>Sudden full eruption, maximum volume,
genuine exasperation barely disguised as a bit</action>
OH MY GOD, THERE'S TOO MANY BLOODY SHOWS.
<action>Dropping back into casual reviewer mode,
almost like he didn't just yell</action>
Drops of God, an anime about expert sommelier wine lovers.

The action tags here contain narrative stage directions (“genuine exasperation barely disguised as a bit,” “almost like he didn’t just yell”) instead of vocal descriptors. The fix is straightforward: action tags describe how the voice sounds, not what the character is doing or thinking.

Corrected:
<action>Loud, exasperated, explosive</action>
OH MY GOD, THERE'S TOO MANY BLOODY SHOWS.
<action>Casual, amused, deadpan</action>
Drops of God, an anime about expert sommelier wine lovers.

Here is the corrected version, generated with proper action tags:

YouTube Reviewer (Corrected)

The same script with action tags rewritten as vocal descriptors instead of stage directions.

Show prompt
Voice A man, late 20s. British accent, London. Energetic, expressive range from shouting to quiet sincerity. Natural comedic timing, shifts between theatrical outbursts and genuine warmth. YouTube reviewer energy.
Loud, exasperated, explosive
OH MY GOD, THERE'S TOO MANY BLOODY SHOWS.
Casual, amused, deadpan
Drops of God, an anime about expert sommelier wine lovers. And to everyone who says that's just alcoholics written in cursive.
Dramatic, reverent, mock-serious
JACOB'S CREEK CHARDONNAY 1991. DRINK.
Quieter, genuine, reflective
When I first read the manga of this I had no interest in wine. But I don't know, something about the concept of having a Japanese king of wine who died and will leave his one treasure, his vast wine collection, to the one person that can go out into the world and find the singular best bottle of wine that may or may not exist... appeal to me.

Vague voice descriptions. “A narrator” or “a friendly voice” gives the model almost no information. Always include gender, age, accent, and at least one vocal quality descriptor. The more specific the description, the more distinct and consistent the output.

Over-directing, under-writing. If every other line is an action tag but the script text between them is short and flat, the model has more direction than material. Action tags work best when the script gives the model enough text to actually perform the shift. A single sentence after an action tag is often too little for the delivery change to register fully.

Describing emotions instead of performing them. “He sighs heavily” as script text will be spoken as the words “he sighs heavily.” If you want a sigh, write something like “Hhhhhh…” before the next line and use an action tag like <action>Exhausted, deflated</action> to set the tone. The same applies to crying, gasping, stuttering, and any other vocal effect. Write the sound, not the description.

Multilingual

The same voice description and action tag system works in any supported language. Write the script in the target language, set the language field, and the model generates speech in that language with the same voice identity.

Spanish Storyteller

Native Castilian Spanish with poetic cadence and action tag transitions.

Show prompt
Voice A man, early 40s. Rich warm baritone. Native Castilian Spanish accent from Madrid. Poetic cadence, unhurried. A natural storyteller. Scene Beach at sunset, waves Language ES
Calm, warm, reflective
Mi abuela siempre decia que el mar guarda los secretos de todos los que lo miran. Yo no le creia. Hasta que una noche, sentado en la arena, escuche su voz entre las olas.
Slower, nostalgic
Me dijo que el amor no se pierde, que se transforma. Que cada ola que llega a la orilla trae consigo un recuerdo de alguien que amo ese mismo mar. Y yo, sin entender del todo, me quede ahi, escuchando, hasta que el sol desaparecio detras del agua.

Japanese Narrator

Formal, contemplative Japanese narration with measured pacing.

Show prompt
Voice A woman, 40s. Native Japanese accent. Formal, contemplative narrator. Precise diction, measured rhythm. The warmth of a museum guide with quiet elegance. Language JA
Quiet authority, measured
京都の古い寺院の庭に立つと、時間の流れが変わるのを感じます。何百年も前に植えられた松の木が、今も静かに空へ向かって伸びています。
Reflective, softer
石庭の白い砂に描かれた模様は、毎朝僧侶たちの手によって新しく作り直されます。完璧を求めるのではなく、その瞬間の美しさを受け入れること。それが、この庭が私たちに教えてくれることなのかもしれません。

Supported languages include English, Spanish, French, German, Japanese, Portuguese, Korean, Chinese, Hindi, Arabic, and more. The full list is available in the Audio Studio language selector.

Voice Cloning

Scenema Audio Pro supports zero-shot expressive voice cloning. Upload a short reference clip of any voice, and the model generates new speech in that voice performing your script. The reference only needs to be a few seconds of clean speech.

What makes this different from other voice cloning systems is that emotional range is not locked to the reference recording. If your reference clip is calm, the output can still shout, whisper, laugh, or cry. The emotion comes from the prompt and action tags, not from the reference. The reference provides the voice identity. The script provides the performance.

This means a single calm reference clip is enough to produce a full range of emotional performances in that voice. You do not need multiple reference recordings at different emotional registers.

Advanced Techniques

Pacing through the prompt. Scenema Audio Pro has no pace slider. All pacing control comes from the prompt itself. Write the way you want the voice to sound. Ellipses (“I just… I do not know…”) slow the delivery. Short punchy sentences speed it up. Action tags like “speaking faster, more urgent” or “slow, deliberate” give the model explicit pacing direction. The text and the action tags together determine the rhythm of the output.

Long-form consistency. For scripts longer than a few paragraphs, the voice description anchors identity across the entire duration. The model maintains the same speaker characteristics from beginning to end. Action tags handle moment-level variation within that consistent identity. For projects exceeding 30 minutes, splitting into chapter-length segments produces the most reliable results.

Shot type. The shot type selector (closeup, medium, wide) affects the perceived proximity of the voice. Closeup produces an intimate, present sound as if the speaker is right next to the microphone. Wide produces a more distant, environmental quality. For most speech generation, closeup is the right choice.

Scene description. The scene field adds environmental context. While Scenema Audio Pro does not generate sound effects (use the dedicated Foley service for that), the scene description can subtly influence the vocal performance. A scene like “quiet library” may produce softer, more restrained delivery, while “packed stadium” may produce something louder and more projected.


Scenema Audio Pro is available now in the Audio Studio. Select “Scenema Audio Pro” from the model dropdown to get started.

For technical details on the engine, read the Scenema Audio Pro announcement. For the base audio diffusion model with scene-aware sound effects, read about Scenema Audio.