Name the sound you want
A useful video prompt can include four kinds of sound:
You do not need all four. A quiet room with one line of dialogue may only need speech and room tone. A product clip may need effects and music but no voice.
Distant traffic and rain ticking against a metal awning gives the model more to work with than moody ambience.
Name the sound source you want rather than asking for silence. Rain against the window or a soft piano bed gives the model something concrete; quiet room tone may collapse into static or dead air.
Dialogue and voiceover
Put the exact spoken words in quotation marks and say who delivers them. For a visible speaker, describe the person speaking on camera: For an off-screen line, call it a voiceover: A quoted line without a visible speaker or a voiceover cue may be treated as text that belongs in the frame. Make the distinction explicit when it matters.A visible speaker gives FLUX 3 a face to lip-sync. An off-screen voice needs a clear
voiceover or narration cue.Direct the speaker
Phrases such asprofessional, warm, and engaging leave most of the voice to the model. They often produce the same polished read: careful diction, even pauses, and more energy than the scene needs.
Give the voice a few concrete anchors instead:
Use only the details that change the read. Long stacks of personality adjectives can fight each other. Directions such as
one audible breath can also make that breath too prominent. If you want natural timing, ask for a relaxed read and judge the result by ear.
Write dialogue people can say
Voice direction cannot rescue stiff copy. Read the line aloud before generating it.- Use contractions when the character would use them.
- Cut setup the listener can already see.
- Avoid a slogan at the end of every line.
- Give the speaker a reason to say the words to someone in the scene.
- Keep punctuation simple. Too many pauses can turn into a sing-song rhythm.
Leave room for the line
Speech takes time, and the line may not start right away. A short line in a longer clip is safer than copy written to fill every second. If the final word gets cut off, shorten the line or increase the clip duration. Timing instructions can give the model a target, but they are not exact:Build the mix around the scene
Name sounds that have a source in the frame or just outside it. This keeps the audio tied to the picture. If speech is the focus, keep other voices out of the background. Crowd conversation, a talking radio, or another narrator can compete with the main line. Weather, machinery, footsteps, and traffic are easier to layer under speech because they do not introduce more words.Languages and accents
FLUX 3 can generate speech in many languages and accents, and it can switch languages within one performance. Name each language directly, place it beside the line it belongs to, and describe the delivery as you would for any other speaker. You can provide dialogue in native script, in a romanized or transliterated form, or as a plain-language instruction that names the target language and meaning. Use quoted dialogue when the exact words matter. No one format is always best, so use the one that fits how you write and review the result by ear. For one speaker changing languages:- Label each line with its language.
- Put the lines in the intended order.
- Say that the same speaker continues across the switch.
- Keep each segment short enough to leave room for a natural pause.
natural Hindi delivery, warm and relieved, or French spoken softly over a suit radio, with quiet wonder.
Language, accent, line order, speaker attribution, and timing are directable targets rather than exact controls. Review each take by ear when the performance matters.
Switch languages in one voice
The language labels, exact order, andsame voice cue make the handoff explicit without turning the prompt into a timing sheet.
Give each speaker a language
Visible roles such asorange-suited astronaut and blue-suited astronaut connect each language and line to the intended person.
Use native-script dialogue
Native-script dialogue can make the intended words explicit. Name the language as well as the quoted line, then describe how the speaker should deliver it.Keeping a voice across clips
Reuse the full voice direction when a character returns. Keep the person, register, recording setup, and delivery wording stable, then replace only the script and scene details. This can preserve the same kind of voice, but it does not guarantee the same performer on every generation. Treat the prompt as casting direction rather than a fixed speaker identity. Compare takes by ear before cutting them into the same sequence.Troubleshooting
Related pages
Text-to-Video
Build the shot, action, camera movement, pacing, and scene around the audio.
Video Generation Overview
Choose the FLUX 3 video workflow that fits the source material and shot.
FLUX 3 Video
See how synchronized audio works in the API and how to turn it off.
Camera Terms, Prompts & Examples
Pair sound direction with clear framing and camera movement.

