flux-3-audio-dialogue

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

FLUX 3 Audio and Dialogue

FLUX 3 音频与对话

Name each layer separately: speech, voiceover, ambience, effects, music, or deliberate silence. One blurred description gives up control of all of them; every sound needs a physical source or narrative role.
Speech. Quote the exact line, name the visible speaker (or label the line
voiceover
/
narration
so it is not searching for a mouth to belong to), and add
no on-screen text, no subtitles
when text is unwanted:
text
A weather presenter on camera in front of a stylized storm map, speaking directly to
the lens: "Storm season is here, and this time, we're ready." Confident delivery,
clean studio lighting. No on-screen text, no subtitles.
Voice anchors: age range, accent when relevant, register, energy, recording distance. Reusing the same direction preserves a kind of voice, not the same performer across generations.
Speakability. Write for the clip's real duration: short sentences, one thought per line, room before and after the payoff; spell unusual names phonetically; shorten the line before speeding the delivery. A line that cannot finish comfortably needs a shorter script or a longer clip.
Effects are causal, not a detached list:
text
As the cup hits the tile, it cracks with one sharp ceramic snap.
Mix. Say what leads and what stays under it. Keep background voices out when one line matters:
text
Her line is foreground and fully intelligible. Café chatter and espresso hiss remain
low and diffuse. A restrained piano pulse enters beneath the final words without
masking them.
Silence and post. Set
generate_audio: false
for a deliberately silent source clip. Reserve for deterministic post: final loudness, EQ, ducking, and fades; guaranteed wording or speaker identity; frame-accurate sync; subtitles and captions; continuity across separately generated clips.
When a take misses, change one dimension at a time: speaker ownership, line length, delivery anchors, competing layers, action-to-effect causality, or generation versus post.
请分别命名每个音轨层:对话、旁白、氛围音、音效、音乐或刻意静音。模糊的描述会让你失去对所有音轨的控制权;每种声音都需要明确的物理来源或叙事作用。
对话:引用确切台词,指明可见说话人(或将台词标记为
voiceover
/
narration
,避免系统寻找对应说话的嘴部画面),若不需要文字则添加
no on-screen text, no subtitles
text
A weather presenter on camera in front of a stylized storm map, speaking directly to
the lens: "Storm season is here, and this time, we're ready." Confident delivery,
clean studio lighting. No on-screen text, no subtitles.
声音锚点:年龄范围、相关口音、语调、活力、录制距离。重复使用相同的指导会保持一种声音特质,而非固定同一表演者。
口语化适配:根据视频片段的实际时长撰写台词:短句、每句表达一个想法、在关键内容前后预留停顿;生僻名称标注音标;若要加快语速,先缩短台词。无法舒适说完的台词,要么缩短脚本,要么延长片段时长。
音效需具备因果性,而非孤立的列表:
text
As the cup hits the tile, it cracks with one sharp ceramic snap.
混音:明确说明哪个音轨为主,哪个为背景。当某段台词至关重要时,排除背景人声:
text
Her line is foreground and fully intelligible. Café chatter and espresso hiss remain
low and diffuse. A restrained piano pulse enters beneath the final words without
masking them.
静音与后期制作:对于刻意静音的源片段,设置
generate_audio: false
。将以下操作留作确定性后期处理:最终响度调整、均衡器(EQ)、闪避处理、淡入淡出;确保措辞或说话人身份准确;帧级精准同步;字幕与隐藏式字幕;不同生成片段间的连贯性。
当生成效果不理想时,每次仅调整一个维度:说话人归属、台词长度、声音锚点、竞争音轨层、动作与音效的因果关系,或是生成环节与后期制作的选择。