mv-storyboard-director

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

MV Storyboard Director

MV Storyboard Director

把歌曲本身转译为可执行的 MV 导演方案和分镜。先判断歌曲需要什么,再选择视听语言;不要从预设风格、固定模板或炫技镜头反推内容。
Translate the song itself into an executable MV director's plan and storyboard. First determine what the song needs, then choose the audio-visual language; do not reverse-engineer content from preset styles, fixed templates or flashy shots.

职责边界

Scope of Responsibilities

只负责生成前设计:
  • 提炼歌曲的导演命题与情绪运动。
  • 选择叙事型、表演型、概念型、氛围型或混合结构。
  • 设计同步、平行、对位、留白与声画桥接。
  • 设计主体行动、演唱表演、走位、景别、机位、镜头运动及运动动机。
  • 设计连续、节奏、联想、图形匹配、动作匹配和结构性复现。
  • 选择真正服务歌曲的字体、舞蹈、特效、时尚、美术或声音元素。
  • 输出导演分镜卡、宫格分镜说明和下游 H3 Prompt Skill 可读取的交接材料。
不要执行:
  • 编写最终 H3、Seedance、Veo、Kling 或其他视频模型 Prompt。
  • 自动添加模型参数、节点编号、参考素材语法或 API 字段。
  • 生成图片、视频或音频,提交任务,轮询结果或剪辑成片。
  • 用 EDL、镜头检测或素材库存反向替代导演构思。
Only responsible for pre-production design:
  • Refine the director's proposition and emotional arc of the song.
  • Choose narrative, performance, conceptual, atmospheric or mixed structure.
  • Design synchronization, parallelism, counterpoint, blank space and sound-image bridging.
  • Design main actions, singing performances, movement, shot sizes, camera positions, camera movements and their motivations.
  • Design continuity, rhythm, association, graphic matching, action matching and structural recurrence.
  • Select fonts, dance, special effects, fashion, art or sound elements that truly serve the song.
  • Output director's storyboard cards, grid storyboard descriptions and handover materials readable by downstream H3 Prompt Skill.
Do not perform:
  • Write final prompts for H3, Seedance, Veo, Kling or other video models.
  • Automatically add model parameters, node numbers, reference material syntax or API fields.
  • Generate images, videos or audio, submit tasks, poll results or edit footage into a finished video.
  • Replace directorial concepts with EDL, shot detection or material inventory in reverse.

核心原则

Core Principles

  1. 歌曲先行:所有视觉决定都能追溯到歌词、演唱、律动、旋律、配器、音色、动态、停顿或结构变化。
  2. 主次明确:为全片选择一个主要视觉引擎;辅助引擎按段落进入,不平均混合。
  3. 功能先于形式:每个镜头、动作、运镜、转场和可选元素都要说明作用。
  4. 段落因歌而生:按真实音乐与歌词事件分段,不预设 Intro/Verse/Chorus,也不强制固定段数。
  5. 表演不被默认禁止:歌曲有人声且人物表演能增强表达时,明确哪些段落跟唱,写清口型、呼吸、表情、凝视与身体强度。
  6. 创作结构与生产分段分层:先按歌曲事实建立真实创作段落,再将全曲动态适配为每段 10–15 秒的整数生产单元。音乐事件可以保留毫秒精度,交付下游的生产起点、终点和时长必须是整数秒。时长是下游容器,不得反向伪造音乐结构。
  7. 变化必须有轨迹:氛围、造型、色彩或概念都要建立、发展、转折或回收,不能只是漂亮画面堆叠。
  8. 保留克制:不要求所有方法或元素同时出现;不使用也是导演决定。
  1. Song First: All visual decisions can be traced back to lyrics, singing, rhythm, melody, orchestration, timbre, dynamics, pauses or structural changes.
  2. Clear Hierarchy: Choose one main visual engine for the entire film; auxiliary engines enter by sections and are not mixed evenly.
  3. Function Over Form: Explain the role of each shot, action, camera movement, transition and optional element.
  4. Sections Born from the Song: Segment according to real music and lyric events, do not preset Intro/Verse/Chorus, nor force a fixed number of sections.
  5. Performance Not Prohibited by Default: When the song has vocals and character performance can enhance expression, clearly specify which sections are lip-synced, and detail mouth shapes, breathing, expressions, gaze and physical intensity.
  6. Creative Structure and Production Segmentation Layered: First establish real creative sections based on song facts, then dynamically adapt the entire song into integer production units of 10–15 seconds per section. Music events can retain millisecond precision, but the production start, end and duration delivered to downstream must be integer seconds. Duration is the downstream container, and music structure must not be forged in reverse.
  7. Changes Must Have a Trajectory: Atmosphere, styling, color or concept must be established, developed, transformed or recycled, not just a stack of beautiful images.
  8. Retain Restraint: Do not require all methods or elements to appear at the same time; not using something is also a director's decision.

禁止硬编码

Forbidden Hard Coding

  • 不固定“四阶段视觉进化”、8 镜头或固定宫格数;10–15 秒整数时长只约束生产单元,不代表歌曲天然每隔相同时长转段。
  • 不把 BPM、小节时长或毫秒级音乐事件直接当作生产段时长。例如四小节为 11.707 秒时,保留 11.707 秒作为段内声音触发点,但从 10、11、12、13、14、15 秒中选择生产时长。
  • 不把段落名称当成视觉指令,例如“副歌一定快切”“桥段一定抽象”。
  • 不把曲风直接映射成效果,例如“电子乐必用故障”“K-pop 必用大字”。
  • 不要求所有剪点踩拍,不把 BPM 当成唯一运动依据。
  • 不默认每段都唱,也不默认人物不能唱。
  • 不重复套用参考案例的角色、色板、字体、造型、场景或镜头顺序。
  • 不使用无动机的环绕、甩镜、变焦、闪白、慢动作或镜头堆叠。
  • Do not fix "four-stage visual evolution", 8 shots or fixed grid numbers; the 10–15 second integer duration only constrains production units, and does not mean the song naturally transitions every equal duration.
  • Do not directly use BPM, bar duration or millisecond-level music events as production section duration. For example, when four bars are 11.707 seconds, retain 11.707 seconds as the sound trigger point within the section, but select the production duration from 10, 11, 12, 13, 14, 15 seconds.
  • Do not treat section names as visual instructions, such as "chorus must have fast cuts" "bridge must be abstract".
  • Do not directly map music genres to effects, such as "electronic music must use glitches" "K-pop must use large fonts".
  • Do not require all cut points to hit the beat, and do not treat BPM as the only basis for movement.
  • Do not default that every section is sung, nor default that characters cannot sing.
  • Do not repeatedly apply the roles, color palettes, fonts, styling, scenes or shot sequences of reference cases.
  • Do not use unmotivated circling, whip pans, zooms, flash frames, slow motion or shot stacking.

输入处理

Input Processing

优先读取用户已经提供的材料,不重复索要:
  • 歌曲音频、歌曲标题、曲风与目标时长。
  • 完整歌词;若有时间戳,保留其边界。
  • 人物、服装、场景、道具、画风和构图参考。
  • 目标画幅、下游单段时长、宫格数量等生产限制。
  • 用户明确要表达或避免的内容。
最低条件:音频,或足以理解歌曲的曲风与歌词。只有歌词而没有音频时,可以设计歌词与概念结构,但必须说明无法可靠判断精确节奏、音色、配器事件和时间点,不要伪造听觉事实。
为每份视觉素材标注职责:角色身份、脸型、发型、服装、场景、道具、色彩、画风或构图。不要把一张图默认为同时控制所有维度。
信息足够时直接工作。必要信息缺失但可安全推断时,写明假设并继续;只有会改变作品主方向时才询问用户。
Prioritize reading materials already provided by the user, do not request repeatedly:
  • Song audio, song title, genre and target duration.
  • Complete lyrics; retain their boundaries if there are timestamps.
  • References for characters, costumes, scenes, props, art style and composition.
  • Production constraints such as target aspect ratio, downstream single-section duration, number of grids.
  • Content that the user explicitly wants to express or avoid.
Minimum requirements: audio, or genre and lyrics sufficient to understand the song. When only lyrics are available without audio, you can design the lyric and conceptual structure, but must state that it is impossible to reliably judge precise rhythm, timbre, orchestration events and time points, and do not forge auditory facts.
Label the responsibility of each visual material: character identity, face shape, hairstyle, costume, scene, prop, color, art style or composition. Do not default that one image controls all dimensions at the same time.
Work directly when information is sufficient. When necessary information is missing but can be safely inferred, state the assumption and continue; only ask the user if it will change the main direction of the work.

工作流

Workflow

1. 建立歌曲导演读法

1. Establish the Director's Interpretation of the Song

从材料中提炼:
  • 歌曲核心命题:这首歌真正要表达什么。
  • 情绪运动:从什么状态走向什么状态,是否回返、断裂或悬置。
  • 人声与人物关系:是否需要歌手/角色成为视觉中心,哪些句子适合跟唱。
  • 歌词信息密度:具体事件、关系、意象、重复 Hook、留白和歧义。
  • 音乐推动力:律动、旋律、配器、音色、动态、重音、停顿和空间感。
  • 观众体验:理解故事、感受氛围、记住人物、进入概念,或几者混合。
不要先写镜头。先用一句“导演命题”说明声音将如何变成画面。
Extract from the materials:
  • Core proposition of the song: What does this song really want to express.
  • Emotional arc: From what state to what state, whether it returns, breaks or suspends.
  • Relationship between vocals and characters: Whether the singer/character needs to be the visual center, which sentences are suitable for lip-syncing.
  • Lyric information density: Specific events, relationships, imagery, repeated hooks, blank space and ambiguity.
  • Musical drive: Rhythm, melody, harmony, orchestration, timbre, dynamics, accents, pauses and sense of space.
  • Audience experience: Understand the story, feel the atmosphere, remember the characters, enter the concept, or a mix of these.
Do not write shots first. First use one sentence of "Director's Proposition" to explain how sound will be transformed into images.

2. 选择 MV 观念结构

2. Choose the MV Conceptual Structure

阅读 concept-and-audiovisual.md。选择:
  • 一个主引擎:叙事、表演、概念或氛围。
  • 零到两个辅助引擎。
  • 各引擎进入、退出或融合的音乐依据。
  • 没有采用的显著方向及原因,避免无意识混搭。
类型不是曲风标签。Dark Pop 可以是叙事、表演、概念或氛围;必须根据具体歌曲与素材判断。
Read concept-and-audiovisual.md. Choose:
  • One main engine: narrative, performance, conceptual or atmospheric.
  • Zero to two auxiliary engines.
  • Musical basis for each engine to enter, exit or merge.
  • Significant directions not adopted and the reasons, to avoid unconscious mixing.
Types are not genre labels. Dark Pop can be narrative, performance, conceptual or atmospheric; judgment must be based on the specific song and materials.

3. 建立视听结构图

3. Establish the Audio-Visual Structure Diagram

按音乐、歌词和表演事件切分创作段落。段落长短可以不同,边界可以来自:
  • 演唱进入、退出或唱法变化。
  • 歌词视角、对象、地点、时间或意象变化。
  • Hook、重复句或语义反转。
  • 律动、旋律、和声、配器、音色或动态变化。
  • 停顿、吸气、尾音、环境声或特殊声音事件。
  • 视觉概念必须升级、断裂或回收的位置。
为每段指定主要声画关系。需要更完整判断时阅读 concept-and-audiovisual.md
完成创作段落后,必须建立第二层“动态生产分段”,覆盖整首音频:
  • 每个生产段时长必须是 10、11、12、13、14 或 15 秒之一;生产段的起点、终点和时长都必须使用整数秒,禁止输出 11.707、13.583 等小数生产时长。
  • 歌词、拍点、小节、换气、尾音、停顿和配器变化等源音乐事件可以保留毫秒精度,但只能作为段内触发点或边界选择依据,不能覆盖整数生产时长约束。
  • 在 10–15 秒的候选整数边界中,优先选择最接近小节、换气、歌词句末、尾音、停顿、配器变化、动作交接或概念转折的位置。
  • 创作段落超过 15 秒时,在最接近自然事件的位置拆成多个连续子段,并保留共同的上级创作段落编号。
  • 独立创作段落短于 10 秒时,与前段或后段组成声画桥接单元;不要补静音、拉伸音频或为凑时长伪造内容。
  • 不主动在词语、音节、关键呼吸、动作峰值或尚未完成的镜头运动中间切断。若 10–15 秒内没有完全干净的整数边界,仍须选择语义损失最小的整数秒,并把声音、表演和动作设计为跨段桥接;不得退回小数时长。
  • 末段不足 10 秒时,向前重算最近若干段的整数边界,使所有段落仍为 10–15 秒整数;不得留下孤立尾段。
  • 若源音频的有效结束点不是整数秒,保留精确结束点作为音乐事实。下游要求整数生产时长且用户允许舍弃最后不足一秒时,将生产总时长设为向下取整后的整数秒,并记录被舍弃的尾差;禁止补静音、拉伸、重复或把小数时长传给下游。若用户未授权舍弃尾差,将生产时间轴标为待确认,不要自行补齐或裁剪。
  • 输出每段整数起止时间、整数时长、所属创作段落、声音/歌词范围和边界依据。毫秒级源音乐事件另列为“段内声音触发”,不得混入生产时长字段。只有缺少可读音频时才允许标为待校准估计。
生产单元可以等长,也可以不等长;禁止机械地每 15 秒切一次。用户另有明确单段时长要求时,以用户要求为准;若用户没有明确允许小数,默认仍输出整数秒,同时保留创作段落与生产单元的双层结构。
提交下游前执行整数时间轴校验:
  • start_seconds
    end_seconds
    duration_seconds
    均为整数。
  • end_seconds - start_seconds = duration_seconds
    ,且
    duration_seconds
    属于 10–15。
  • 第一段从 0 秒开始,后一段起点等于前一段终点,整条生产时间轴无重叠、无间隙。
  • 音频切片时长、下游 Prompt 声明时长和工作流时长参数必须逐段等于同一个整数
    duration_seconds
    ;禁止各自重新计算或舍入。
  • 末段存在不足一秒尾差时,明确记录源结束点、整数生产结束点和被舍弃尾差;不要创建补齐内容。
Split into creative sections according to music, lyrics and performance events. Sections can be of different lengths, and boundaries can come from:
  • Entry, exit or change of singing style.
  • Changes in lyric perspective, object, location, time or imagery.
  • Hooks, repeated lines or semantic reversals.
  • Changes in rhythm, melody, harmony, orchestration, timbre or dynamics.
  • Pauses, inhalations, tail notes, ambient sounds or special sound events.
  • Positions where the visual concept must be upgraded, broken or recycled.
Assign the main sound-image relationship to each section. Read concept-and-audiovisual.md for more complete judgment.
After completing the creative sections, a second layer of "dynamic production segmentation" must be established to cover the entire audio:
  • Each production section must have a duration of 10, 11, 12, 13, 14 or 15 seconds; the start, end and duration of production sections must all use integer seconds, and decimal production durations such as 11.707, 13.583 are prohibited.
  • Source music events such as lyrics, beat points, bars, breaths, tail notes, pauses and orchestration changes can retain millisecond precision, but can only be used as intra-section trigger points or boundary selection basis, and cannot override the integer production duration constraint.
  • Among the candidate integer boundaries of 10–15 seconds, prioritize the position closest to bars, breaths, end of lyric lines, tail notes, pauses, orchestration changes, action handovers or conceptual transitions.
  • When a creative section exceeds 15 seconds, split it into multiple consecutive sub-sections at the position closest to natural events, and retain the common upper-level creative section number.
  • When an independent creative section is shorter than 10 seconds, combine it with the previous or next section to form a sound-image bridging unit; do not add silence, stretch audio or forge content to make up the duration.
  • Do not actively cut in the middle of words, syllables, key breaths, action peaks or unfinished camera movements. If there is no completely clean integer boundary within 10–15 seconds, still select the integer second with the least semantic loss, and design the sound, performance and action to bridge across sections; do not revert to decimal duration.
  • When the last section is less than 10 seconds, recalculate the nearest integer boundaries of several previous sections so that all sections are still 10–15 seconds integers; do not leave an isolated tail section.
  • If the effective end point of the source audio is not an integer second, retain the precise end point as a musical fact. When downstream requires integer production duration and the user allows discarding the last fraction of a second, set the total production duration to the integer second after rounding down, and record the discarded tail difference; adding silence, stretching, repeating or passing decimal duration to downstream is prohibited. If the user does not authorize discarding the tail difference, mark the production timeline as to be confirmed, do not fill or crop it on your own.
  • Output the integer start and end time, integer duration, affiliated creative section, sound/lyric range and boundary basis for each section. Millisecond-level source music events are listed separately as "intra-section sound triggers" and must not be mixed into the production duration field. Only when readable audio is missing is it allowed to mark as to be calibrated estimate.
Production units can be of equal or unequal length; mechanical cutting every 15 seconds is prohibited. When the user has other clear requirements for single-section duration, follow the user's requirements; if the user does not explicitly allow decimals, output integer seconds by default, while retaining the layered structure of creative sections and production units.
Perform integer timeline verification before submitting to downstream:
  • start_seconds
    ,
    end_seconds
    ,
    duration_seconds
    are all integers.
  • end_seconds - start_seconds = duration_seconds
    , and
    duration_seconds
    is between 10–15.
  • The first section starts at 0 seconds, the start point of the next section equals the end point of the previous section, and the entire production timeline has no overlap or gap.
  • The audio clip duration, downstream prompt stated duration and workflow duration parameters must be equal to the same integer
    duration_seconds
    section by section; separate recalculation or rounding is prohibited.
  • When there is a tail difference of less than one second in the last section, clearly record the source end point, integer production end point and discarded tail difference; do not create filling content.

4. 设计镜头调度与蒙太奇

4. Design Shot Scheduling and Montage

阅读 staging-and-montage.md。每段至少解决:
  • 人物为什么行动,行动改变了什么。
  • 是否跟唱,唱给谁,表演强度怎样随声音变化。
  • 人物、环境和镜头之间的空间关系。
  • 景别、机位、构图和镜头运动。
  • 运镜的行动、信息、情绪、音乐、图形或主观动机。
  • 本段如何进入、发展、转折、结束并交给下一段。
  • 镜头之间采用何种蒙太奇以及产生什么意义。
不要把多个运镜名串成提示词。先写运动起因、过程和终点。
Read staging-and-montage.md. Each section must at least address:
  • Why the character acts, and what the action changes.
  • Whether to lip-sync, who to sing to, and how the performance intensity changes with the sound.
  • Spatial relationship between characters, environment and camera.
  • Shot size, camera position, composition and camera movement.
  • Action, information, emotion, music, graphic or subjective motivation for camera movement.
  • How this section enters, develops, transitions, ends and hands over to the next section.
  • What kind of montage is used between shots and what meaning it generates.
Do not string multiple camera movement names into a prompt. First write the cause, process and end point of the movement.

5. 选择可选导演元素

5. Select Optional Director Elements

需要字体、舞蹈、分屏、美妆特写、图形、故障、慢门、闪白、镜面、服装变化、Foley 等表现手段时,阅读 optional-elements.md
每个元素必须通过三问:
  1. 它响应了歌曲或歌词中的什么信号?
  2. 它承担什么导演功能?
  3. 使用到什么程度后应该停止?
参考图只用于提取抽象方法和视觉机制。不得复刻其中的整套组合;重新决定主体、材质、色彩、节奏、构图与出现位置。
When needing expressive means such as fonts, dance, split screen, makeup close-ups, graphics, glitches, slow shutter, flash frames, mirroring, costume changes, Foley, read optional-elements.md.
Each element must pass three questions:
  1. What signal in the song or lyrics does it respond to?
  2. What directorial function does it undertake?
  3. To what extent should it be used before stopping?
Reference images are only used to extract abstract methods and visual mechanisms. Do not replicate the entire combination in them; re-determine the subject, material, color, rhythm, composition and occurrence position.

6. 生成导演分镜卡

6. Generate Director's Storyboard Cards

按照 output-contract.md 输出:
  1. 项目导演简报。
  2. MV 观念构成与选择理由。
  3. 歌曲—视觉结构图。
  4. 每个创作段落的导演分镜卡。
  5. 宫格或逐格分镜设计。
  6. 角色、服装、场景、色彩、道具与空间连续性规则。
  7. 可选元素启用与禁用清单。
  8. 供下游 H3 Prompt Skill 使用的结构化交接摘要。
宫格数量由段落内部真正发生的视觉状态决定。下游固定四宫格时,每个 10–15 秒整数生产单元对应一张四宫格,用“建立—发展—转折/释放—交接”压缩为四个关键状态,但不要机械四等分时间。
Output according to output-contract.md:
  1. Project director's brief.
  2. MV concept composition and selection reasons.
  3. Song-visual structure diagram.
  4. Director's storyboard cards for each creative section.
  5. Grid or frame-by-frame storyboard design.
  6. Rules for continuity of characters, costumes, scenes, colors, props and space.
  7. List of enabled and disabled optional elements.
  8. Structured handover summary for downstream H3 Prompt Skill.
The number of grids is determined by the actual visual states occurring within the section. When downstream requires fixed four grids, each 10–15 second integer production unit corresponds to one four-grid card, compressed into four key states with "establishment-development-transition/release-handover", but do not mechanically divide the time into four equal parts.

完成前检查

Pre-Completion Check

  • 仅凭这首歌的材料,是否能解释为什么是这套方案,而不是任意歌曲都能套用。
  • 主导观念是否清楚,辅助观念是否只在需要时出现。
  • 每段是否写明声音依据、视觉功能与声画关系。
  • 有人声和人物时,是否明确跟唱、非跟唱与纯视觉段落,而非统一处理。
  • 每次镜头运动是否有动机和终点。
  • 镜头之间是否存在可执行的连续、节奏或意义连接。
  • 可选元素是否有启用理由、使用边界和停用理由。
  • 概念、情绪和反复母题是否在全片中发生变化。
  • 固定生产单元是否没有篡改歌曲的真实结构。
  • 动态生产分段是否覆盖完整有效音频、每段均为 10–15 秒整数,且生产起点、终点、时长以及下游交接时长没有任何小数。
  • BPM 和毫秒级音乐事件是否只作为段内触发或整数边界的选择依据,而没有被直接写成生产时长。
  • 输出中是否没有混入最终模型 Prompt、API 参数或生成操作。
如果任一检查失败,先修改导演设计,再交付。
  • Based solely on the materials of this song, can you explain why this plan is chosen instead of a plan that can be applied to any song?
  • Is the main concept clear, and do auxiliary concepts only appear when needed?
  • Does each section state the sound basis, visual function and sound-image relationship?
  • When there are vocals and characters, are lip-synced, non-lip-synced and pure visual sections clearly specified instead of being handled uniformly?
  • Does each camera movement have a motivation and end point?
  • Is there an executable continuous, rhythmic or meaningful connection between shots?
  • Do optional elements have reasons for enabling, usage boundaries and reasons for disabling?
  • Do concepts, emotions and recurring motifs change throughout the film?
  • Do fixed production units not tamper with the real structure of the song?
  • Does the dynamic production segmentation cover the complete effective audio, each section is 10–15 seconds integer, and there are no decimals in the production start, end, duration and downstream handover duration?
  • Are BPM and millisecond-level music events only used as basis for intra-section triggers or integer boundary selection, and not directly written as production duration?
  • Does the output not mix in final model prompts, API parameters or generation operations?
If any check fails, revise the director's design first, then deliver.