h3-prompt-writing

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

H3 提示词写作(底座)

H3 Prompt Writing (Base)

本技能集的通用底座:只负责把意图写成 H3 可直接吃的提示词
风格题材(产品片、3D、纸艺等)见同仓库其他 skill;它们产出草稿后,可用本 skill 压成标准字段。
This is the universal base of the skill set: it only converts intentions into prompts that H3 can directly use.
For styles and themes (product videos, 3D, paper art, etc.), refer to other skills in the same repository; after they generate drafts, you can use this skill to compress them into standard fields.

何时用 / 不用

When to Use / Not to Use

用: 写或改 H3 提示词;图/视频/音频参考要标签化;跨镜连续;道具交接等精确物理。
不用: 与 H3 无关的文案、长剧剧本、后期工程说明。
Use: Writing or revising H3 prompts; tagging image/video/audio references; continuous cross-shot sequences; precise physical interactions like prop handovers.
Do Not Use: Copy unrelated to H3, long drama scripts, post-production project instructions.

输出规格

Output Specifications

时长4–15s
画幅21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16
分辨率默认短边 768,可 2K
帧率/音频24fps;立体声 32kHz
模式输入写法要点
T2VA纯文本完整视听时间线
I2VA1 图首帧从图起锚向前发展
FL2VA首+尾 2 图首→尾连续路径;默认单镜插值
L2VA1 图尾帧合理前态 → 收敛尾帧
Ref2VA图≤9 视频≤3 音频≤3 总≤12六段结构;标签全局同义

ItemValue
Duration4–15s
Aspect Ratio21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16
ResolutionDefault short side 768, optional 2K
Frame Rate/Audio24fps; Stereo 32kHz
ModeInputKey Writing Points
T2VAText-onlyComplete audio-visual timeline
I2VA1 first-frame imageAnchor from the image and develop forward
FL2VA2 images (first + last)Continuous path from first to last frame; default single-shot interpolation
L2VA1 last-frame imageReasonable previous state → converge to the last frame
Ref2VA≤9 images, ≤3 videos, ≤3 audios, total ≤12Six-section structure; globally synonymous tags

写法流程

Writing Process

  1. 判模式。
  2. references/base-zh.txt
    (基础)或
    ref-zh.txt
    (全参考);英文原版
    base-en.txt
    /
    ref-en.txt
  3. 描述太粗 →「八步扩展」;用户已给完整分镜 → 只查时长/时间戳/编号/矛盾,不擅自压缩。
  4. 强连续 / 多有序图 / 精确物理 → 先写中文分镜规划,再压终稿。
  5. 终稿:字段名、镜头标记、关系标记用英文固定写法;正文默认英文(更稳);对白/歌词/画面字保留原文。用户要中文正文时字段名不变。
  1. Determine the mode.
  2. Read
    references/base-zh.txt
    (basic) or
    ref-zh.txt
    (full reference); English versions are
    base-en.txt
    /
    ref-en.txt
    .
  3. If the description is too vague → use the "Eight-Step Expansion"; if the user has provided complete storyboards, only check duration/timestamps/numbering/contradictions, do not compress without permission.
  4. For strong continuity / multiple ordered images / precise physics → first write a Chinese storyboard plan, then compress into the final draft.
  5. Final Draft: Field names, shot markers, and relationship markers use fixed English expressions; the main text is in English by default (more stable); keep the original text for dialogue/lyrics/on-screen text. When the user requests Chinese main text, keep field names unchanged.

八步扩展

Eight-Step Expansion

输出目标 → 主体与参考资产 → 时间线 → 场景 → 镜头 → 视觉风格 → 声音 → 约束(必保/禁止)。
不臆造品牌、对白、不安全内容。

Output target → Subject and reference assets → Timeline → Scene → Shot → Visual style → Sound → Constraints (required/prohibited).
Do not invent brands, dialogue, or unsafe content.

基础模式终稿

Final Draft for Basic Modes

模式指令(首行;T2VA 无)

Mode Instruction (First line; none for T2VA)

I2VA
text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
FL2VA
N
=末镜号,
S.SS
=时长两位小数)
text
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.
L2VA
text
How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.
空一行后写三字段:
text
integrated_multimodal_description: [Shot 1] ...

overall_soundscape: ...

non_diegetic_music: ...
字段写什么
integrated_multimodal_description
时间线上的画面、动作、镜头、说话人、对白/唱、画内声
overall_soundscape
全片环境/物理/非语言人声;1–4 句;勿重复对白
non_diegetic_music
仅观众可听配乐(乐器、速度、动态);无则
N/A
I2VA
text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.
FL2VA (
N
=last shot number,
S.SS
=duration with two decimal places)
text
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.
L2VA
text
How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.
Leave a blank line and write three fields:
text
integrated_multimodal_description: [Shot 1] ...

overall_soundscape: ...

non_diegetic_music: ...
FieldContent to Write
integrated_multimodal_description
On-screen visuals, actions, shots, speakers, dialogue/singing, in-screen sounds on the timeline
overall_soundscape
Full-video environment/physical/non-verbal human sounds; 1–4 sentences; do not repeat dialogue
non_diegetic_music
Only background music audible to the audience (instruments, tempo, dynamics); write
N/A
if none

关键帧进描述

Keyframes in Description

  • I2VA
    <Picture 1>
    =0.00s 首帧 → 锚点 → 起势 → 发展 → 结果
  • FL2VA:首态 → 可观测中间变化 → 收窄 → 尾态;尾落在末镜片尾
  • L2VA:前态 → 路径 → 末镜对齐图
  • I2VA:
    <Picture 1>
    = 0.00s first frame → anchor → momentum → development → result
  • FL2VA: Initial state → observable intermediate changes → narrowing → final state; end at the last shot's end frame
  • L2VA: Previous state → path → align with the last shot's frame

镜头与运镜

Shots and Camera Movements

  • [Shot 1]
    无时间戳;后镜
    [Shot N] At 00:MM.SSS, ...
    ,切点递增且 ≤ 片长
  • 普切:
    the camera cuts to
    等;叠化仅用户明确要求
  • 运镜写进句内:类型 + 需要时的幅度/速度(Zoom/Push/Pan/Truck/Tilt/Pedestal/Arc/Tracking/Static/Shake/POV/Roll;
    with small|large amplitude
    at slow|fast speed
  • [Shot 1]
    has no timestamp; subsequent shots use
    [Shot N] At 00:MM.SSS, ...
    , with increasing cut points ≤ video length
  • Standard cuts:
    the camera cuts to
    etc.; cross-dissolve only if explicitly requested by the user
  • Write camera movements within the sentence: type + amplitude/speed if needed (Zoom/Push/Pan/Truck/Tilt/Pedestal/Arc/Tracking/Static/Shake/POV/Roll;
    with small|large amplitude
    ;
    at slow|fast speed
    )

说话人与字

Speakers and Text

  • (S1)
    (S2)
    跨镜稳定;合唱
    (S1,S2)
  • 对白只在
    <d>[Language] ...</d>
    ,原文不改写
  • 画外音:
    says in an off-screen voiceover
    + 在画闭嘴
  • 跨切
    <scenetrans>
    ;片尾截断
    <cutoff>
  • 画面字:英文双引号包原文
  • (S1)
    (S2)
    remain stable across shots; chorus uses
    (S1,S2)
  • Dialogue is only placed within
    <d>[Language] ...</d>
    , do not rewrite the original text
  • Voiceover:
    says in an off-screen voiceover
    + ensure the character's mouth is closed on-screen
  • Cross-cut:
    <scenetrans>
    ; end truncation:
    <cutoff>
  • On-screen text: Wrap original text in English double quotes

模式骨架

Mode Frameworks

模式骨架
T2VA风格+构图 → 动作与声 → 切镜补信息
I2VA首帧全锁 → 起势 → 发展 → 结果
FL2VA首态 → 中间变化 → 尾态
L2VA前态 → 路径 → 落尾帧
松散单段:一段紧凑即可。多有序图/强物理:用下方分镜规划。

ModeFramework
T2VAStyle + composition → actions and sounds → cut shots to supplement information
I2VALock first frame completely → momentum → development → result
FL2VAInitial state → intermediate changes → final state
L2VAPrevious state → path → end at last frame
Loose single segment: A compact single segment is sufficient. For multiple ordered images/strong physics: use the storyboard plan below.

全参考 Ref2VA

Full Reference Ref2VA

顺序固定:
text
subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music:
标签用途
<Subject N>
可复用可见内容(非源文件本身)
<Picture N>
图作具体帧/构图锚时单独建;仅定义角色则写进 Subject
<Video N>
剪辑源、续写起点、整片结构
<Audio N>
拷贝或参考的音频
summary
前缀:
keyframe completion
/
reference generation
/
video editing
/
video continuation
/
audio reuse
/
audio reference
,多类型用
+
画面保留标记:
fully_preserved
|
partially_preserved
|
attribute_transfer
|
weak_reference

音频:
fully_copy
|
partially_copy
|
reference
|
weak_reference
detailed_description
:风格 1–2 句 → 逐镜;生成任务约 350–500 英文词;标签首次出现写清,后镜只复用。
细则与完整例见
references/ref-zh.txt

Fixed order:
text
subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music:
TagUsage
<Subject N>
Reusable visible content (not the source file itself)
<Picture N>
Create separately when using images as specific frames/composition anchors; only write into Subject if defining roles
<Video N>
Editing source, continuation starting point, full-video structure
<Audio N>
Audio to copy or reference
summary
prefixes:
keyframe completion
/
reference generation
/
video editing
/
video continuation
/
audio reuse
/
audio reference
, use
+
for multiple types.
Visual retention markers:
fully_preserved
|
partially_preserved
|
attribute_transfer
|
weak_reference

Audio:
fully_copy
|
partially_copy
|
reference
|
weak_reference
detailed_description
: 1–2 sentences about style → shot-by-shot description; approximately 350–500 English words for generation tasks; clarify tags when first mentioned, reuse them in subsequent shots.
For detailed rules and complete examples, see
references/ref-zh.txt
.

中文分镜规划(再压终稿)

Chinese Storyboard Plan (Then Compress into Final Draft)

适用:≥2 有序图、跨镜连续、精确交接。
  • 4–15s 内物理可行;每镜一个主动作节拍
  • 时间戳连续,末点 =
    D
    ;默认 0.5s 精度
  • 两级时间轴:镜级范围 + 镜内微节拍(建立/预备/核心/稳定)
  • 镜 N 结束状态锁定 = 镜 N+1 起始(姿、视、双手、道具主/位/连接、机位侧)
  • 交接因果:归属 → 接触拿稳 → 松手 → 新归属;勿压成一句糊话
Applicable for: ≥2 ordered images, continuous cross-shot sequences, precise handovers.
  • Physically feasible within 4–15s; one main action beat per shot
  • Timestamps are continuous, end point =
    D
    ; default precision is 0.5s
  • Two-level Timeline: Shot-level range + micro-beats within the shot (setup/preparation/core/stabilization)
  • End state of Shot N = starting state of Shot N+1 (posture, gaze, hands, prop owner/location/connection, camera side)
  • Handover causality: Ownership → hold firmly upon contact → release → new ownership; do not compress into ambiguous sentences

模板

Template

text
【输出规格】{duration}s · {ratio} · {用途} · {N}张有序分镜图 · 不跳镜 · 锁人物服装位置道具

【整体风格】{媒介} · {视觉} · {光色纹理节奏运镜} · 场景{…} · 声音{…}

【人物与空间】人物1/2:外观与固定位置 · 硬约束{脸发型服装座位朝向…}

【道具】全片仅{名称数量} · 初态{主/手/位/连接} · 禁凭空增减跳位暗改归属

【两级时间轴】
镜头1|0:00-{T1}|{D1}s|图1
运镜:…  起始:…  微轴:建立→预备→核心→保持  结束锁定:…
镜头2|… 起始=镜头1结束锁定 …

【动作顺序】A→B→C  每镜一事  交接写清谁不动/谁拿稳/谁松手
【负面】禁身份漂移变装换位闪烁畸形多余肢物体跳变跳切字幕水印
交接微时序(耳机)见写作时按 0.5s 级拆「取出 → 停在两人之间 → 对方夹稳后才松手 → 未塞耳」。
规划后:基础模式写入
integrated_multimodal_description
;全参考写入
subject_definitions
+
detailed_description
+
retention_analysis

text
【Output Specifications】{duration}s · {ratio} · {purpose} · {N} ordered storyboard images · No shot skipping · Lock character clothing, positions, and props

【Overall Style】{medium} · {visuals} · {lighting, color, texture, rhythm, camera movement} · Scene{…} · Sound{…}

【Characters and Space】Character 1/2: Appearance and fixed position · Hard constraints{face, hairstyle, clothing, seat, orientation…}

【Props】Only {name, quantity} in the whole video · Initial state{owner/hand/location/connection} · 禁凭空增减跳位暗改归属

【Two-level Timeline】
Shot 1|0:00-{T1}|{D1}s|Image 1
Camera Movement:…  Start:…  Micro-axis: Setup→Preparation→Core→Hold  End Lock:…
Shot 2|… Start = Shot 1 End Lock …

【Action Sequence】A→B→C  One action per shot  Clearly write who stays still/who holds firmly/who releases during handover
【Negative Constraints】Forbid identity drift, costume changes, position shifts, flickering, deformities, extra limbs, object jumps, jump cuts, subtitles, watermarks
For micro-timing of handovers (e.g., headphones), split into 0.5s steps during writing: "Take out → stop between two people → release only after the other party holds firmly → not inserted into ear".
After planning: For basic modes, write into
integrated_multimodal_description
; for full references, write into
subject_definitions
+
detailed_description
+
retention_analysis
.

终稿自检

Final Draft Self-Check

  • 模式与指令行正确
  • 字段顺序/标签格式对
  • 时间连续,末点=时长
  • 无悬空标签;镜间状态可对账
  • 对白/可见字原文;未擅自加品牌对白
  • 声景与配乐字段分工正确
参考内容
base-zh.txt
/
ref-zh.txt
中文指南 + 英文范例
base-en.txt
/
ref-en.txt
官方英文原版
  • Correct mode and instruction line
  • Correct field order/tag format
  • Continuous timeline, end point = duration
  • No dangling tags; inter-shot states are reconcilable
  • Original dialogue/visible text preserved; no unauthorized brands or dialogue added
  • Correct division of labor between soundscape and music fields
ReferenceContent
base-zh.txt
/
ref-zh.txt
Chinese guide + English examples
base-en.txt
/
ref-en.txt
Official English original