h3-prompt-writing
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseH3 提示词写作(底座)
H3 Prompt Writing (Base)
本技能集的通用底座:只负责把意图写成 H3 可直接吃的提示词。
风格题材(产品片、3D、纸艺等)见同仓库其他 skill;它们产出草稿后,可用本 skill 压成标准字段。
风格题材(产品片、3D、纸艺等)见同仓库其他 skill;它们产出草稿后,可用本 skill 压成标准字段。
This is the universal base of the skill set: it only converts intentions into prompts that H3 can directly use.
For styles and themes (product videos, 3D, paper art, etc.), refer to other skills in the same repository; after they generate drafts, you can use this skill to compress them into standard fields.
For styles and themes (product videos, 3D, paper art, etc.), refer to other skills in the same repository; after they generate drafts, you can use this skill to compress them into standard fields.
何时用 / 不用
When to Use / Not to Use
用: 写或改 H3 提示词;图/视频/音频参考要标签化;跨镜连续;道具交接等精确物理。
不用: 与 H3 无关的文案、长剧剧本、后期工程说明。
不用: 与 H3 无关的文案、长剧剧本、后期工程说明。
Use: Writing or revising H3 prompts; tagging image/video/audio references; continuous cross-shot sequences; precise physical interactions like prop handovers.
Do Not Use: Copy unrelated to H3, long drama scripts, post-production project instructions.
Do Not Use: Copy unrelated to H3, long drama scripts, post-production project instructions.
输出规格
Output Specifications
| 项 | 值 |
|---|---|
| 时长 | 4–15s |
| 画幅 | 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16 |
| 分辨率 | 默认短边 768,可 2K |
| 帧率/音频 | 24fps;立体声 32kHz |
| 模式 | 输入 | 写法要点 |
|---|---|---|
| T2VA | 纯文本 | 完整视听时间线 |
| I2VA | 1 图首帧 | 从图起锚向前发展 |
| FL2VA | 首+尾 2 图 | 首→尾连续路径;默认单镜插值 |
| L2VA | 1 图尾帧 | 合理前态 → 收敛尾帧 |
| Ref2VA | 图≤9 视频≤3 音频≤3 总≤12 | 六段结构;标签全局同义 |
| Item | Value |
|---|---|
| Duration | 4–15s |
| Aspect Ratio | 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16 |
| Resolution | Default short side 768, optional 2K |
| Frame Rate/Audio | 24fps; Stereo 32kHz |
| Mode | Input | Key Writing Points |
|---|---|---|
| T2VA | Text-only | Complete audio-visual timeline |
| I2VA | 1 first-frame image | Anchor from the image and develop forward |
| FL2VA | 2 images (first + last) | Continuous path from first to last frame; default single-shot interpolation |
| L2VA | 1 last-frame image | Reasonable previous state → converge to the last frame |
| Ref2VA | ≤9 images, ≤3 videos, ≤3 audios, total ≤12 | Six-section structure; globally synonymous tags |
写法流程
Writing Process
- 判模式。
- 读 (基础)或
references/base-zh.txt(全参考);英文原版ref-zh.txt/base-en.txt。ref-en.txt - 描述太粗 →「八步扩展」;用户已给完整分镜 → 只查时长/时间戳/编号/矛盾,不擅自压缩。
- 强连续 / 多有序图 / 精确物理 → 先写中文分镜规划,再压终稿。
- 终稿:字段名、镜头标记、关系标记用英文固定写法;正文默认英文(更稳);对白/歌词/画面字保留原文。用户要中文正文时字段名不变。
- Determine the mode.
- Read (basic) or
references/base-zh.txt(full reference); English versions areref-zh.txt/base-en.txt.ref-en.txt - If the description is too vague → use the "Eight-Step Expansion"; if the user has provided complete storyboards, only check duration/timestamps/numbering/contradictions, do not compress without permission.
- For strong continuity / multiple ordered images / precise physics → first write a Chinese storyboard plan, then compress into the final draft.
- Final Draft: Field names, shot markers, and relationship markers use fixed English expressions; the main text is in English by default (more stable); keep the original text for dialogue/lyrics/on-screen text. When the user requests Chinese main text, keep field names unchanged.
八步扩展
Eight-Step Expansion
输出目标 → 主体与参考资产 → 时间线 → 场景 → 镜头 → 视觉风格 → 声音 → 约束(必保/禁止)。
不臆造品牌、对白、不安全内容。
不臆造品牌、对白、不安全内容。
Output target → Subject and reference assets → Timeline → Scene → Shot → Visual style → Sound → Constraints (required/prohibited).
Do not invent brands, dialogue, or unsafe content.
Do not invent brands, dialogue, or unsafe content.
基础模式终稿
Final Draft for Basic Modes
模式指令(首行;T2VA 无)
Mode Instruction (First line; none for T2VA)
I2VA
text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.FL2VA(=末镜号,=时长两位小数)
NS.SStext
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.L2VA
text
How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.空一行后写三字段:
text
integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...| 字段 | 写什么 |
|---|---|
| 时间线上的画面、动作、镜头、说话人、对白/唱、画内声 |
| 全片环境/物理/非语言人声;1–4 句;勿重复对白 |
| 仅观众可听配乐(乐器、速度、动态);无则 |
I2VA
text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.FL2VA (=last shot number, =duration with two decimal places)
NS.SStext
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.L2VA
text
How the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.Leave a blank line and write three fields:
text
integrated_multimodal_description: [Shot 1] ...
overall_soundscape: ...
non_diegetic_music: ...| Field | Content to Write |
|---|---|
| On-screen visuals, actions, shots, speakers, dialogue/singing, in-screen sounds on the timeline |
| Full-video environment/physical/non-verbal human sounds; 1–4 sentences; do not repeat dialogue |
| Only background music audible to the audience (instruments, tempo, dynamics); write |
关键帧进描述
Keyframes in Description
- I2VA:=0.00s 首帧 → 锚点 → 起势 → 发展 → 结果
<Picture 1> - FL2VA:首态 → 可观测中间变化 → 收窄 → 尾态;尾落在末镜片尾
- L2VA:前态 → 路径 → 末镜对齐图
- I2VA: = 0.00s first frame → anchor → momentum → development → result
<Picture 1> - FL2VA: Initial state → observable intermediate changes → narrowing → final state; end at the last shot's end frame
- L2VA: Previous state → path → align with the last shot's frame
镜头与运镜
Shots and Camera Movements
- 无时间戳;后镜
[Shot 1],切点递增且 ≤ 片长[Shot N] At 00:MM.SSS, ... - 普切:等;叠化仅用户明确要求
the camera cuts to - 运镜写进句内:类型 + 需要时的幅度/速度(Zoom/Push/Pan/Truck/Tilt/Pedestal/Arc/Tracking/Static/Shake/POV/Roll;;
with small|large amplitude)at slow|fast speed
- has no timestamp; subsequent shots use
[Shot 1], with increasing cut points ≤ video length[Shot N] At 00:MM.SSS, ... - Standard cuts: etc.; cross-dissolve only if explicitly requested by the user
the camera cuts to - Write camera movements within the sentence: type + amplitude/speed if needed (Zoom/Push/Pan/Truck/Tilt/Pedestal/Arc/Tracking/Static/Shake/POV/Roll; ;
with small|large amplitude)at slow|fast speed
说话人与字
Speakers and Text
(S1)跨镜稳定;合唱(S2)(S1,S2)- 对白只在 ,原文不改写
<d>[Language] ...</d> - 画外音:+ 在画闭嘴
says in an off-screen voiceover - 跨切 ;片尾截断
<scenetrans><cutoff> - 画面字:英文双引号包原文
(S1)remain stable across shots; chorus uses(S2)(S1,S2)- Dialogue is only placed within , do not rewrite the original text
<d>[Language] ...</d> - Voiceover: + ensure the character's mouth is closed on-screen
says in an off-screen voiceover - Cross-cut: ; end truncation:
<scenetrans><cutoff> - On-screen text: Wrap original text in English double quotes
模式骨架
Mode Frameworks
| 模式 | 骨架 |
|---|---|
| T2VA | 风格+构图 → 动作与声 → 切镜补信息 |
| I2VA | 首帧全锁 → 起势 → 发展 → 结果 |
| FL2VA | 首态 → 中间变化 → 尾态 |
| L2VA | 前态 → 路径 → 落尾帧 |
松散单段:一段紧凑即可。多有序图/强物理:用下方分镜规划。
| Mode | Framework |
|---|---|
| T2VA | Style + composition → actions and sounds → cut shots to supplement information |
| I2VA | Lock first frame completely → momentum → development → result |
| FL2VA | Initial state → intermediate changes → final state |
| L2VA | Previous state → path → end at last frame |
Loose single segment: A compact single segment is sufficient. For multiple ordered images/strong physics: use the storyboard plan below.
全参考 Ref2VA
Full Reference Ref2VA
顺序固定:
text
subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music:| 标签 | 用途 |
|---|---|
| 可复用可见内容(非源文件本身) |
| 图作具体帧/构图锚时单独建;仅定义角色则写进 Subject |
| 剪辑源、续写起点、整片结构 |
| 拷贝或参考的音频 |
summarykeyframe completionreference generationvideo editingvideo continuationaudio reuseaudio reference+画面保留标记: | | |
音频: | | |
fully_preservedpartially_preservedattribute_transferweak_reference音频:
fully_copypartially_copyreferenceweak_referencedetailed_description细则与完整例见
references/ref-zh.txtFixed order:
text
subject_definitions:
summary:
retention_analysis:
detailed_description:
overall_soundscape:
non_diegetic_music:| Tag | Usage |
|---|---|
| Reusable visible content (not the source file itself) |
| Create separately when using images as specific frames/composition anchors; only write into Subject if defining roles |
| Editing source, continuation starting point, full-video structure |
| Audio to copy or reference |
summarykeyframe completionreference generationvideo editingvideo continuationaudio reuseaudio reference+Visual retention markers: | | |
Audio: | | |
fully_preservedpartially_preservedattribute_transferweak_referenceAudio:
fully_copypartially_copyreferenceweak_referencedetailed_descriptionFor detailed rules and complete examples, see
references/ref-zh.txt中文分镜规划(再压终稿)
Chinese Storyboard Plan (Then Compress into Final Draft)
适用:≥2 有序图、跨镜连续、精确交接。
- 4–15s 内物理可行;每镜一个主动作节拍
- 时间戳连续,末点 = ;默认 0.5s 精度
D - 两级时间轴:镜级范围 + 镜内微节拍(建立/预备/核心/稳定)
- 镜 N 结束状态锁定 = 镜 N+1 起始(姿、视、双手、道具主/位/连接、机位侧)
- 交接因果:归属 → 接触拿稳 → 松手 → 新归属;勿压成一句糊话
Applicable for: ≥2 ordered images, continuous cross-shot sequences, precise handovers.
- Physically feasible within 4–15s; one main action beat per shot
- Timestamps are continuous, end point = ; default precision is 0.5s
D - Two-level Timeline: Shot-level range + micro-beats within the shot (setup/preparation/core/stabilization)
- End state of Shot N = starting state of Shot N+1 (posture, gaze, hands, prop owner/location/connection, camera side)
- Handover causality: Ownership → hold firmly upon contact → release → new ownership; do not compress into ambiguous sentences
模板
Template
text
【输出规格】{duration}s · {ratio} · {用途} · {N}张有序分镜图 · 不跳镜 · 锁人物服装位置道具
【整体风格】{媒介} · {视觉} · {光色纹理节奏运镜} · 场景{…} · 声音{…}
【人物与空间】人物1/2:外观与固定位置 · 硬约束{脸发型服装座位朝向…}
【道具】全片仅{名称数量} · 初态{主/手/位/连接} · 禁凭空增减跳位暗改归属
【两级时间轴】
镜头1|0:00-{T1}|{D1}s|图1
运镜:… 起始:… 微轴:建立→预备→核心→保持 结束锁定:…
镜头2|… 起始=镜头1结束锁定 …
【动作顺序】A→B→C 每镜一事 交接写清谁不动/谁拿稳/谁松手
【负面】禁身份漂移变装换位闪烁畸形多余肢物体跳变跳切字幕水印交接微时序(耳机)见写作时按 0.5s 级拆「取出 → 停在两人之间 → 对方夹稳后才松手 → 未塞耳」。
规划后:基础模式写入 ;全参考写入 + + 。
integrated_multimodal_descriptionsubject_definitionsdetailed_descriptionretention_analysistext
【Output Specifications】{duration}s · {ratio} · {purpose} · {N} ordered storyboard images · No shot skipping · Lock character clothing, positions, and props
【Overall Style】{medium} · {visuals} · {lighting, color, texture, rhythm, camera movement} · Scene{…} · Sound{…}
【Characters and Space】Character 1/2: Appearance and fixed position · Hard constraints{face, hairstyle, clothing, seat, orientation…}
【Props】Only {name, quantity} in the whole video · Initial state{owner/hand/location/connection} · 禁凭空增减跳位暗改归属
【Two-level Timeline】
Shot 1|0:00-{T1}|{D1}s|Image 1
Camera Movement:… Start:… Micro-axis: Setup→Preparation→Core→Hold End Lock:…
Shot 2|… Start = Shot 1 End Lock …
【Action Sequence】A→B→C One action per shot Clearly write who stays still/who holds firmly/who releases during handover
【Negative Constraints】Forbid identity drift, costume changes, position shifts, flickering, deformities, extra limbs, object jumps, jump cuts, subtitles, watermarksFor micro-timing of handovers (e.g., headphones), split into 0.5s steps during writing: "Take out → stop between two people → release only after the other party holds firmly → not inserted into ear".
After planning: For basic modes, write into ; for full references, write into + + .
integrated_multimodal_descriptionsubject_definitionsdetailed_descriptionretention_analysis终稿自检
Final Draft Self-Check
- 模式与指令行正确
- 字段顺序/标签格式对
- 时间连续,末点=时长
- 无悬空标签;镜间状态可对账
- 对白/可见字原文;未擅自加品牌对白
- 声景与配乐字段分工正确
| 参考 | 内容 |
|---|---|
| 中文指南 + 英文范例 |
| 官方英文原版 |
- Correct mode and instruction line
- Correct field order/tag format
- Continuous timeline, end point = duration
- No dangling tags; inter-shot states are reconcilable
- Original dialogue/visible text preserved; no unauthorized brands or dialogue added
- Correct division of labor between soundscape and music fields
| Reference | Content |
|---|---|
| Chinese guide + English examples |
| Official English original |