dramake

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Dramake

Dramake

把短剧制作当成一条连续生产线,而不是一堆提示词。让没有编剧、导演或剪辑经验的用户,也能从一段文字开始,在明确的验收门禁下完成故事、剧本、分镜、生成、配音、剪辑和质检;让专业用户可以从任意阶段接管。
本 Skill 不绑定单一平台、网站、API 或 MCP,执行平台由用户选择。使用用户能合法访问并愿意使用的模型与执行入口;每次执行前核实当前参数、价格、素材上传与保存规则。
project-state.md 作为唯一生产状态机。提示词、接口参数和时间线都从已批准的创意状态与已检查的真实素材派生。
Treat short drama production as a continuous production line, not a bunch of prompts. Allow users without screenwriting, directing, or editing experience to start from a piece of text and complete story creation, script writing, storyboarding, generation, dubbing, editing, and quality inspection under clear acceptance gates; enable professional users to take over at any stage.
This Skill is not bound to a single platform, website, API, or MCP; the execution platform is chosen by the user. Use models and execution entrances that the user can legally access and is willing to use; verify current parameters, prices, material upload and storage rules before each execution.
Use project-state.md as the sole production state machine. Prompts, interface parameters, and timelines are derived from approved creative states and inspected real materials.

十条铁律

Ten Iron Rules

  1. 主角必须因果性改变结果,人、动物、产品或物体都不能只是装饰。
  2. 静音观看也应能理解主要因果与重点笑点;对白、字幕和群演笑声不能替代可见铺垫、原因、反应与揭晓。
  3. 模型支持时优先使用更少、更长的剧情段,但必须限制事件负载:段内可有多个节拍,不能让一个提示词同时承担多条高风险物理链、道具变种和无关任务。
  4. 未实际下载和验收上游素材,不得生成依赖其状态的下游段落。
  5. 锁定身份、材质、服装结构、轴线、道具、天气、湿度、光线和声音边界。
  6. 接触必须有接近、接触、施力、响应和稳定结果;拒绝穿模与不可能解剖。
  7. 默认不在视频里生成手机 UI、信件、价签、倒计时和关键文字;后期叠加。
  8. 原生对白可以解决口型,但每句必须锁定唯一说话者、可见嘴型与非说话者闭嘴状态;原生音乐不能自动解决终片混音。
  9. 技术干净不等于剧情能发;60–120 秒故事必须有贯穿全片的宏观因果链和有铺垫的主反转,并进行未读梗概的冷观众检查。
  10. 没有真实看过、听过或抽帧检查,就不得声称素材已通过。
  1. The protagonist must causally change the outcome; humans, animals, products, or objects cannot be mere decorations.
  2. The main causality and key punchlines should be understandable even when watching without sound; dialogue, subtitles, and audience laughter cannot replace visible foreshadowing, causes, reactions, and reveals.
  3. When supported by the model, prioritize fewer, longer plot segments, but must limit event load: multiple beats can be included in a segment, but a single prompt cannot simultaneously handle multiple high-risk physical chains, prop variants, and unrelated tasks.
  4. Do not generate downstream segments that depend on the state of upstream materials without actually downloading and accepting those materials.
  5. Lock identity, material, clothing structure, axis, props, weather, humidity, lighting, and sound boundaries.
  6. Contact must include approach, contact, force application, response, and stable result; reject mesh penetration and impossible anatomy.
  7. By default, do not generate mobile UI, letters, price tags, countdowns, and key text in the video; add them in post-production.
  8. Native dialogue can solve lip-sync issues, but each line must lock a unique speaker, visible lip movements, and the non-speaker's closed-mouth state; native music cannot automatically solve final film mixing.
  9. Technical cleanliness does not mean the plot is publishable; a 60–120 second story must have a macro causal chain running through the entire film and a foreshadowed main reversal, and undergo a cold audience check without reading the outline.
  10. Do not claim materials have passed inspection without actually watching, listening to, or frame-by-frame checking them.

从用户会什么开始

Start from What the User Knows

先判断用户输入深度,不要要求小白先学会专业术语:
  • **一句话/脑洞:**主动给出 2–3 个可拍方向,补齐主角、欲望、阻碍、升级与结果;
  • **故事/小说/文案:**抽取人物、关系、时间线、道具、秘密、转折与可视化动作;
  • **剧本:**检查因果、铺垫、反转、对白密度、场景数量和时长;
  • **分镜/提示词:**检查节拍、连续性、物理、文字风险和模型可执行性;
  • **已有素材/粗剪:**实际检查画面与声音,定位最小失败单元,不无故推倒重来。
再选任务深度:
  • **诊断/建议:**只分析,不生成、不付费;
  • **快速修复:**只修剧本、提示词、镜头、接缝、声线或混音的最小单元;
  • **生产包:**输出简报、剧本、母版、分镜、清单、提示词、费用与剪辑方案;
  • **执行/续作:**从当前已验收状态继续;
  • **完整制作:**从文字到发布执行整条状态机。
局部问题不得被强制重走整条流程。
First judge the depth of the user's input, do not require beginners to learn professional terminology first:
  • One-sentence/idea: Proactively provide 2–3 shootable directions, supplementing the protagonist, desire, obstacle, upgrade, and outcome;
  • Story/novel/copywriting: Extract characters, relationships, timelines, props, secrets, twists, and visualizable actions;
  • Script: Check causality, foreshadowing, reversal, dialogue density, number of scenes, and duration;
  • Storyboard/prompt: Check beats, continuity, physics, text risks, and model executability;
  • Existing materials/rough cut: Actually check the footage and sound, locate the smallest failed unit, and do not restart production without reason.
Then choose the task depth:
  • Diagnosis/suggestion: Only analyze, do not generate or pay;
  • Quick fix: Only repair the smallest units of scripts, prompts, shots, seams, voices, or mixing;
  • Production package: Output brief, script, master version, storyboard, checklist, prompts, cost, and editing plan;
  • Execution/sequel: Continue from the current accepted state;
  • Full production: Execute the entire state machine from text to release.
Local issues cannot be forced to go through the entire process again.

选择内容形态与画幅

Choose Content Form and Aspect Ratio

先选择当前交付的主形态,详见 format-routing.md
  • **15–45 秒独立小短片:**一个钩子、一次升级、一个回报;
  • **45–120 秒独立短片:**2–4 个剧情段,完成情绪或喜剧弧;
  • **1–4 分钟单元剧:**固定角色/世界规则,每集问题独立解决;
  • **1–4 分钟连续短剧:**单集目标与回报之外,还维护季线状态。
支持
9:16
16:9
1:1
和自定义比例。记录发布渠道、画幅、主体安全区、字幕区、封面裁切区与交付编码。
需要横竖双版本时,共用剧情、人物、声线、道具与连续性母版,分别设计构图、景别、移动路线和安全区,分别验收。不得默认中心裁切或机械补边。
First select the main form of the current delivery, see format-routing.md for details:
  • 15–45 second independent short clips: One hook, one upgrade, one reward;
  • 45–120 second independent short films: 2–4 plot segments, completing an emotional or comedic arc;
  • 1–4 minute anthology series: Fixed characters/world rules, each episode's problem is solved independently;
  • 1–4 minute serial short dramas: In addition to the single-episode goal and reward, maintain the season-line state.
Supports
9:16
,
16:9
,
1:1
, and custom ratios. Record the release channel, aspect ratio, subject safe area, subtitle area, cover cropping area, and delivery encoding.
When both horizontal and vertical versions are needed, share the plot, characters, voices, props, and continuity master version, design composition, shot size, movement path, and safe area separately, and accept separately. Do not default to center cropping or mechanical edge filling.

建立生产简报

Establish Production Brief

使用 production-brief.md。从对话和素材推断安全默认值,只询问会实质改变结果的选择。
至少记录:
  • 原始文字与改编边界;
  • 渠道、画幅、时长、受众、语言和情绪承诺;
  • 人物、动物、产品、地点、参考素材与使用权;
  • 独立/单元/连续形式与内容配置;
  • 可用模型/执行入口和大致预算上限;
  • 对白策略:无对白、原生、参考音频、外部配音或混合;
  • 视觉介质:真人、写实宠物、动画、漫剧、UGC 或混合。
公开示例不得暴露密钥、签名链接、请求 ID、私人路径、真实脸/声音或私有参考素材。
Use production-brief.md. Infer safe default values from conversations and materials, only ask for choices that will substantially change the result.
Record at least:
  • Original text and adaptation boundaries;
  • Channel, aspect ratio, duration, audience, language, and emotional commitment;
  • Characters, animals, products, locations, reference materials, and usage rights;
  • Independent/anthology/serial form and content configuration;
  • Available models/execution entrances and approximate budget ceiling;
  • Dialogue strategy: no dialogue, native, reference audio, external dubbing, or hybrid;
  • Visual medium: live-action, realistic pets, animation, comic dramas, UGC, or hybrid.
Public examples must not expose keys, signed links, request IDs, private paths, real faces/voices, or private reference materials.

先写故事,再写提示词

Write the Story First, Then the Prompt

阅读 story-and-pacing.mddirecting-camera-and-pacing.md;喜剧、萌宠、拟人或动画化表演项目同时阅读 comedy-and-animated-performance.md。交付:
  1. 一句话观众承诺;
  2. 0–3 秒视觉钩子;
  3. 只属于该主角的动机;
  4. 可见目标与阻碍;
  5. 每 4–8 秒一次信息、行动、反应、关系或风险变化;
  6. 关系或状态变化;
  7. 有铺垫的回报/反转;
  8. 最后一记情绪、笑点、悬念或角色性动作。
60–120 秒喜剧先写
触发事件 → 主角误解/欲望 → 目标 → 主动选择 → 代价升级 → 主反转 → 后果/回收
,再放局部笑点。触发优先由角色在生活流中亲历,而不是开场站定听完整段说明;每个笑点必须改变主线状态,并按
铺垫 → 可见原因 → 观众信息优势 → 反应停顿 → 揭晓 → 后果
验收。
每个剧情段定义
entry_state
、可观察节拍和
exit_state
。20–30 秒长段通常写 4–6 个节拍;10–12 秒对白段通常写 2–3 个节拍。持续加入呼吸、眨眼、视线、表情、物种合适的反应、道具互动与背景生活,不能让角色说完一句话后静止。
目标为 60–120 秒抖音高密度剧情时,启用对应预设:3 秒内出现明确冲突,10 秒内讲清即时目标,90 秒以 18–24 个有效节拍为起点,每 15–25 秒闭合一个微任务;约 30 秒连续配音场景包含多轮推进对话和 4–7 个可见动作/反应。一个固定宏观目标贯穿全片,换场必须有因果桥。有效节拍不等于切镜或生成请求,不得重新拆成碎片硬拼。
出现以下任一情况,生成前退回修改:
  • 删除主角仍不改变剧情;
  • 动机只靠台词解释;
  • 反转没有提前种下可见线索;
  • 人物反复表达同一个意思;
  • 地点、时间、目标或人物知识突然改变;
  • 冷观众无法说出目标、阻碍、选择和结果。
Read story-and-pacing.md and directing-camera-and-pacing.md; for comedy, cute pet, anthropomorphic, or animated performance projects, also read comedy-and-animated-performance.md. Deliver:
  1. One-sentence audience commitment;
  2. 0–3 second visual hook;
  3. Motivation unique to the protagonist;
  4. Visible goal and obstacle;
  5. A change in information, action, reaction, relationship, or risk every 4–8 seconds;
  6. Relationship or state change;
  7. Foreshadowed reward/reversal;
  8. Final emotional beat, punchline, suspense, or character-specific action.
For 60–120 second comedies, first write
trigger event → protagonist's misunderstanding/desire → goal → active choice → escalating cost → main reversal → consequence/payoff
, then add local punchlines. Triggers should be experienced by the character in daily life, rather than standing at the opening listening to a full explanation; each punchline must change the main line state, and be accepted according to
foreshadowing → visible cause → audience information advantage → reaction pause → reveal → consequence
.
Define
entry_state
, observable beats, and
exit_state
for each plot segment. A 20–30 second long segment usually has 4–6 beats; a 10–12 second dialogue segment usually has 2–3 beats. Continuously add breathing, blinking, eye contact, expressions, species-appropriate reactions, prop interactions, and background life; do not let characters stand still after finishing a line.
When targeting high-density 60–120 second Douyin plots, enable the corresponding preset: clear conflict within 3 seconds, clear immediate goal within 10 seconds, start with 18–24 effective beats for 90 seconds, close a micro-task every 15–25 seconds; a continuous dubbing scene of about 30 seconds includes multiple rounds of advancing dialogue and 4–7 visible actions/reactions. A fixed macro goal runs through the entire film, and scene transitions must have a causal bridge. Effective beats do not equal cuts or generation requests; do not re-split into fragmented stitching.
Return for modification before generation if any of the following occurs:
  • The plot does not change when the protagonist is removed;
  • Motivation is only explained by lines;
  • Reversal has no visible clues planted in advance;
  • Characters repeatedly express the same meaning;
  • Location, time, goal, or character knowledge changes suddenly;
  • Cold audience cannot state the goal, obstacle, choice, and outcome.

锁定内容配置、角色、风格和世界

Lock Content Configuration, Characters, Style, and World

阅读 content-profiles.md,选择一个主配置,必要时叠加专项规则。
为每个常驻人物/动物创建母版与 character-bible.md,记录身份、解剖、脸/身体比例、标记、服装结构、允许行为、禁止变异和声音。使用 identity-voice-and-reference-locks.md 为每个控制维度指定唯一权威参考,并为每个常驻说话人建立独立声线锚点。
在分镜前盘点所有重复、跨场景、剧情关键或带文字/Logo 的道具,为每件建立 prop-bible.md。先用用户选择并验证的高保真图像生成/编辑路线(可包括 GPT Image 类模型)固定中性光多视角几何、尺度、材质、连接件与状态;精确文字另建
text_plate_id
,默认后期跟踪合成。每个镜头引用同一
prop_id + state_id + text_plate_id
,不得让秤、印章、头盔、手机或包装跨幕重新设计。
真人项目锁定真实摄影、皮肤/毛发、实际服装材质、真实尺度、动机光和物理融合特效。动画项目锁定设计语法、线条、阴影、比例与运动语言。真人环境中的宠物可以选择自然写实、加强写实或“动画演员混合”:环境与材质写实,角色使用锁定的拟人姿势、表情库、蓄力、过冲、回稳和受控 squash/stretch;道具与接触仍服从真实连续性。不得把“动画夸张”当作穿模、瞬移或道具漂移的借口。阅读 style-and-medium-locks.mdcomedy-and-animated-performance.md
使用 continuity-ledger.csv 维护正式状态,每个验收段都要更新。
长小说、连续剧或跨集 IP 同时阅读 series-memory-and-source-adaptation.md,使用分层事实、人物知识、时间线与揭示状态,而不是把全文反复塞入提示词或依赖对话记忆。
多角色调度、真实地点、追逐、镜面、巨物、设施、复杂出入口或异常空间同时阅读 spatial-blocking-and-world-state.md。先锁地点锚点与连通关系,再写角色
start → path → end
、机位、视线、遮挡、尺度和反射拓扑;提示词不得临时发明位置。
在高风险或强连续项目中,视频生成前阅读 previsualization-and-coverage.md:先做角色/地点/道具长期母版,再为关键剧情时刻试拍少量覆盖方案,最后只冻结真正会约束请求的首帧、尾帧或状态帧。简单反应镜头不强制做完整故事板;多角色、巨物尺度、镜面规则、真实地标、精确走位和关键接触不得跳过预演。
需要生成角色设定图、场景图、道具图或分镜图时,同时阅读 reference-graph-and-storyboard-images.md。把素材组织成有职责的依赖图,先验收一个构图母帧,再生成后续分镜;每个参考明确
controls
must_not_transfer
,多姿势拼版不得复制成额外角色。
Read content-profiles.md, select a main configuration, and overlay special rules if necessary.
Create a master version and character-bible.md for each permanent character/animal, recording identity, anatomy, face/body proportions, markings, clothing structure, allowed behaviors, prohibited variations, and voice. Use identity-voice-and-reference-locks.md to assign a unique authoritative reference for each control dimension, and establish an independent voice anchor for each permanent speaker.
Before storyboarding, inventory all repeated, cross-scene, plot-critical, or text/Logo-bearing props, and create prop-bible.md for each. First use the user-selected and verified high-fidelity image generation/editing route (which may include GPT Image-like models) to fix neutral-light multi-angle geometry, scale, material, connectors, and state; precise text is separately assigned a
text_plate_id
, which is tracked and synthesized in post-production by default. Each shot references the same
prop_id + state_id + text_plate_id
; do not let scales, seals, helmets, mobile phones, or packaging be redesigned across acts.
Live-action projects lock real photography, skin/hair, actual clothing materials, real scale, motivated lighting, and physically integrated special effects. Animation projects lock design grammar, lines, shadows, proportions, and movement language. Pets in live-action environments can choose natural realism, enhanced realism, or "animated actor hybrid": the environment and materials are realistic, while the character uses locked anthropomorphic poses, expression libraries, anticipation, overshoot, settle, and controlled squash/stretch; props and contact still follow real continuity. Do not use "animation exaggeration" as an excuse for mesh penetration, teleportation, or prop drift. Read style-and-medium-locks.md and comedy-and-animated-performance.md.
Use continuity-ledger.csv to maintain formal status, updating it for each accepted segment.
For long novels, serial dramas, or cross-episode IPs, also read series-memory-and-source-adaptation.md, using layered facts, character knowledge, timelines, and reveal states instead of repeatedly stuffing the full text into prompts or relying on dialogue memory.
For multi-character scheduling, real locations, chases, mirrors, giants, facilities, complex entrances/exits, or abnormal spaces, also read spatial-blocking-and-world-state.md. First lock location anchors and connectivity, then write the character's
start → path → end
, camera position, line of sight, occlusion, scale, and reflection topology; prompts must not invent locations temporarily.
In high-risk or strongly continuous projects, read previsualization-and-coverage.md before video generation: first create long-term master versions of characters/locations/props, then shoot a small number of coverage options for key plot moments, and finally only freeze the first frame, last frame, or state frame that truly constrains the request. Simple reaction shots do not require a complete storyboard; multi-character, giant scale, mirror rules, real landmarks, precise movement, and key contact cannot skip previsualization.
When needing to generate character setting images, scene images, prop images, or storyboard images, also read reference-graph-and-storyboard-images.md. Organize materials into a responsible dependency graph, first accept a composition master frame, then generate subsequent storyboards; each reference clearly defines
controls
and
must_not_transfer
, and multi-pose collages must not be copied into additional characters.

设计生成段与剪辑节拍

Design Generation Segments and Editing Beats

阅读 generation-blocks-and-edit-beats.mddirecting-camera-and-pacing.md。模型稳定支持时使用 15–30 秒段;精确对白、复杂接触、高运动、变形或易漂移任务使用更短段。
每个长段必须有:
  • 3–6 个带时间的节拍;
  • 每个节拍一个主要动作或信息变化;
  • 需要时有 2–5 次有动机的内部切镜/视角变化;
  • 稳定的时间、地点和轴线范围;
  • 重复道具的
    prop_id/state_id/text_plate_id
    与唯一允许变化;
  • 精确物理事件的起点、控制输入、世界/屏幕向量、阶段序列和终态;
  • macro_goal_id
    、本段对主线的唯一推进,以及事件负载预算;
  • 喜剧段的笑点合同、视线/遮挡逻辑、反应角色和 0.6–1.0 秒结果停顿;
  • 钩子/碰撞/揭示/高潮需要时的单一
    impact_role
    与可见后果;
  • 末尾 0.5–0.8 秒仍有动作的剪辑余量;
  • 下一段能承接的明确退出动作。
每个外部接缝:
  • 支持时使用已验收尾帧,但同时提供身份/材质母版;
  • 保持屏幕方向与动作阶段;
  • 必要时重叠约 0.3–2 秒动作;
  • 用动作、遮挡、甩镜、前景经过、声音、视线或形状匹配过渡;
  • 不得同时无解释改变地点、动作、景别、方向、服装和天气。
详见 continuity-and-transitions.md
长段不是“把更多事情塞进一次请求”。每段优先一个主道具系统、一个高风险物理链和一个主要喜剧/情绪回报;若需要多个地点、食物种类、精确交接、多人轮流说话或多次道具状态变化,拆段或改用内部建立/峰值/后果镜头。
正式提示词必须从已批准状态编译,阅读 prompt-compiler.md。使用稳定的
场景上下文 → 参考职责 → 空间与首状态 → 镜头 → 时间节拍 → 表演/物理 → 灯光/声音 → MUST-HIT/终态 → 高损失负面约束
骨架,再按 T2V、首帧、首尾帧、R2V、续写、编辑或多镜头模式裁剪。付费前展示完整请求预览;任何影响画面或费用的 prompt、参数、参考或 schema 变化都会使旧确认失效。
Read generation-blocks-and-edit-beats.md and directing-camera-and-pacing.md. Use 15–30 second segments when the model stably supports them; use shorter segments for precise dialogue, complex contact, high movement, deformation, or drift-prone tasks.
Each long segment must have:
  • 3–6 timed beats;
  • One main action or information change per beat;
  • 2–5 motivated internal cuts/perspective changes when needed;
  • Stable time, location, and axis range;
  • prop_id/state_id/text_plate_id
    for repeated props and only allowed changes;
  • Start point, control input, world/screen vectors, phase sequence, and final state for precise physical events;
  • macro_goal_id
    , unique advancement of the main line by this segment, and event load budget;
  • Punchline contract, line-of-sight/occlusion logic, reaction characters, and 0.6–1.0 second result pause for comedy segments;
  • Single
    impact_role
    and visible consequences when hooks/collisions/reveals/climaxes are needed;
  • 0.5–0.8 second editing margin with ongoing action at the end;
  • Clear exit action that can be承接 by the next segment.
Each external seam:
  • Use accepted last frames when supported, but also provide identity/material master versions;
  • Maintain screen direction and action phase;
  • Overlap actions by approximately 0.3–2 seconds if necessary;
  • Transition using action, occlusion, whip pan, foreground pass, sound, line of sight, or shape matching;
  • Do not change location, action, shot size, direction, clothing, and weather simultaneously without explanation.
See continuity-and-transitions.md for details.
Long segments are not "stuffing more things into one request". Prioritize one main prop system, one high-risk physical chain, and one main comedy/emotional reward per segment; if multiple locations, food types, precise handovers, multiple people speaking in turn, or multiple prop state changes are needed, split the segment or use internal setup/peak/consequence shots.
Formal prompts must be compiled from approved states, read prompt-compiler.md. Use a stable skeleton of
scene context → reference responsibilities → space and initial state → shot → timed beats → performance/physics → lighting/sound → MUST-HIT/final state → high-loss negative constraints
, then crop according to T2V, first frame, first-last frame, R2V, continuation, editing, or multi-shot mode. Show a complete request preview before payment; any change to prompts, parameters, references, or schemas that affects the footage or cost invalidates the old confirmation.

按能力选模型

Choose Models by Capability

阅读 model-router.md,提交前核实当前入口。
  • **Wan 3 类长段:**优先测试 20–30 秒多节拍段、参考驱动多角色原生对白、减少外部接缝和 720p 经济/均衡任务;混合 R2V 与严格首尾帧二选一,重点查角色/声线绑定和后半段漂移。
  • **Seedance 2.5 类:**高价值对白、多图/视频/音频参考、参考声线、编辑、续写和修复。
  • **Seedance 2.0:**4–15 秒对白、动作和过渡。
  • **Seedance 2.0 Fast:**预演、覆盖、备选与经济档;严格验收后才可进入终片。
  • **MiniMax H3 类:**短时高运动、首尾帧控制和重点音画镜头。
  • **Kling 3 Omni 类:**元素绑定、多镜头内部编排和首尾关键帧桥接;先验收子镜头负载与内部切镜。
  • **高保真图像模型:**角色/地点/道具母版、覆盖方案和冻结关键帧;静态图通过后再进入运动生成。
这些只是起始假设,不是永久排名。实际接口暴露的能力高于品牌印象。
Read model-router.md, verify the current entrance before submission.
  • Wan 3-like long segments: Prioritize testing 20–30 second multi-beat segments, reference-driven multi-character native dialogue, reducing external seams, and 720p economy/balanced tasks; mix R2V and strict first-last frame selection, focus on checking character/voice binding and drift in the second half.
  • Seedance 2.5-like: High-value dialogue, multi-image/video/audio references, reference voices, editing, continuation, and repair.
  • Seedance 2.0: 4–15 second dialogue, action, and transitions.
  • Seedance 2.0 Fast: Previsualization, coverage, alternatives, and economy tier; can only enter the final film after strict acceptance.
  • MiniMax H3-like: Short-duration high movement, first-last frame control, and key audio-visual shots.
  • Kling 3 Omni-like: Element binding, multi-shot internal orchestration, and first-last keyframe bridging; first accept sub-shot load and internal cuts.
  • High-fidelity image models: Character/location/prop master versions, coverage options, and frozen keyframes; enter motion generation only after static images are accepted.
These are only starting assumptions, not permanent rankings. The actual capabilities exposed by the interface are higher than brand impressions.

选择质量与付费授权

Choose Quality and Paid Authorization

阅读 budget-and-quality.md,付费前建立费用计划。价格永远是项目配置,不写成永久事实。
  • **Sketch:**最低可用预演;
  • **Economy:**通常 720p、长段优先、定点修复;
  • **Balanced:**把强参考/音频模型用在真正提高通过率的段;
  • **Premium:**富多模态、参考音频、编辑/续写和选择性高分辨率;
  • **Hero Insert:**集中预算做钩子、高潮、复杂接触或情绪特写;
  • **Final Polish:**只放大和修复已验收素材。
授权另选 authorization-and-batch.md 中的
plan-only
single-test
approved-batch
autopilot-with-cap
。质量档不代表付费许可。用户已预先授权平台、上传规则、预算与重试上限时,不要逐条重复询问;触及边界必须停止。
正式执行同时阅读 production-operations-and-recovery.md,用 generation-jobs.csv 保存请求指纹、幂等键、任务状态、依赖、费用、恢复和候选谱系。超时先进入
unknown
并找回任务,不能把未知状态当失败后盲目重提;人工修改上游状态后,使受影响的确认和下游任务失效。
Read budget-and-quality.md, establish a cost plan before payment. Prices are always project configurations, not permanent facts.
  • Sketch: Minimum usable previsualization;
  • Economy: Usually 720p, long segment priority, targeted repair;
  • Balanced: Use strong reference/audio models on segments that truly improve pass rate;
  • Premium: Rich multimodal, reference audio, editing/continuation, and selective high resolution;
  • Hero Insert: Concentrate budget on hooks, climaxes, complex contact, or emotional close-ups;
  • Final Polish: Only upscale and repair accepted materials.
Authorization options include
plan-only
,
single-test
,
approved-batch
, or
autopilot-with-cap
from authorization-and-batch.md. Quality tiers do not represent paid licenses. Do repeatedly ask for confirmation item by item if the user has pre-authorized the platform, upload rules, budget, and retry limits; must stop when touching boundaries.
Read production-operations-and-recovery.md during formal execution, use generation-jobs.csv to save request fingerprints, idempotency keys, task status, dependencies, costs, recovery, and candidate lineage. Enter
unknown
state first for timeouts and retrieve the task; do not blindly resubmit by treating unknown status as failure; invalidate affected confirmations and downstream tasks after manually modifying upstream states.

选择对白与声音路线

Choose Dialogue and Audio Route

阅读 audio-and-dubbing.md,按场景选择,而不是整片盲目统一。
  • **无对白:**统一环境底、音乐结构与精选拟音;
  • **原生对白:**台词短、说话人明确,说话时有动作,之后有反应和下一节拍;
  • **参考音频:**每个常驻角色一条经授权的声线锚点,锁定语言、语速、音色与强度;
  • **外部配音:**使用口型工具,或通过反应、侧脸、远景、插镜和画外音保护可信度;
  • **混合:**保留好的原生英雄台词,旁白、电话声、环境、音乐与问题音频后期统一。
所有模式都要分轨。连续场景跨镜头保持房间底噪/雨声,使用 J/L cut,压低 BGM 避让对白,避免每次切画面都重启声音。
多人原生对白为每句记录
speaker_id
、可见/画外声源、镜头中的嘴型主体、非说话者嘴部状态和轮次。工作人员讲规则时必须由工作人员本人可见开口,或明确使用有来源的广播/画外音;其他角色闭嘴并只做倾听反应。错人开口、声音串角色或非说话者同步嘴型为 P0,不得靠字幕解释。
Read audio-and-dubbing.md, choose by scene instead of blindly unifying the entire film.
  • No dialogue: Unified ambient background, music structure, and selected foley;
  • Native dialogue: Short lines, clear speakers, actions while speaking, followed by reactions and next beats;
  • Reference audio: One authorized voice anchor per permanent character, locking language, speaking rate, timbre, and intensity;
  • External dubbing: Use lip-sync tools, or protect credibility through reactions, side shots, long shots, insert shots, and voiceover;
  • Hybrid: Keep good native hero lines, unify narration, phone sounds, environment, music, and problematic audio in post-production.
All modes require separate tracks. Maintain room background noise/rain across shots in continuous scenes, use J/L cuts, lower BGM to avoid dialogue, and avoid restarting sound every time the footage cuts.
For multi-person native dialogue, record
speaker_id
, visible/off-screen sound source, lip-sync subject in the shot, non-speaker's mouth state, and turn order for each line. When staff explain rules, they must be visible opening their mouths, or clearly use sourced broadcast/voiceover; other characters keep their mouths closed and only make listening reactions. Wrong person speaking, voice cross-character, or non-speaker syncing lip movements is P0, cannot be explained by subtitles.

门禁式生成顺序

Gate-Based Generation Sequence

  1. 验收故事和剧情段计划;
  2. 验收人物、风格、世界与道具母版;
  3. 先生成钩子段或最高风险段;
  4. 下载并直接检查真实输出;
  5. 30 秒段制作 12 帧接触表,至少看开头、25%、50%、75% 和结尾;
  6. 慢放检查接触峰值、穿模和物理;
  7. 有音频时转写并人工听;
  8. 用证据标记
    accepted
    repair
    rejected
  9. 提取验收尾帧,更新连续性状态;
  10. 之后才解锁依赖段。
可并行:独立母版、音乐探索、声线、非依赖插镜与备选钩子。不可并行:尚未获得已验收父状态的连续剧情段。
  1. Accept the story and plot segment plan;
  2. Accept character, style, world, and prop master versions;
  3. Generate the hook segment or highest-risk segment first;
  4. Download and directly check the real output;
  5. Create a 12-frame contact sheet for 30-second segments, at least watch the beginning, 25%, 50%, 75%, and end;
  6. Slow down to check contact peaks, mesh penetration, and physics;
  7. Transcribe and manually listen if there is audio;
  8. Mark
    accepted
    ,
    repair
    , or
    rejected
    with evidence;
  9. Extract the accepted last frame, update continuity status;
  10. Only unlock dependent segments after that.
Parallelizable: independent master versions, music exploration, voices, non-dependent insert shots, and alternative hooks. Non-parallelizable: continuous plot segments that have not obtained the accepted parent state.

物理与解剖门

Physics and Anatomy Gate

阅读 physics-and-failure-modes.md。生成前列出可见
physical_contact_points
,用符合主体解剖的动作原语。人手、动物爪、机器、液体、布料、车辆和魔法效果采用不同接触规则。
要求顺序:
接近 → 接触 → 施力 → 响应 → 稳定结果
“倒车、转向、脱落、碰撞”必须编译成
physical_start_state → control_input → world/screen motion vectors → phase_sequence → physical_end_state
。一个生成段只承担一个高风险精确物理事件;复杂事件拆成建立、峰值和后果镜头。提示词意图正确但画面实际方向错误,仍判失败。
拒绝身体/物体穿透、额外或缺失动物、不可能支撑、道具瞬移、手机方向反转、服装结构改变、无解释复位或因果不可见。
Read physics-and-failure-modes.md. List visible
physical_contact_points
before generation, use action primitives that conform to the subject's anatomy. Human hands, animal claws, machines, liquids, fabrics, vehicles, and magical effects use different contact rules.
Required sequence:
approach → contact → force application → response → stable result
.
"Reverse, turn, fall off, collide" must be compiled into
physical_start_state → control_input → world/screen motion vectors → phase_sequence → physical_end_state
. One generation segment only handles one high-risk precise physical event; complex events are split into setup, peak, and consequence shots. If the prompt intent is correct but the actual footage direction is wrong, it is still judged as failed.
Reject body/object penetration, extra or missing animals, impossible support, prop teleportation, mobile phone direction reversal, clothing structure change, unexplained reset, or invisible causality.

质检、修复与剪辑

Quality Inspection, Repair, and Editing

使用 qa-report.mdqa-and-repair.md,分别运行:
  1. **剧情:**冷观众能复述目标 → 阻碍 → 选择/行动 → 结果;
  2. **身份/画风:**人物、动物、产品、服装、材质和介质保持母版;
  3. **物理/连续性:**接触、方向、状态和接缝可信;
  4. **声音:**对白、声线、响度、环境、BGM 和桥接连续。
喜剧/动画化表演另做静音验收:观众能指出笑点正常状态、可见原因、信息差、反应帧、揭晓和对主线的后果;角色在 15–30 秒段内有至少三个可区分表演状态。若只能靠台词说明“东西去了哪里”或“为什么好笑”,判为剧情/表演失败。
只修最小失败单元:局部缺陷用编辑,接缝问题用桥接/续写,身份/解剖/因果崩坏才重生,纯音频问题优先重混。不得用音乐掩盖剧情失败。
多个候选先过 MUST-HIT 二元门,再用 candidate-scorecard.csv 比较剧情、身份、空间、物理、表演、声音和可剪辑性;关键候选尽量盲化模型、种子、费用与生成顺序。评审者先负责发现和分级,创作/修复者再决定改法,避免同一轮为自己的结果辩护。
进入成片阶段阅读 editing-delivery-and-release.md,用 edit-timeline.csv 只装入已验收的真实可用区间,统一工作格式与交付配置,保留字幕/艺术字和音频分轨,并检查缺失媒体、时间线空洞、黑/冻/重复帧、响度、字幕安全区和横竖版本。生成完成不等于剪辑完成,上传成功不等于发布通过。
Use qa-report.md and qa-and-repair.md, run separately:
  1. Plot: Cold audience can retell goal → obstacle → choice/action → result;
  2. Identity/style: Characters, animals, products, clothing, materials, and medium maintain the master version;
  3. Physics/continuity: Contact, direction, state, and seams are credible;
  4. Sound: Dialogue, voice, loudness, environment, BGM, and bridging are continuous.
Comedy/animated performance also requires silent acceptance: the audience can point out the normal state of the punchline, visible cause, information gap, reaction frame, reveal, and consequence to the main line; the character has at least three distinguishable performance states in a 15–30 second segment. If the plot can only be explained by lines like "where the thing went" or "why it's funny", it is judged as plot/performance failure.
Only repair the smallest failed unit: use editing for local defects, bridging/continuation for seam problems, regeneration only for identity/anatomy/causality collapse, and prioritize remixing for pure audio issues. Do not use music to cover plot failures.
Multiple candidates first pass the MUST-HIT binary gate, then compare plot, identity, space, physics, performance, sound, and editability using candidate-scorecard.csv; blind the model, seed, cost, and generation order as much as possible for key candidates. Reviewers are responsible for discovering and grading first, then creators/repairers decide the modification method, avoiding defending their own results in the same round.
Read editing-delivery-and-release.md when entering the final film stage, use edit-timeline.csv to only load accepted real usable intervals, unify working format and delivery configuration, retain subtitle/art text and audio tracks, and check for missing media, timeline gaps, black/frozen/repeated frames, loudness, subtitle safe area, and horizontal/vertical versions. Generation completion does not equal editing completion; upload success does not equal release approval.

按任务范围交付

Deliver by Task Scope

不要把完整制作包强加给局部请求。只创建下一步确实会读取或验证的文件:
  • **诊断/建议:**问题证据、优先级和最小修复;不生成空白 CSV 或生产目录。
  • **故事/剧本:**生产简报、剧情合同、剧本/节拍;只在身份、道具或季线确有连续性时建对应母版。
  • **分镜/生产包:**加入覆盖方案/冻结关键帧、生成段、剪辑节拍、参考权威、模型能力卡、费用与声音路线。
  • **执行:**加入请求预览与指纹、幂等键、任务/恢复记录、候选谱系、费用账本、真实输出、QA 证据、验收尾帧和连续性状态。
  • **成片/发布:**加入带来源链的正式时间线、媒体标准化、混音、字幕/艺术字、封面、编码与冷观众报告。
已有素材或上游状态已足够时复用,不为“流程完整”重复生成。统一命名:
BLOCK_ID_vN.mp4
_tail.png
_contact_12.jpg
_audio.wav
_QA.md
Do not impose a full production package on local requests. Only create files that will actually be read or verified in the next step:
  • Diagnosis/suggestion: Problem evidence, priority, and minimal repair; do not generate blank CSV or production directories.
  • Story/script: Production brief, plot contract, script/beats; only create corresponding master versions when identity, props, or season lines have definite continuity.
  • Storyboard/production package: Add coverage options/frozen keyframes, generation segments, editing beats, reference authorities, model capability cards, costs, and audio routes.
  • Execution: Add request preview and fingerprint, idempotency key, task/recovery record, candidate lineage, cost ledger, real output, QA evidence, accepted last frame, and continuity status.
  • Final film/release: Add formal timeline with source chain, media standardization, mixing, subtitle/art text, cover, encoding, and cold audience report.
Reuse when existing materials or upstream states are sufficient, do not regenerate for "complete process". Unified naming:
BLOCK_ID_vN.mp4
,
_tail.png
,
_contact_12.jpg
,
_audio.wav
,
_QA.md
.

证据边界

Evidence Boundaries

项目文档必须区分:
  • **文档能力:**从当前正式参数页或接口结构核实;
  • **实测结果:**特定模型、入口、日期与配置下观察到;
  • **创作偏好:**为当前故事主动选择;
  • **推断:**有可能但未经直接验证。
不得把动态价格、时长限制或模型质量写成永恒事实。
Project documents must distinguish between:
  • Documented capabilities: Verified from the current official parameter page or interface structure;
  • Measured results: Observed under specific models, entrances, dates, and configurations;
  • Creative preferences: Actively selected for the current story;
  • Inferences: Possible but not directly verified.
Do not write dynamic prices, duration limits, or model quality as eternal facts.