captions-overlay

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Captions Overlay Doctrine

字幕叠加原则

Overlay doctrine — supplements the upstream
embedded-captions
skill. Applies ON TOP of it; do not expect it folded into the upstream skill.
Two ideas combine here. First, the caption model — every spoken phrase is
drop
,
rail
, or
embed
, and embed is the scarce earned peak, not the default. Second, the overlay law — a caption line is composited ON TOP of the film as an overlay; it is NOT a reserved zone, so you never shift content up or leave a dead band to "make room" for it. The two reinforce each other: because captions ride as an overlay (the verbatim rail in front, the occasional embed behind the subject), the composition keeps its full frame and centers on the true vertical center.
叠加原则——是对上游
embedded-captions
skill的补充。需应用于该skill之上;请勿期望将其整合到上游skill中。
该原则包含两个核心理念。其一为字幕模型——每一段口语内容都属于
drop
rail
embed
中的一种,其中embed是稀缺的、需精心打造的高潮内容,而非默认选项。其二为叠加法则——字幕行作为叠加层合成在影片上方;它并非预留区域,因此永远无需上移内容或留出空白区域来“为字幕腾空间”。这两点相辅相成:由于字幕以叠加形式呈现(逐字rail在前景,偶尔的embed在主体后方),画面构图可保留完整帧,并以真实垂直中心为基准居中。

The caption model — drop / rail / embed

字幕模型——drop / rail / embed

Every spoken phrase is one of three things (verbatim from
embedded-captions
):
WhatHow it's shown
dropfiller — um/uh, stutters, self-correctionsnot shown
railthe default — ordinary spoken content (verbatim)clean lower-third subtitle, in front, readable. A punch word can get an inline
emphasis
highlight (accent colour / active-word pop) — it stays on the rail.
embeda promoted peak — the headline beatone big word composited behind the subject (matte occlusion), designed entrance + exit
The rail carries most of the text; embed is the scarce, earned peak — ≤1 per beat, never two adjacent/co-visible, spaced ≥ a beat apart. A short clip → usually one embed; a long explainer → ~one per section. Embedding every word is the common mistake.
This is the Standard mode shape (rail = the verbatim lower-third; embed = the climax composited behind the subject). Cinematic mode drops the rail and makes everything embed-style — use it only for pure-cinematic asks, never for explainer / voiceover where the words must read.
每一段口语内容都属于以下三类之一(直接引用自
embedded-captions
):
说明展示方式
drop填充内容——嗯/呃、口吃、自我修正内容不显示
rail默认类型——普通口语内容(逐字)简洁的下部三分之一字幕,置于前景,易于阅读。重点词汇可添加行内
emphasis
高亮(强调色/动态词汇突出显示)——仍保留在rail上。
embed升级的高潮内容——核心亮点单个大词汇合成在主体后方(遮罩遮挡),设计专属入场和退场动画
rail承载大部分文本;embed是稀缺的、需精心打造的高潮内容——每段节拍最多1个embed,绝不出现相邻或同时可见的两个embed,间隔至少一个节拍。短剪辑→通常仅1个embed;长讲解视频→每部分约1个embed。将每个词汇都设为embed是常见误区。
这是标准模式的形态(rail=逐字下部三分之一字幕;embed=合成在主体后方的高潮内容)。电影模式会移除rail,所有内容都采用embed样式——仅适用于纯电影类需求,绝不能用于需要清晰阅读文字的讲解/旁白视频。

Rail-first, embed-scarce (the load-bearing rules)

优先使用rail,embed需稀缺(核心规则)

Quoted from the
embedded-captions
non-negotiables:
  • Rail-first for talking-head / explainer. Don't embed the whole transcript — most text is the rail; embed only peaks. Embedding everything is the default mistake.
  • Embed is scarce + spaced. ≤1 embed per sentence/beat, never two adjacent or co-visible, ≥ a beat apart, at most one
    apex
    . climax = per-beat peak, not "the single payoff of the entire clip."
引用自
embedded-captions
的不可协商规则:
  • 访谈/讲解视频优先使用rail。不要将整个脚本设为embed——大部分文本应使用rail;仅将高潮内容设为embed。将所有内容都设为embed是典型的错误做法。
  • embed需稀缺且间隔分布。每句/每节拍最多1个embed,绝不出现相邻或同时可见的两个embed,间隔至少一个节拍,最多1个
    apex
    。高潮内容=每节拍的亮点,而非“整个剪辑的单一回报点”。

The overlay law — captions are NOT a reserved band

叠加法则——字幕并非预留区域

In a generated launch composition, when captions are enabled, finalize composites a small, minimal word-by-word caption line as an overlay layer ON TOP of the whole film (a single text line, bottom-centered, roughly the bottom ~5-8% of canvas height). It is an overlay, not a reserved zone (verbatim from constraint #13 of the product-launch-video scene agent):
  • Center the composition on the TRUE vertical center — y = H / 2 (landscape 540, portrait 960). Do not shift content up to "make room" for captions; a composition centered at 0.42 × H with a dead lower band is the bug, not the fix.
  • Content may extend to the canvas bottom. Full-bleed subjects, rails, and backgrounds all welcome.
  • One soft courtesy rule: avoid parking critical small readable text (a URL line, a legal line, a sub-caption) exactly in the bottom ~80px center span where the caption line sits — the overlay would fight it. Large imagery / cards / ambient content under the captions is fine; the caption skin is designed to read over content.
  • There is no machine keep-out gate (the old
    captions.mjs keepout
    check is retired). Finalize snapshot QA judges caption-over-content legibility visually.
When captions are disabled: identical positioning freedom — the overlay simply doesn't exist.
在生成的发布类视频构图中,当启用字幕时,需将小型、极简的逐字字幕行作为叠加层合成在整个影片上方(单行文本,底部居中,约占画布高度的5-8%)。它是叠加层,而非预留区域(直接引用自动画片发布视频场景agent的约束条件#13):
  • 将构图居中于真实垂直中心——y = H / 2(横屏为540,竖屏为960)。请勿上移内容来“为字幕腾空间”;将构图居中于0.42×H并留有底部空白区域是错误做法,而非解决方案。
  • 内容可延伸至画布底部。欢迎全bleed主体、rail和背景。
  • **一条软性礼貌规则:**避免将关键的小型可读文本(URL行、法律声明行、子字幕)恰好放在字幕行所在的底部约80px中心区域——叠加层会与之冲突。字幕下方的大尺寸图像/卡片/背景内容则无问题;字幕样式设计为可在内容上方清晰阅读。
  • 不再有机器避让检查(旧版
    captions.mjs keepout
    检查已停用)。最终快照QA需通过视觉判断字幕与内容叠加后的可读性。
**当字幕禁用时:**构图定位自由度相同——仅叠加层不存在而已。

Why these two rules are one doctrine

为何这两条规则同属一个原则

The model says the rail rides in front and an embed is a rare word composited behind the subject — both are layers added to footage that ships untouched. The overlay law says the caption line is a layer composited on top of the whole film, not a band carved out of the layout. So in both the captioning pipeline and the launch-video pipeline, captions are an overlay you add, not a zone you reserve:
  • Keep the full frame; center on true center; let content run to the edges.
  • Make the rail (or the small overlay caption line) carry the verbatim words.
  • Promote a word to an embed only at a genuine peak — scarce, spaced, never two at once.
  • Reserve nothing; judge legibility of captions-over-content visually, not by a keep-out gate.
字幕模型指出,rail置于前景,而embed是合成在主体后方的稀有词汇——两者都是添加到原始素材上的图层。叠加法则指出,字幕行是合成在整个影片上方的图层,而非从布局中划出的区域。因此,无论是字幕制作流程还是发布视频制作流程,字幕都是需要添加的叠加层,而非预留区域:
  • 保留完整帧;以真实中心为基准居中;让内容延伸至边缘。
  • 让rail(或小型叠加字幕行)承载逐字文本。
  • 仅在真正的高潮时刻将词汇升级为embed——稀缺、间隔分布,绝不同时出现两个。
  • 不预留任何区域;通过视觉判断字幕与内容叠加后的可读性,而非依赖避让检查。