music-caption-rewriter

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Music Caption Rewriter

音乐标题重写器

Transform the user's musical intent into a new, generation-oriented structured caption. Find useful references through progressive disclosure: route to a small style family, compare compact cards, then read only the selected complete templates.
Use natural-language reasoning and local text files only. Do not execute scripts, build a database, calculate embeddings, call external APIs, or scan all 1,000 templates.
将用户的音乐创作意图转换为面向生成的全新结构化标题。通过渐进式披露找到有用的参考:先定位到小型风格类别,对比精简卡片,再仅阅读选中的完整模板。
仅使用自然语言推理和本地文本文件。不得执行脚本、构建数据库、计算嵌入向量、调用外部API或扫描全部1000个模板。

Inputs

输入参数

Accept:
  • Caption
    : required natural-language music description.
  • Lyrics
    : optional lyrics containing bracketed section or control tags.
  • Additional constraints: optional length, format, exclusions, or creative direction.
Use lyric text only to infer broad emotional context and narrative intensity. Never quote, paraphrase, summarize, or reproduce it. Treat only bracketed tags as executable structural, musical, vocal, or production directives.
接受以下内容:
  • Caption
    :必填的自然语言音乐描述。
  • Lyrics
    :可选的包含方括号分段或控制标签的歌词。
  • 额外约束:可选的长度、格式、排除项或创作方向。
仅使用歌词文本推断宽泛的情感背景和叙事强度。不得引用、改写、总结或重现歌词内容。仅将方括号中的标签视为可执行的结构、音乐、人声或制作指令。

Workflow

工作流程

Follow these stages in order:
  1. Build a private Music Brief from the inputs.
  2. Resolve explicit constraints and section-local directives.
  3. Read references/genre-router.md.
  4. Read one primary family index and, only when useful, one secondary family index.
  5. Select up to three references with distinct roles.
  6. Read only the complete template files named by those cards.
  7. Design a coherent section-by-section timeline.
  8. Render and validate the new caption.
Do not expose the Music Brief, routing choices, scores, or template IDs unless the user requests diagnostics.
按以下步骤依次执行:
  1. 根据输入内容构建私人音乐简报。
  2. 解析明确的约束条件和分段本地指令。
  3. 阅读references/genre-router.md
  4. 阅读一个主类别索引,仅在需要时阅读一个次要类别索引。
  5. 选择最多三个具有不同作用的参考模板。
  6. 仅阅读这些卡片指定的完整模板文件。
  7. 设计连贯的分段时间线。
  8. 生成并验证新标题。
除非用户要求诊断信息,否则不得披露音乐简报、路由选择、评分或模板ID。

Build the Music Brief

构建音乐简报

Extract only supported or reasonably inferred values:
  • macro genre, subgenres, and cultural or market style
  • mood and emotional arc
  • approximate tempo, meter, and groove
  • vocal presence, gender, register, timbre, and delivery
  • core instruments and production texture
  • section structure and section-specific changes
  • spatial character and explicit exclusions
Classify each value internally as
explicit
,
tagged
,
inferred
, or
unspecified
.
Do not invent a precise key, BPM, vocal gender, melodic interval, or production technique when a broader description is sufficient.
Preserve an explicit instrumental request. Do not add vocals. If vocal presence is unspecified, choose a conservative treatment supported by the user's description and the closest style family.
仅提取支持或合理推断的信息:
  • 宏观流派、子流派以及文化或市场风格
  • 情绪和情感弧线
  • 大致速度、节拍和律动
  • 人声存在形式、性别、音域、音色和演唱方式
  • 核心乐器和制作质感
  • 分段结构和分段特定变化
  • 空间特征和明确排除项
将每个信息在内部分类为
explicit
(明确指定)、
tagged
(标签指定)、
inferred
(推断得出)或
unspecified
(未指定)。
当宽泛描述足够时,不得编造精确的调号、BPM、人声性别、旋律音程或制作技巧。
保留明确的器乐要求,不得添加人声。如果人声存在形式未指定,选择符合用户描述和最接近风格类别的保守处理方式。

Resolve Constraints

约束条件解析

Apply this precedence:
  1. Explicit user requirements and exclusions.
  2. Section-local directives from lyric tags, within that section.
  3. Strong implications from the user's Caption.
  4. Selected reference characteristics.
  5. Conservative musical defaults.
A section tag may change its local arrangement without replacing the song's global genre. Preserve a hard user exclusion when a tag conflicts with it.
When two explicit instructions conflict, prefer the more specific and later instruction if the intent remains clear. Otherwise make the smallest musically coherent compromise.
Never silently reverse an explicit vocal gender, instrumental requirement, tempo limit, required instrument, or prohibited element.
遵循以下优先级:
  1. 用户明确的要求和排除项。
  2. 来自歌词标签的分段本地指令(仅适用于对应分段)。
  3. 用户
    Caption
    中的强烈暗示。
  4. 选中参考模板的特征。
  5. 保守的音乐默认设置。
分段标签可改变其所在分段的编排,但不得替换歌曲的全局流派。当标签与用户明确排除项冲突时,保留用户的排除要求。
当两个明确指令冲突时,如果意图清晰,优先选择更具体且较晚提出的指令。否则做出最小的音乐连贯性妥协。
不得擅自推翻明确的人声性别、器乐要求、速度限制、必填乐器或禁止元素。

Route by Progressive Disclosure

渐进式披露路由

Read the genre router first. Choose:
  • one primary family for a clear genre request
  • one primary and one secondary family for an explicit fusion
  • at most two plausible families for an ambiguous genre
  • the general pop and ballad family when only mood or imagery is available
Use genre, groove, instrumentation, and cultural context as stronger routing signals than generic adjectives such as
emotional
,
epic
,
dark
, or
modern
.
Read only the family indexes selected by the router. Do not inspect every family index, reconstruct a global catalog, or scan every template filename.
首先阅读流派路由文件。选择:
  • 对于明确的流派请求,选择一个主类别
  • 对于明确的风格融合请求,选择一个主类别和一个次要类别
  • 对于模糊的流派,选择最多两个合理的类别
  • 当仅提供情绪或意象时,选择通用流行和民谣类别
将流派、律动、乐器配置和文化背景作为比
emotional
(情绪化)、
epic
(史诗感)、
dark
(黑暗风)或
modern
(现代感)等通用形容词更强的路由信号。
仅阅读路由选择的类别索引。不得检查每个类别索引、重建全局目录或扫描所有模板文件名。

Select References

选择参考模板

Compare cards in the selected family indexes using this priority:
  1. Genre and subgenre compatibility.
  2. Explicit requirements and exclusions.
  3. Groove and tempo compatibility, including plausible half-time or double-time relationships.
  4. Vocal configuration.
  5. Instrumentation.
  6. Mood and emotional arc.
  7. Production character.
Apply a strong penalty to direct conflicts. Prefer a close musical family over a card that merely shares mood vocabulary.
Select up to three references with different responsibilities:
  • Foundation
    : closest overall identity, groove, and songwriting language.
  • Modifier
    : best source for a requested secondary genre, vocal character, cultural color, or production texture.
  • Arrangement
    : best source for section development, energy contour, transitions, and instrument lifecycle.
Use one or two references when the request is simple. Do not select a weak match merely to reach three.
按照以下优先级对比选中类别索引中的卡片:
  1. 流派和子流派兼容性。
  2. 明确的要求和排除项。
  3. 律动和速度兼容性,包括合理的半速或双倍速关系。
  4. 人声配置。
  5. 乐器配置。
  6. 情绪和情感弧线。
  7. 制作特征。
对直接冲突的模板施加严重惩罚。优先选择相近的音乐类别,而非仅共享情绪词汇的卡片。
选择最多三个具有不同职责的参考模板:
  • Foundation
    (基础模板):最匹配整体风格、律动和创作语言的模板。
  • Modifier
    (修饰模板):最适合提供所需次要流派、人声特征、文化色彩或制作质感的模板。
  • Arrangement
    (编排模板):最适合提供分段发展、能量曲线、过渡和乐器生命周期的模板。
当请求简单时,使用一到两个参考模板。不得为凑够三个而选择匹配度低的模板。

Use Templates Safely

安全使用模板

Use the Foundation for broad musical identity, the Modifier only for its matched dimension, and the Arrangement reference only for timeline logic.
Do not inherit unsupported details such as a template's exact key, BPM, vocalist, instruments, emotional story, or section order.
Do not copy sentences, distinctive phrases, or a template's complete structure. Synthesize a new caption around the user's brief.
使用基础模板确定广泛的音乐风格,仅使用修饰模板的匹配维度,仅使用编排模板的时间线逻辑。
不得继承不支持的细节,例如模板的精确调号、BPM、歌手、乐器、情感故事或分段顺序。
不得复制句子、独特短语或模板的完整结构。围绕用户的音乐简报合成新标题。

Plan the Timeline

编排时间线规划

Build around the user's section tags when present. Otherwise choose only sections appropriate to the style, for example:
Intro → Verse → Pre-Chorus → Chorus → Verse → Chorus → Bridge → Final Chorus → Outro
For every included section, state what enters, exits, changes, or intensifies. Keep instrument behavior continuous and make transitions musically plausible.
Create a readable energy arc rather than a static equipment list or a stack of production terminology.
当存在用户提供的分段标签时,围绕这些标签构建时间线。否则仅选择符合风格的分段,例如:
Intro → Verse → Pre-Chorus → Chorus → Verse → Chorus → Bridge → Final Chorus → Outro
对于每个包含的分段,说明新增、移除、变化或增强的元素。保持乐器行为连贯,确保过渡符合音乐逻辑。
创建清晰可读的能量曲线,而非静态设备列表或堆砌制作术语。

Output Contract

输出规范

Write the final caption in English unless the user explicitly requests another language.
Return exactly these three top-level headings in this order:
除非用户明确要求其他语言,否则最终标题使用英文撰写。
严格按照以下顺序返回三个顶级标题:

Global Metadata

Global Metadata

Include genre and subgenres, tempo, emotional progression, and overall sonic and production profile. Use an exact BPM only when explicit or strongly justified; otherwise use a range or qualitative tempo. Include key and scale only when explicit or musically useful.
包含流派和子流派、速度、情感变化以及整体声音和制作概况。仅在明确指定或有充分理由时使用精确BPM;否则使用范围或定性描述。仅在明确指定或对音乐有用时包含调号和音阶。

Vocal Details

Vocal Details

For vocal music, describe the lead configuration, timbre, register, delivery, harmony or backing vocals, and restrained vocal effects.
For instrumental music, state that the piece is instrumental and identify the instrument or texture carrying the lead melodic role.
Do not invent lyrical subject matter or reproduce lyrics.
对于声乐作品,描述主唱配置、音色、音域、演唱方式、和声或伴唱,以及克制的人声效果。
对于器乐作品,说明作品为纯器乐,并指出承担主旋律的乐器或音色。
不得编造歌词主题或重现歌词内容。

Arrangement

Arrangement

Describe the song as a section-by-section timeline. Explain primary and secondary instrument lifecycles, groove development, transitions, embellishments, texture, and spatial effects only where relevant.
Prefer concrete musical changes over decorative prose. Default to approximately 250–450 English words unless the user requests another length.
Do not include a song title, track ID, selected template ID, reasoning trace, or copied lyric line.
按分段时间线描述歌曲。仅在相关时说明主要和次要乐器的生命周期、律动发展、过渡、装饰、质感和空间效果。
优先使用具体的音乐变化描述,而非装饰性散文。默认篇幅约为250-450英文单词,除非用户要求其他长度。
不得包含歌曲标题、曲目ID、选中模板ID、推理过程或复制的歌词行。

Machine-Readable Output

机器可读输出

Return JSON or JSONL only when explicitly requested. Include original inputs and
rewritten_caption
. Include routing diagnostics or selected template IDs only when explicitly requested.
Never include complete template contents in machine-readable output unless the user specifically asks for them.
仅在明确要求时返回JSON或JSONL格式。包含原始输入和
rewritten_caption
。仅在明确要求时包含路由诊断信息或选中模板ID。
除非用户特别要求,否则不得在机器可读输出中包含完整模板内容。

Validate Before Returning

返回前验证

Verify that:
  • every explicit user constraint is preserved
  • every actionable section tag appears in the matching section
  • no quoted, paraphrased, or summarized lyric content, title, or track ID appears
  • an instrumental request remains instrumental
  • vocal gender is not contradicted
  • genre and local modifiers coexist coherently
  • the three required headings are present
  • the arrangement follows a readable timeline
  • instruments have coherent entrances, changes, and exits
  • exact BPM, key, and technical details are not fabricated
  • no template sentence or complete template structure is copied
  • the caption is specific enough to guide generation without becoming an essay
Revise once when any check fails, then return only the corrected result.
验证以下内容:
  • 所有明确的用户约束均已保留
  • 所有可执行的分段标签均出现在对应分段中
  • 未出现引用、改写或总结的歌词内容、标题或曲目ID
  • 器乐请求仍保持纯器乐
  • 人声性别未被违背
  • 流派和本地修饰符连贯共存
  • 三个必填标题均存在
  • 编排遵循清晰可读的时间线
  • 乐器的入场、变化和退场连贯合理
  • 未编造精确的BPM、调号和技术细节
  • 未复制模板句子或完整模板结构
  • 标题足够具体以指导生成,且不会冗长成篇
当任何检查未通过时,修改一次,然后仅返回修正后的结果。

Static Library Maintenance

静态库维护

Keep the library entirely text-based. When adding a template:
  1. Add one complete Caption file under
    templates/
    .
  2. Add one compact card to exactly one family index linked from the genre router.
  3. Record compatible secondary families in that card instead of duplicating it.
  4. Confirm that the card ID matches the template filename and that its path exists.
  5. Update the family count in that index.
Do not add scripts, generated catalogs, embeddings, vector stores, databases, or external service configuration.
保持库完全基于文本。添加模板时:
  1. templates/
    下添加一个完整的标题文件。
  2. 在流派路由链接的恰好一个类别索引中添加一张精简卡片。
  3. 在该卡片中记录兼容的次要类别,而非重复添加卡片。
  4. 确认卡片ID与模板文件名匹配,且路径存在。
  5. 更新该索引中的类别计数。
不得添加脚本、生成的目录、嵌入向量、向量存储、数据库或外部服务配置。