text-splitter

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

文本分割工具

Text Splitting Tool

功能

Features

将文本分割成指定大小的块,保持语义完整性,便于后续的批处理和分析。
Split text into chunks of specified size while maintaining semantic integrity, facilitating subsequent batch processing and analysis.

使用场景

Usage Scenarios

  • 处理超长文本,将其拆分为智能体可处理的片段。
  • 进行文本预处理,为后续的内容分析、摘要或评估任务提供标准化输入。
  • 优化文本处理流程,确保每个处理单元的大小可控且语义完整。
  • Process ultra-long text and split it into fragments that can be handled by agents.
  • Perform text preprocessing to provide standardized input for subsequent content analysis, summarization, or evaluation tasks.
  • Optimize text processing workflows to ensure the size of each processing unit is controllable and semantically complete.

核心能力

Core Capabilities

  • 语义完整性优先: 在分割时尽量避免截断句子或段落,优先在自然边界处(如句号、换行符)分割,确保每个块的语义完整性。
  • 保持原始格式: 在分割过程中,努力保留文本的原始格式和结构,例如Markdown、HTML标签等。
  • 精确控制块大小: 严格按照指定的大小限制进行分割,确保每个文本块不超过预设的最大长度。
  • 多种分割策略: 支持基于字符数、token数、段落数或特定分隔符等多种分割策略。
  • Semantic Integrity Priority: When splitting, try to avoid truncating sentences or paragraphs, and prioritize splitting at natural boundaries (such as periods, line breaks) to ensure the semantic integrity of each chunk.
  • Preserve Original Format: Strive to retain the original format and structure of the text during splitting, such as Markdown, HTML tags, etc.
  • Precise Chunk Size Control: Split strictly according to the specified size limit to ensure each text chunk does not exceed the preset maximum length.
  • Multiple Splitting Strategies: Support multiple splitting strategies based on character count, token count, paragraph count, or specific delimiters.

输入要求

Input Requirements

  • 文本内容: 待分割的原始文本。
  • 块大小限制: 每个文本块的最大长度(建议提供字符数或 token 数)。
  • 分割策略(可选): 指定分割时优先考虑的策略,如按句分割、按段落分割、按特定分隔符分割等。
  • Text Content: The original text to be split.
  • Chunk Size Limit: The maximum length of each text chunk (it is recommended to provide character count or token count).
  • Splitting Strategy (Optional): Specify the priority strategy for splitting, such as splitting by sentence, splitting by paragraph, splitting by specific delimiter, etc.

输出格式

Output Format

【文本分割报告】

- 原始文本长度: [整数] 字/Token
- 目标块大小: [整数] 字/Token
- 实际分割块数: [整数] 块
[Text Splitting Report]

- Original Text Length: [Integer] characters/Token
- Target Chunk Size: [Integer] characters/Token
- Actual Number of Chunks: [Integer] chunks

分割结果

Split Results

  • 块1 (长度: [整数] 字/Token): "[内容预览...]"
  • 块2 (长度: [整数] 字/Token): "[内容预览...]"
  • 块3 (长度: [整数] 字/Token): "[内容预览...]" ...
undefined
  • Chunk 1 (Length: [Integer] characters/Token): "[Content Preview...]"
  • Chunk 2 (Length: [Integer] characters/Token): "[Content Preview...]"
  • Chunk 3 (Length: [Integer] characters/Token): "[Content Preview...]" ...
undefined

约束条件

Constraints

  • 分割结果必须严格符合指定的块大小限制。
  • 确保在分割时最大限度地保持文本的语义完整性。
  • 输出格式必须结构化,清晰展示每个文本块的内容和长度。
  • 避免在输出中引入任何额外信息或解释,只提供分割结果。
  • The split results must strictly comply with the specified chunk size limit.
  • Ensure maximum preservation of text semantic integrity during splitting.
  • The output format must be structured, clearly displaying the content and length of each text chunk.
  • Avoid introducing any additional information or explanations in the output, only provide split results.

示例

Examples

参见
{baseDir}/references/examples.md
目录获取更多详细示例:
  • examples.md
    - 包含不同长度、不同分割策略和复杂文本结构的分割示例。
Refer to the
{baseDir}/references/examples.md
directory for more detailed examples:
  • examples.md
    - Contains splitting examples of different lengths, different splitting strategies, and complex text structures.

详细文档

Detailed Documentation

参见
{baseDir}/references/examples.md
获取关于文本分割工具的详细指导与案例。

Refer to
{baseDir}/references/examples.md
for detailed guidance and cases about the text splitting tool.

版本历史

Version History

版本日期变更
2.1.02026-01-11优化 description 字段,使其更精简并符合命令式语言规范;模型更改为 opus;优化功能、核心能力、输入要求、输出格式的描述,使其更符合命令式语言规范;添加使用场景、约束条件、示例和详细文档部分。
2.0.02026-01-11按官方规范重构
1.0.02026-01-10初始版本
VersionDateChanges
2.1.02026-01-11Optimized the description field to make it more concise and compliant with imperative language specifications; changed the model to opus; optimized the descriptions of functions, core capabilities, input requirements, and output formats to comply with imperative language specifications; added sections for usage scenarios, constraints, examples, and detailed documentation.
2.0.02026-01-11Restructured according to official specifications
1.0.02026-01-10Initial version