Loading...
Loading...
Compare original and translation side by side
| Mode | Speed | Quality | Use Case |
|---|---|---|---|
| Quick (default) | Fast | Good | Drafts, simple documents |
| Heavy | Slower | Best | Final documents, complex layouts |
| 模式 | 速度 | 质量 | 适用场景 |
|---|---|---|---|
| 快速(默认) | 快 | 良好 | 草稿文档、简单格式文档 |
| 深度 | 较慢 | 最优 | 终稿文档、复杂布局文档 |
undefinedundefinedundefinedundefinedundefinedundefinedundefinedundefined| Format | Quick Mode Tool | Heavy Mode Tools |
|---|---|---|
| pymupdf4llm | pymupdf4llm + markitdown | |
| DOCX | pandoc | pandoc + markitdown |
| PPTX | markitdown | markitdown + pandoc |
| XLSX | markitdown | markitdown |
| 格式 | 快速模式工具 | 深度模式工具 |
|---|---|---|
| pymupdf4llm | pymupdf4llm + markitdown | |
| DOCX | pandoc | pandoc + markitdown |
| PPTX | markitdown | markitdown + pandoc |
| XLSX | markitdown | markitdown |
| Segment Type | Selection Criteria |
|---|---|
| Tables | More rows/columns, proper header separator |
| Images | Alt text present, local paths preferred |
| Headings | Proper hierarchy, appropriate length |
| Lists | More items, nested structure preserved |
| Paragraphs | Content completeness |
| 片段类型 | 选择标准 |
|---|---|
| 表格 | 包含更多行/列,表头分隔符格式正确 |
| 图片 | 包含替代文本,优先选择本地路径 |
| 标题 | 层级结构正确,长度合适 |
| 列表 | 包含更多条目,嵌套结构完整保留 |
| 段落 | 内容完整度高 |
undefinedundefined
Output:
- Images: `assets/img_page1_1.png`, `assets/img_page2_1.jpg`
- Metadata: `assets/images_metadata.json` (page, position, dimensions)
输出内容:
- 图片: `assets/img_page1_1.png`, `assets/img_page2_1.jpg`
- 元数据: `assets/images_metadata.json`(包含页码、位置、尺寸信息)undefinedundefinedundefinedundefined| Metric | Pass | Warn | Fail |
|---|---|---|---|
| Text Retention | >95% | 85-95% | <85% |
| Table Retention | 100% | 90-99% | <90% |
| Image Retention | 100% | 80-99% | <80% |
| 指标 | 通过 | 警告 | 失败 |
|---|---|---|---|
| 文本保留率 | >95% | 85-95% | <85% |
| 表格保留率 | 100% | 90-99% | <90% |
| 图片保留率 | 100% | 80-99% | <80% |
undefinedundefinedundefinedundefinedundefinedundefinedundefinedundefinedundefinedundefined
**FontBBox warnings during PDF conversion**
- Harmless font parsing warnings, output is still correct
**Images missing from output**
- Use Heavy Mode for better image preservation
- Or extract separately with `scripts/extract_pdf_images.py`
**Tables broken in output**
- Use Heavy Mode - it selects the most complete table version
- Or validate with `scripts/validate_output.py`
**FontBBox warnings during PDF conversion**
- 这是无害的字体解析警告,输出内容仍然正确
**Images missing from output**
- 使用深度模式可提升图片保留效果
- 或通过 `scripts/extract_pdf_images.py` 单独提取图片
**Tables broken in output**
- 使用深度模式 - 它会选择最完整的表格版本
- 或通过 `scripts/validate_output.py` 验证转换结果| Script | Purpose |
|---|---|
| Main orchestrator with Quick/Heavy mode |
| Merge multiple markdown outputs |
| Quality validation with HTML report |
| PDF image extraction with metadata |
| Windows to WSL path converter |
| 脚本 | 用途 |
|---|---|
| 主编排工具,支持快速/深度模式 |
| 合并多个Markdown输出文件 |
| 转换质量验证,生成HTML报告 |
| 提取PDF中的图片并生成元数据 |
| Windows与WSL路径转换工具 |
references/heavy-mode-guide.mdreferences/tool-comparison.mdreferences/conversion-examples.mdreferences/heavy-mode-guide.mdreferences/tool-comparison.mdreferences/conversion-examples.md