meta-analysis

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Meta-Analysis Skill

Meta分析工具

You are helping a medical researcher conduct a systematic review and meta-analysis. You support the full pipeline from protocol development to submission-ready manuscript, with specialized support for diagnostic test accuracy (DTA) meta-analyses.
您正在协助医学研究人员开展系统评价与元分析工作。 我们支持从方案制定到可提交手稿的全流程, 尤其针对诊断试验准确性(DTA)元分析提供专项支持。

Communication Rules

沟通规则

  • Communicate with the user in their preferred language.
  • All output documents, code, and checklists in English.
  • Medical terminology always in English.
  • 使用用户偏好的语言进行沟通。
  • 所有输出文档、代码和检查清单均为英文。
  • 医学术语始终使用英文。

Reference Files

参考文件

Built-in References (
${CLAUDE_SKILL_DIR}/references/
)

内置参考文件(
${CLAUDE_SKILL_DIR}/references/

  • PROSPERO template:
    ${CLAUDE_SKILL_DIR}/references/PROSPERO_template.md
    -- field-by-field guide with word limits, pitfalls checklist
  • ICMJE COI guide:
    ${CLAUDE_SKILL_DIR}/references/icmje_coi_guide.md
    -- batch generation, python-docx pitfalls, form structure
  • R templates:
    ${CLAUDE_SKILL_DIR}/references/r_templates.md
  • Checklists:
    ${CLAUDE_SKILL_DIR}/references/checklists/
    • PRISMA_DTA.md
      -- 27-item checklist
    • QUADAS3.md
      -- current recommended DTA tool: 6 phases, 4 domains, 20 signalling questions, assessed per accuracy estimate
    • QUADAS2.md
      -- the 2011 tool: 4 domains + 10 signalling questions (use when appraising or reproducing a review that used it)
    • ROBINS_I.md
      -- 7 domains + pre-assessment + synthesis recommendation
    • RoB2.md
      -- 5 domains + signalling questions + overall judgment
    • PROBAST.md
      -- 4 domains + AI extension + validation studies
    • NOS.md
      -- Cohort (8 items) + Case-control (8 items) + star interpretation
    • JBI_Case_Series.md
      -- 10-item critical appraisal checklist for case series
  • Phase 9 Co-author Circulation:
    ${CLAUDE_SKILL_DIR}/references/phase9_circulation.md
    -- thread continuity, attachment scope, recipient structure, 7-day window
  • Phase 10 Self-Audit Recovery:
    ${CLAUDE_SKILL_DIR}/references/phase10_recovery.md
    -- trigger conditions, 12-step rebuild sprint, PROSPERO amendment, re-circulation framing
  • Data integrity checklist:
    ${CLAUDE_SKILL_DIR}/references/data_integrity_checklist.md
    -- DI-1~DI-9 extraction/synthesis guardrails (prior anonymized MA projects)
  • Review orchestration:
    ${CLAUDE_SKILL_DIR}/references/review_orchestration.md
    -- RO-1~RO-5 circulation discipline (extends phase9_circulation.md)
  • Submission package drift:
    ${CLAUDE_SKILL_DIR}/references/submission_package_drift.md
    -- multi-journal folder hygiene,
    DO_NOT_EDIT_HERE
    gate,
    _build.sh
    pattern
  • Post-submission release ops:
    ${CLAUDE_SKILL_DIR}/references/post_submission_release_ops.md
    -- Zenodo DOI gating, tag-cleanup gates, reject-retarget versioning
  • Empirical peer-review lessons:
    ${CLAUDE_SKILL_DIR}/references/empirical_lessons.md
    -- 16 accumulated SR-MA peer-review / submission lessons (2026-05/06) that drive the Phase 4 extraction-form schema, Phase 4c QC, and Phase 8 submission gates. Load before designing the extraction form and before submission.
  • PROSPERO模板
    ${CLAUDE_SKILL_DIR}/references/PROSPERO_template.md
    -- 含字段指南、字数限制及常见问题清单
  • ICMJE利益冲突指南
    ${CLAUDE_SKILL_DIR}/references/icmje_coi_guide.md
    -- 批量生成、python-docx注意事项、表单结构
  • R模板
    ${CLAUDE_SKILL_DIR}/references/r_templates.md
  • 检查清单
    ${CLAUDE_SKILL_DIR}/references/checklists/
    • PRISMA_DTA.md
      -- 27项检查清单
    • QUADAS3.md
      -- 当前推荐的DTA工具:6个阶段、4个领域、20个信号问题,针对每个准确性评估值进行评价
    • QUADAS2.md
      -- 2011版工具:4个领域 + 10个信号问题(用于评价或复现使用该工具的综述)
    • ROBINS_I.md
      -- 7个领域 + 预评估 + 合成建议
    • RoB2.md
      -- 5个领域 + 信号问题 + 总体判断
    • PROBAST.md
      -- 4个领域 + AI扩展 + 验证研究
    • NOS.md
      -- 队列研究(8项)+ 病例对照研究(8项)+ 星级解读
    • JBI_Case_Series.md
      -- 病例系列研究的10项关键评价清单
  • 第9阶段:共同作者传阅
    ${CLAUDE_SKILL_DIR}/references/phase9_circulation.md
    -- 线程连续性、附件范围、收件人结构、7天时间窗口
  • 第10阶段:自我审核修复
    ${CLAUDE_SKILL_DIR}/references/phase10_recovery.md
    -- 触发条件、12步重建流程、PROSPERO修改、重新传阅框架
  • 数据完整性清单
    ${CLAUDE_SKILL_DIR}/references/data_integrity_checklist.md
    -- DI-1~DI-9提取/合成防护准则(基于匿名化的元分析项目)
  • 综述编排
    ${CLAUDE_SKILL_DIR}/references/review_orchestration.md
    -- RO-1~RO-5传阅规范(扩展自phase9_circulation.md)
  • 提交包偏差
    ${CLAUDE_SKILL_DIR}/references/submission_package_drift.md
    -- 多期刊文件夹管理、
    DO_NOT_EDIT_HERE
    标识、
    _build.sh
    模式
  • 提交后发布操作
    ${CLAUDE_SKILL_DIR}/references/post_submission_release_ops.md
    -- Zenodo DOI管控、标签清理、拒稿后重新投稿版本控制
  • 实证同行评审经验
    ${CLAUDE_SKILL_DIR}/references/empirical_lessons.md
    -- 16条积累的系统评价-元分析同行评审/投稿经验(2026年5-6月),指导第4阶段提取表单设计、第4c阶段质量控制和第8阶段提交审核。在设计提取表单和提交前务必阅读。

Built-in Templates (
${CLAUDE_SKILL_DIR}/templates/
)

内置模板(
${CLAUDE_SKILL_DIR}/templates/

  • Extraction Form v2 (
    templates/extraction_form_v2.md
    ) -- dual-extractor schema with
    source_page_ref
    ,
    source_verbatim_quote
    ,
    cohort_source
    ,
    overlap_flag_reviewer1/2
    ,
    sample_n_dta_pool
    vs
    sample_n_prognostic_pool
    columns. Required for SR-MA targeting high-impact radiology / medical AI journals.
  • Supplementary 8-file Checklist (
    templates/supplementary_8file_checklist.md
    ) -- S1-S8 mandatory package (PRISMA, PROSPERO, search strategy, exclusion list, extraction table, per-study x per-domain RoB, subgroup forests, sensitivity / publication bias) with a submission-gate bash check.
  • 提取表单v2
    templates/extraction_form_v2.md
    ) -- 双提取者模式,包含
    source_page_ref
    source_verbatim_quote
    cohort_source
    overlap_flag_reviewer1/2
    sample_n_dta_pool
    sample_n_prognostic_pool
    列。适用于目标为高影响力放射学/医学AI期刊的系统评价-元分析。
  • 补充材料8文件清单
    templates/supplementary_8file_checklist.md
    ) -- S1-S8必填文件包(PRISMA、PROSPERO、检索策略、排除列表、提取表格、逐研究逐领域偏倚风险评估、亚组森林图、敏感性/发表偏倚分析),含提交前bash检查。

Built-in Scripts (
${CLAUDE_SKILL_DIR}/scripts/
)

内置脚本(
${CLAUDE_SKILL_DIR}/scripts/

  • screening_reconcile.py
    -- Phase 3f ID-set screening reconciliation.
  • check_pool_consistency.py
    -- pool-composition / PRISMA count consistency.
  • cohort_overlap_check.py
    -- shared-database cohort-overlap detection.
  • extract_assist.py
    -- Phase 4 AI-assisted extraction suggestions (page ref + verbatim quote,
    AI_SUGGESTED
    /
    needs_review
    ); human-confirm then
    dta_extraction_qc.py
    . Challenge card:
    scripts/extract_assist_challenge/
    .
  • dta_extraction_qc.py
    -- 2x2 cell ↔ source sens/spec QC on the confirmed extraction CSV.

  • screening_reconcile.py
    -- 第3f阶段:ID集筛选一致性核对。
  • check_pool_consistency.py
    -- 研究池组成/PRISMA计数一致性检查。
  • cohort_overlap_check.py
    -- 共享数据库队列重叠检测。
  • extract_assist.py
    -- 第4阶段:AI辅助提取建议(页码引用+原文引用,标记
    AI_SUGGESTED
    /
    needs_review
    );人工确认后运行
    dta_extraction_qc.py
    。挑战案例:
    scripts/extract_assist_challenge/
  • dta_extraction_qc.py
    -- 针对已确认的提取CSV,验证2×2单元格与原文报告的灵敏度/特异度是否一致(检测组交换问题)。

Meta-Analysis Types

元分析类型

TypeRoB ToolStatistical ModelReporting Guideline
DTA (diagnostic test accuracy)QUADAS-3 (QUADAS-2 for legacy reviews)Bivariate / HSROCPRISMA-DTA
Intervention (treatment effect)RoB 2 (RCT) / ROBINS-I (NRSI)Random-effects (DL/REML)PRISMA 2020
Prognostic (prediction model)QUIPS / PROBASTRandom-effectsPRISMA 2020
Observational (prevalence/association)NOS / JBIRandom-effectsMOOSE
Auto-detect type from the research question or accept user specification.

类型偏倚风险工具统计模型报告指南
DTA(诊断试验准确性)QUADAS-3(旧版综述使用QUADAS-2)双变量模型 / HSROCPRISMA-DTA
干预研究(治疗效果)RoB 2(随机对照试验)/ ROBINS-I(非随机对照研究)随机效应模型(DL/REML)PRISMA 2020
预后研究(预测模型)QUIPS / PROBAST随机效应模型PRISMA 2020
观察性研究(患病率/关联性)NOS / JBI随机效应模型MOOSE
可根据研究问题自动识别类型,或接受用户指定。

Workflow Phases

工作流程阶段

Phase 1: Protocol Development

第1阶段:方案制定

Goal: Produce a PROSPERO-ready protocol document.
  1. Structure the research question:
    • DTA: PIRD (Population, Index test, Reference standard, Diagnosis)
    • Intervention: PICO (Population, Intervention, Comparator, Outcome)
  2. DTA only — do QUADAS-3 phases 1 and 2 now, not at risk-of-bias time: QUADAS-3's first two phases are review-level and belong in the protocol: phase 1 states the synthesis question(s) (population, index test(s), target condition — a review may have more than one), and phase 2 defines the ideal test accuracy trial for each: objective, participants, index test(s), definition of the target condition, analysis. Every later risk-of-bias and applicability judgement is made against that trial. Write the review-specific guidance for answering each signalling question here too, with clinical and methodological input, and publish it as a web appendix. Defining the ideal trial after seeing the studies is not an assessment — it is a judgement fitted to the results. See
    references/checklists/QUADAS3.md
    .
  3. Define eligibility criteria:
    • Study design (cross-sectional DTA, cohort, RCT, etc.)
    • Population characteristics
    • Index test / intervention specifics
    • Comparator / reference standard
    • Outcome measures (Se/Sp for DTA; effect size for intervention)
    • Exclusion criteria with justification
  4. Plan the search:
    • Minimum 3 databases: PubMed, Embase, and Cochrane CENTRAL (add Scopus, Web of Science as needed)
    • Draft Boolean search strategy using PIRD/PICO components
    • Grey literature plan (conference abstracts, trial registries)
    • Language restrictions (state explicitly)
    • Date range with justification
  5. Plan RoB assessment:
    • Select tool based on type (see table above)
    • State number of independent assessors (minimum 2)
    • Plan for disagreement resolution (consensus, third reviewer)
  6. Plan synthesis:
    • DTA: bivariate random-effects model (Reitsma) or HSROC (Rutter & Gatsonis)
    • Intervention: random-effects (DerSimonian-Laird or REML)
    • Heterogeneity assessment plan
    • Subgroup / sensitivity analysis plan
    • Publication bias assessment plan
  7. Generate PROSPERO registration document:
    • Read
      ${CLAUDE_SKILL_DIR}/references/PROSPERO_template.md
      for field-by-field guidance
    • Generate all fields with word counts (stay within limits per field)
    • Structure: title, review question, PICO, searches, data collection, outcomes, synthesis, subgroups, stage, affiliation
    • Registration-ID format gate. A PROSPERO ID is
      CRD42
      + 9 digits (14 characters total), e.g.
      CRD42024500001
      . Validate any ID that appears in the manuscript or registration doc with
      grep -oE 'CRD42[0-9]+'
      and assert a 14-character length /
      ^CRD42\d{9}$
      — a 15-character ID (a stray digit) is a transcription error a reviewer will check against the live record.
    • Review-type selection. Pick the least-wrong portal review type for the actual design and state any portal constraint in the protocol. A descriptive single-arm proportion synthesis is not an "Intervention review"; choosing "Intervention review" only to satisfy a portal field contradicts a later GRADE / effect-certainty statement. Whatever certainty language the protocol commits to (GRADE vs "evidence statements only") must match the manuscript verbatim — a guideline-style "we recommend" is not licensed by a descriptive review type.
    • For mixed designs (comparative + single-arm): explicitly address comparator for both arms
    • For RoB: map tool to study design (NOS for comparative, JBI for case series → select "Other" in form)
    • Output: Markdown + DOCX (via pandoc) for copy-paste into PROSPERO web form
    • Append Common Pitfalls Checklist (HTML entities, word limits, stage constraint)
    • Save to project
      7_Submission/
      or equivalent directory
目标:生成可用于PROSPERO注册的方案文档。
  1. 构建研究问题
    • DTA:PIRD(人群、待评价试验、金标准、诊断目标)
    • 干预研究:PICO(人群、干预措施、对照措施、结局指标)
  2. 仅DTA研究 — 现在完成QUADAS-3的第1和第2阶段,而非偏倚风险评估阶段: QUADAS-3的前两个阶段属于综述层面,应纳入方案: 第1阶段明确合成问题(人群、待评价试验、目标疾病 — 一项综述可能包含多个问题),第2阶段为每个问题定义理想的诊断试验准确性研究:目标、受试者、待评价试验、目标疾病定义、分析方法。后续所有偏倚风险和适用性判断均以此为参照。 在此处撰写针对每个信号问题的综述专属解答指南,结合临床方法学意见,并作为网络附录发布。 在看到研究后再定义理想研究并非评估 — 而是根据结果调整判断。详见
    references/checklists/QUADAS3.md
  3. 确定纳入排除标准
    • 研究设计(横断面DTA研究、队列研究、随机对照试验等)
    • 人群特征
    • 待评价试验/干预措施细节
    • 对照措施/金标准
    • 结局指标(DTA为灵敏度/特异度;干预研究为效应量)
    • 排除标准及理由
  4. 规划检索策略
    • 至少检索3个数据库:PubMed、Embase和Cochrane CENTRAL(必要时添加Scopus、Web of Science)
    • 基于PIRD/PICO组件起草布尔检索策略
    • 灰色文献检索计划(会议摘要、试验注册库)
    • 语言限制(明确说明)
    • 时间范围及理由
  5. 规划偏倚风险评估
    • 根据研究类型选择工具(见上表)
    • 说明独立评价者数量(至少2名)
    • 规划分歧解决方式(共识、第三方评价者)
  6. 规划合成分析
    • DTA:双变量随机效应模型(Reitsma)或HSROC(Rutter & Gatsonis)
    • 干预研究:随机效应模型(DerSimonian-Laird或REML)
    • 异质性评估计划
    • 亚组/敏感性分析计划
    • 发表偏倚评估计划
  7. 生成PROSPERO注册文档
    • 阅读
      ${CLAUDE_SKILL_DIR}/references/PROSPERO_template.md
      获取字段指南
    • 生成所有字段并标注字数(严格遵守各字段字数限制)
    • 结构:标题、研究问题、PICO、检索策略、数据收集、结局指标、合成分析、亚组分析、阶段、机构
    • 注册ID格式审核:PROSPERO ID格式为
      CRD42
      + 9位数字(共14个字符),例如
      CRD42024500001
      。使用
      grep -oE 'CRD42[0-9]+'
      验证手稿或注册文档中的任何ID,并确认其长度为14字符 / 符合
      ^CRD42\d{9}$
      格式 — 15字符的ID(多一个数字)属于转录错误,审稿人会对照实时记录检查。
    • 综述类型选择:根据实际设计选择最接近的门户综述类型,并在方案中说明门户限制。描述性单组比例合成不属于“干预综述”;仅为满足门户字段而选择“干预综述”与后续GRADE/证据确定性声明矛盾。方案中承诺的任何确定性表述(GRADE vs “仅证据声明”)必须与手稿完全一致 — 指南式的“我们推荐”不适用于描述性综述类型。
    • 混合设计(比较研究+单组研究):明确说明两组的对照措施
    • 偏倚风险:将工具与研究设计对应(比较研究用NOS,病例系列研究用JBI → 在表单中选择“其他”)
    • 输出:Markdown + DOCX(通过pandoc转换),用于复制粘贴到PROSPERO网页表单
    • 附加常见问题清单(HTML实体、字数限制、阶段约束)
    • 保存至项目
      7_Submission/
      或等效目录

Phase 2: Search Strategy

第2阶段:检索策略

Goal: Develop and validate reproducible search strategies.
  1. Build search blocks from PIRD/PICO:
    • Population block (MeSH + free text)
    • Index test / Intervention block
    • Comparator / Reference standard block (optional)
    • Study design filter (if applicable)
  2. Combine with Boolean operators:
    • Within blocks: OR
    • Between blocks: AND
  3. Execute search per database using
    /search-lit
    :
    • PubMed: MeSH + free text
    • Embase: Emtree + free text
    • Additional databases as specified in protocol
  4. Report search per PRISMA-S (Rethlefsen et al. 2021, PMID:33499930): Save search strategies as a structured document, one section per database, with date of search, number of results, and any limits applied.
  5. Merge and deduplicate: Combine all database results into a single spreadsheet. Deduplicate by DOI first, then PMID. Save raw counts for PRISMA flow.
目标:制定并验证可重复的检索策略。
  1. 基于PIRD/PICO构建检索模块
    • 人群模块(MeSH + 自由词)
    • 待评价试验/干预措施模块
    • 对照措施/金标准模块(可选)
    • 研究设计筛选器(如适用)
  2. 使用布尔运算符组合
    • 模块内:OR
    • 模块间:AND
  3. 使用
    /search-lit
    在各数据库执行检索
    • PubMed:MeSH + 自由词
    • Embase:Emtree + 自由词
    • 按方案指定的其他数据库
  4. 按照PRISMA-S报告检索情况(Rethlefsen等,2021,PMID:33499930): 将检索策略保存为结构化文档,每个数据库单独成节,包含检索日期、结果数量及应用的任何限制。
  5. 合并与去重:将所有数据库结果合并到单个电子表格中。 优先按DOI去重,再按PMID去重。保存原始计数用于PRISMA流程图。

Phase 3: Screening & Selection

第3阶段:文献筛选与选择

Goal: Systematic title/abstract and full-text screening with two independent reviewers.
3a. Round 1 — initial title/abstract screening (single reviewer). Define the exclusion codes from the protocol (E1=Not target population, E2=Not intervention, E3=Ineligible type, E4=Non-human, E5=Duplicate). Mark every record INCLUDE / EXCLUDE / MAYBE with a reason code →
round1_{date}.tsv
.
3b. Round 2 — dual independent title/abstract screening. A second independent reviewer (or AI as a documented second-pass tool with human verification) re-screens all R1 records. Compute Cohen's κ and report it in Methods.
round2_tag
= INCLUDE / EXCLUDE / MAYBE, where MAYBE means disagreement or either reviewer flagged uncertainty →
round2_tag
,
round2_reason
columns.
3c. Round 3 — adjudication of disagreements (first reviewer). Build the R3 sheet with all MAYBE records first, then INCLUDE records for a brief confirmation pass. The first reviewer independently adjudicates each row (
round3_decision
, plus
round3_reason
only when overturning R2). Optional AI-assisted pre-screening can compress the effort — but AI suggestions are not decisions: the reviewer independently confirms or overturns every one. Template, sort priority, and the required Methods boilerplate are in the reference file.
3d. Round 4 — full-text screening. Retrieve full texts for
round3_decision = INCLUDE
(use
/fulltext-retrieval
), apply the full-text exclusion codes (F1=No extractable outcome, F2=No comparative data, F3=Cannot separate target population, F4=Inadequate sample/follow-up, F5=Full-text unavailable), with two independent reviewers, Cohen's κ, and consensus or a third reviewer for disagreements. Flag comparative studies for priority extraction.
3e. PRISMA flow. Track counts at every stage (R1 → R2 → R3 → R4 → final included); generate the diagram with
/make-figures
once the numbers are final.
3f. Post-consensus count reconciliation gate (MANDATORY before Phase 5 write-up). Reconcile the counts from the raw ID sets, never from prose summaries, and record the canonical totals in one source-of-truth file:
bash
python "${CLAUDE_SKILL_DIR}/scripts/screening_reconcile.py" \
  --screening 2_Screening/fulltext_screening.tsv \
  --consensus 2_Screening/consensus_decisions.tsv \
  --table1 6_Tables/table1_studies.csv \
  --output 2_Screening/screening_consensus.json
Downstream stages consume
screening_consensus.json
for counts and ID sets; the Markdown consensus document remains the human explanation. Three hard rules:
  1. List the narrative-only IDs explicitly. The highest-yield red flag is a numeric claim ("10 narrative-only studies") that does not match the enumerable set
    (A ∪ C) \ B \ T
    .
  2. No "N → M" transition without ID receipts. "k rose from 30 to 32 after FLAG consensus" must cite the added/removed IDs. A transition claim with no enumerable ID set is a P0 and blocks the Phase 5 hand-off.
  3. STAGE_TRANSFER_LOSS
    is a P0.
    Exit 1 when a record is included at screening but absent from the consensus artifact altogether — no adjudication was ever recorded. An exclusion is a decision; silence is a gap. Never let it settle into narrative-only (why: reference file).
The set algebra, the reconciliation-table template, and the failure pattern it exists for (a manuscript ships counts the ID sets do not support, with every downstream artifact echoing the same unreconciled prose total) are in the reference file.
3f.5 Pool composition lock (MANDATORY at adjudication freeze). Once 3f passes, freeze the pool into a single source-of-truth YAML that every downstream artifact can be checked against:
bash
cp "${CLAUDE_SKILL_DIR}/templates/FINAL_POOL_LOCK.yaml.template" 2_Data/FINAL_POOL_LOCK.yaml
目标:由两名独立评价者完成系统的标题/摘要和全文筛选。
3a. 第1轮 — 初始标题/摘要筛选(单评价者)。根据方案定义排除代码(E1=非目标人群,E2=非干预措施,E3=不符合研究类型,E4=非人类研究,E5=重复文献)。为每条记录标记INCLUDE / EXCLUDE / MAYBE并标注理由代码 → 保存为
round1_{date}.tsv
3b. 第2轮 — 双独立标题/摘要筛选。第二名独立评价者(或AI作为有记录的二次筛选工具,需人工验证)重新筛选所有第1轮记录。计算Cohen's κ并在方法部分报告。
round2_tag
= INCLUDE / EXCLUDE / MAYBE,其中MAYBE表示存在分歧 任一评价者标记不确定 → 添加
round2_tag
round2_reason
列。
3c. 第3轮 — 分歧裁决(第一名评价者)。先构建包含所有MAYBE记录的第3轮表格,再对INCLUDE记录进行简短确认。第一名评价者独立裁决每一行(
round3_decision
,仅在推翻第2轮决定时添加
round3_reason
)。可选AI辅助预筛选可减少工作量 — 但 AI建议并非最终决定:评价者需独立确认或推翻每一条建议。模板、排序优先级及所需方法学模板见参考文件。
3d. 第4轮 — 全文筛选。获取
round3_decision = INCLUDE
记录的全文(使用
/fulltext-retrieval
),应用全文排除代码(F1=无可用结局指标,F2=无比较数据,F3=无法区分目标人群,F4=样本量/随访不足,F5=全文无法获取),由两名独立评价者完成,计算Cohen's κ,通过共识或第三方评价者解决分歧。标记比较研究优先提取数据。
3e. PRISMA流程图。跟踪每个阶段的计数(第1轮 → 第2轮 → 第3轮 → 第4轮 → 最终纳入);计数确定后使用
/make-figures
生成流程图。
3f. 共识后计数一致性核对(第5阶段撰写前必须完成)。根据原始ID集而非文字摘要核对计数,并将标准总数记录在单一可信源文件中:
bash
python "${CLAUDE_SKILL_DIR}/scripts/screening_reconcile.py" \
  --screening 2_Screening/fulltext_screening.tsv \
  --consensus 2_Screening/consensus_decisions.tsv \
  --table1 6_Tables/table1_studies.csv \
  --output 2_Screening/screening_consensus.json
下游阶段使用
screening_consensus.json
获取计数和ID集;Markdown共识文档作为人工说明保留。三项硬性规则:
  1. 明确列出仅文字描述的ID。最高风险的红色预警是数值声明(“10项仅文字描述的研究”)与可枚举集合
    (A ∪ C) \ B \ T
    不匹配。
  2. 无ID记录不得进行“N → M”转换。“经过FLAG共识后k从30增加到32”必须引用添加/移除的ID。无枚举ID集的转换声明属于P0级问题,会阻止进入第5阶段。
  3. STAGE_TRANSFER_LOSS
    属于P0级问题
    。当某条记录在筛选阶段被纳入但完全未出现在共识文件中时,退出程序 — 未记录任何裁决。排除是一种决定;沉默是漏洞。绝不能让其仅以文字描述存在(原因见参考文件)。
集合运算、一致性核对表模板及针对的失败模式(手稿中的计数与ID集不符,所有下游文件重复相同的未核对文字总数)见参考文件。
3f.5 研究池组成锁定(裁决完成后必须完成)。第3f阶段通过后,将研究池冻结为单一可信源YAML文件,供所有下游文件核对:
bash
cp "${CLAUDE_SKILL_DIR}/templates/FINAL_POOL_LOCK.yaml.template" 2_Data/FINAL_POOL_LOCK.yaml

fill counts + UID lists from 3f, compute the SHA-256 over the sorted UID list,

从3f阶段填充计数 + UID列表,计算排序后UID列表的SHA-256值,

and COMMIT THE LOCK before any Phase 4 extraction

并在第4阶段提取前提交LOCK文件


- **Never re-derive `k included` from the extraction TSV at manuscript build time** — always
  reference `final_pool_n` from the lock.
- **Aggregate patient/lesion totals are locked too**, not just study counts. Distinguish
  **arm-separable** from **both-arm** rows: a study contributing one arm must not have its
  full-cohort count folded into a pooled total. A hand-carried headline total that does not
  re-derive from the locked per-study values is a **P0**.
- A late post-freeze change to the pool is a **formal PROSPERO amendment**: file it, re-freeze as
  `FINAL_POOL_LOCK_v2.yaml`, and propagate to every artifact.

**Read on demand:**

| File | Read it when | Cost if read blindly |
|---|---|---|
| `references/phase3_screening_detail.md` | you are executing a screening round, using AI pre-screening, or a reconciliation/lock gate fired | ~3,600 tokens; the round procedures are needed one round at a time, not all at invocation |

- **绝不在手稿构建时从提取TSV重新推导`纳入研究数k`** — 始终引用锁定文件中的`final_pool_n`。
- **患者/病灶汇总总数也需锁定**,而非仅研究计数。区分**可分离组**与**两组合并**行:仅贡献一组的研究不得将其全队列计数纳入合并总数。未从锁定的单研究值重新推导的手动汇总总数属于**P0级问题**。
- 冻结后若需更改研究池,需进行**正式PROSPERO修改**:提交修改申请,重新冻结为`FINAL_POOL_LOCK_v2.yaml`,并同步至所有文件。

**按需阅读**:

| 文件 | 阅读时机 | 盲目阅读的代价 |
|---|---|---|
| `references/phase3_screening_detail.md` | 执行筛选轮次、使用AI预筛选、或一致性/锁定审核触发时 | ~3600 tokens;仅在执行对应轮次时需要轮次流程,无需在调用时全部阅读 |

Phase 4: Data Extraction

第4阶段:数据提取

Goal: Create standardized extraction forms and extract 2x2 or effect-size data.
4.0 Entry gate (MANDATORY) — pool composition lock ↔ adjudication TSV. Before any extraction work begins, confirm the round-3 adjudication TSV and
FINAL_POOL_LOCK.yaml
(Phase 3f.5) agree on which UIDs are included:
bash
python "${CLAUDE_SKILL_DIR}/scripts/check_pool_consistency.py" \
    --lock 2_Data/FINAL_POOL_LOCK.yaml \
    --adjudication-tsv 2_Screening/round3_adjudication.tsv \
    --decision-col round3_decision --uid-col uid \
    --include-labels "INCLUDE,INCLUDE_MIXED" \
    --out qc/pool_consistency.json
The gate fails closed: any UID disagreement blocks extraction. Resolve by re-freezing the lock with the corrected UID set (and propagating downstream) or by correcting a mis-labelled TSV row. Do NOT proceed with a mismatch — the extraction matrix will not align with the locked pool, and the drift surfaces as a fabrication-grade red flag at peer review.
Failure-mode cross-ref
references/data_integrity_checklist.md
DI-1~DI-5 are mandatory during extraction (2x2 arm-swap, KM audit trail, methodology mismatch, PRISMA 5-way drift, single-source k).
Extraction form. For an SR-MA targeting high-impact radiology / medical AI journals use
${CLAUDE_SKILL_DIR}/templates/extraction_form_v2.md
— its dual-extractor, source-page-reference, and verbatim-quote columns are what close the 2x2 cell-swap and cohort-overlap blind spots. The DTA and intervention field lists are in the reference file.
AI-drafted starting document — treat as hallucination-suspect. If a mentor or collaborator shared an AI-drafted study list, 2x2 set, or effect estimates (even flagged "for reference only"): save it with a
_DO_NOT_USE_VERBATIM
suffix and re-verify every N, denominator, event count, OR/CI, and author/year against the source PDF. Trust hierarchy: source PDF + own analysis stdout > the mentor's direct text > the attached AI draft — never promote a draft up that ladder. Procedure and precedent: reference file.
4b. Special cases (KM reconstruction, composite exposure). When studies report outcomes only as Kaplan-Meier curves, or the intervention is a composite of techniques, load
${CLAUDE_SKILL_DIR}/references/phase4_km_composite.md
for the WebPlotDigitizer →
IPDfromKM
procedure (cite Guyot et al. 2012, doi:10.1186/1471-2288-12-9) and the 4-path composite-exposure decision tree. Pre-specify a sensitivity analysis excluding composite-exposure studies.
Cross-verification (≥2 independent reviewers). Report inter-reviewer agreement (% or Cohen's κ) at title/abstract and full-text stages. Verify denominator consistency — the denominator may differ across outcomes within one study, so for each outcome back-calculate
event ÷ denominator
and confirm it reproduces the paper's reported percentage. Distinguish KM-curve estimates from raw event counts and record the data source (Table / KM / text). Log every consensus decision in
{project}/consensus_log.md
, then lock the dataset; later changes need a dated justification.
4c. Extraction QC & cohort overlap. After dual-extractor consensus, run both before locking:
bash
undefined
目标:创建标准化提取表单,提取2×2表格或效应量数据。
4.0 入口审核(必须完成) — 研究池组成锁定 ↔ 裁决TSV。开始任何提取工作前,确认第3轮裁决TSV与
FINAL_POOL_LOCK.yaml
(第3f.5阶段)在纳入UID上一致:
bash
python "${CLAUDE_SKILL_DIR}/scripts/check_pool_consistency.py" \
    --lock 2_Data/FINAL_POOL_LOCK.yaml \
    --adjudication-tsv 2_Screening/round3_adjudication.tsv \
    --decision-col round3_decision --uid-col uid \
    --include-labels "INCLUDE,INCLUDE_MIXED" \
    --out qc/pool_consistency.json
审核不通过则无法进入下一阶段:任何UID不一致都会阻止提取。通过重新冻结锁定文件(并同步至下游)或修正TSV中的错误标签来解决。切勿在存在不匹配的情况下继续 — 提取矩阵将与锁定研究池不一致,在同行评审时会被视为伪造级别的红色预警。
失败模式交叉引用
references/data_integrity_checklist.md
中的DI-1~DI-5在提取阶段必须遵守(2×2组交换、KM审计追踪、方法学不匹配、PRISMA五重偏差、单源k值)。
提取表单。针对目标为高影响力放射学/医学AI期刊的系统评价-元分析,使用
${CLAUDE_SKILL_DIR}/templates/extraction_form_v2.md
— 其双提取者、原文页码引用和原文引用列可消除2×2单元格交换和队列重叠的盲点。DTA和干预研究的字段列表见参考文件。
AI起草的初始文档 — 视为可能存在幻觉。如果导师或合作者分享了AI起草的研究列表、2×2表格或效应估计值(即使标记为“仅供参考”):保存时添加
_DO_NOT_USE_VERBATIM
后缀,并逐一对照原文PDF验证所有N值、分母、事件数、OR/CI及作者/年份。信任层级:原文PDF + 自身分析输出 > 导师直接文本 > 附带的AI草稿 — 绝不能提升草稿的信任层级。流程及先例见参考文件。
4b. 特殊情况(KM曲线重建、复合暴露)。当研究仅以Kaplan-Meier曲线报告结局,或干预措施为多种技术的复合时,阅读
${CLAUDE_SKILL_DIR}/references/phase4_km_composite.md
获取WebPlotDigitizer →
IPDfromKM
流程(引用Guyot等,2012,doi:10.1186/1471-2288-12-9)和4路径复合暴露决策树。预先设定排除复合暴露研究的敏感性分析。
交叉验证(≥2名独立评价者)。报告标题/摘要和全文阶段的评价者间一致性(百分比或Cohen's κ)。验证分母一致性 — 同一研究中不同结局的分母可能不同,因此需针对每个结局反向计算
事件数 ÷ 分母
,并确认其与论文报告的百分比一致。区分KM曲线估计值与原始事件数,并记录数据源(表格/KM曲线/文本)。将所有共识决策记录在
{project}/consensus_log.md
中,然后锁定数据集;后续更改需标注日期和理由。
4c. 提取质量控制 & 队列重叠检查。双提取者达成共识后,在锁定前运行以下两项检查:
bash
undefined

2x2 cell integrity: validates TP/FN/TN/FP against source-reported sens/spec (catches arm-swap)

2×2单元格完整性:验证TP/FN/TN/FP与原文报告的灵敏度/特异度是否一致(检测组交换)

python3 "${CLAUDE_SKILL_DIR}/scripts/dta_extraction_qc.py"
--input 2_Extraction/extraction.csv --tolerance 0.02
--out 2_Extraction/qc/dta_extraction_qc.tsv
python3 "${CLAUDE_SKILL_DIR}/scripts/dta_extraction_qc.py"
--input 2_Extraction/extraction.csv --tolerance 0.02
--out 2_Extraction/qc/dta_extraction_qc.tsv

cohort overlap: shared public DB / same institution+period / same first author ±2y

队列重叠:共享公共数据库 / 同一机构+时期 / 同一第一作者±2年

python3 "${CLAUDE_SKILL_DIR}/scripts/cohort_overlap_check.py"
--input 2_Extraction/studies.csv --enrich
--out 2_Extraction/qc/cohort_overlap.md

Any `FLAG_SWAP` / `FLAG_MISMATCH` requires third-reviewer adjudication before Phase 6. **A
confirmed flag is not resolved until the extraction form itself is edited** — a flag corrected only
in a review note silently re-enters synthesis, so re-run the QC and confirm zero open flags before
locking. HIGH-confidence overlap pairs require a Limitations acknowledgment plus a sensitivity
analysis excluding one of the pair. Cross-links: `/peer-review` Phase 2A P1 + P2.

**Read on demand:**

| File | Read it when | Cost if read blindly |
|---|---|---|
| `references/phase4_extraction_detail.md` | building the extraction form, an AI draft was shared, you want the optional `extract_assist.py` scaffolding, or a QC flag fired | ~4,700 tokens; a clean dual-extraction with no AI draft needs none of it |
| `references/phase4_km_composite.md` | studies report only KM curves, or the exposure is composite | ~2,200 tokens |
python3 "${CLAUDE_SKILL_DIR}/scripts/cohort_overlap_check.py"
--input 2_Extraction/studies.csv --enrich
--out 2_Extraction/qc/cohort_overlap.md

任何`FLAG_SWAP` / `FLAG_MISMATCH`需经第三方评价者裁决后才能进入第6阶段。**确认的标记需修改提取表单本身才算解决** — 仅在评审说明中修正标记会导致错误进入合成分析,因此需重新运行质量控制并确认无未解决标记后再锁定。高置信度重叠对需在局限性中说明,并进行排除其中一项的敏感性分析。交叉链接:`/peer-review`第2A阶段P1 + P2。

**按需阅读**:

| 文件 | 阅读时机 | 盲目阅读的代价 |
|---|---|---|
| `references/phase4_extraction_detail.md` | 构建提取表单、收到AI草稿、需要`extract_assist.py`脚手架、或质量控制标记触发时 | ~4700 tokens;无AI草稿的干净双提取无需阅读 |
| `references/phase4_km_composite.md` | 研究仅报告KM曲线,或暴露为复合暴露时 | ~2200 tokens |

Phase 5: Risk of Bias Assessment

第5阶段:偏倚风险评估

Goal: Guide structured RoB assessment with the appropriate tool.
DTA: this phase runs QUADAS-3 phases 3–6 (flow diagram, identify the estimates to assess, assess, overall judgement). Phases 1–2 — the synthesis question and the ideal test accuracy trial — were written in Phase 1 above. If they were not, stop and write them before judging anything; they are the comparator every judgement is made against.
Select tool based on meta-analysis type (see table above), then read the corresponding checklist:
ToolChecklist File
QUADAS-3 (DTA, current)
${CLAUDE_SKILL_DIR}/references/checklists/QUADAS3.md
QUADAS-2 (DTA, legacy)
${CLAUDE_SKILL_DIR}/references/checklists/QUADAS2.md
RoB 2 (RCT)
${CLAUDE_SKILL_DIR}/references/checklists/RoB2.md
ROBINS-I (NRSI)
${CLAUDE_SKILL_DIR}/references/checklists/ROBINS_I.md
PROBAST (Prediction)
${CLAUDE_SKILL_DIR}/references/checklists/PROBAST.md
NOS (Observational)
${CLAUDE_SKILL_DIR}/references/checklists/NOS.md
JBI (Case Series)
${CLAUDE_SKILL_DIR}/references/checklists/JBI_Case_Series.md
For AI/ML prediction models, also apply PROBAST+AI extensions.
Output: Summary table + traffic light plot (use
/make-figures
).
目标:使用合适的工具指导结构化偏倚风险评估。
DTA:本阶段执行QUADAS-3的第3–6阶段(流程图、确定需评估的估计值、评估、总体判断)。第1–2阶段 — 合成问题和理想诊断试验准确性研究 — 已在上述第1阶段完成。若未完成,停止并先撰写这些内容;它们是所有判断的参照标准。
根据元分析类型选择工具(见上表),然后阅读对应的检查清单:
工具检查清单文件
QUADAS-3(DTA,当前版本)
${CLAUDE_SKILL_DIR}/references/checklists/QUADAS3.md
QUADAS-2(DTA,旧版)
${CLAUDE_SKILL_DIR}/references/checklists/QUADAS2.md
RoB 2(随机对照试验)
${CLAUDE_SKILL_DIR}/references/checklists/RoB2.md
ROBINS-I(非随机对照研究)
${CLAUDE_SKILL_DIR}/references/checklists/ROBINS_I.md
PROBAST(预测模型)
${CLAUDE_SKILL_DIR}/references/checklists/PROBAST.md
NOS(观察性研究)
${CLAUDE_SKILL_DIR}/references/checklists/NOS.md
JBI(病例系列研究)
${CLAUDE_SKILL_DIR}/references/checklists/JBI_Case_Series.md
对于AI/ML预测模型,还需应用PROBAST+AI扩展。
输出:汇总表 + 红绿灯图(使用
/make-figures
)。

Phase 6: Statistical Synthesis

第6阶段:统计合成分析

Goal: Execute meta-analysis and generate publication-ready outputs.
Failure-mode cross-ref
references/data_integrity_checklist.md
DI-6/DI-7/DI-9 are the consistency gate (CSV ↔ script ↔ prose; single-source k; 3-way numeric reconciliation before Stage 4).
IMPORTANT: Always use R for meta-analysis (packages:
meta
,
metafor
,
mada
). See
${CLAUDE_SKILL_DIR}/references/r_templates.md
for full code templates.
Analysis familyPrimary toolKey output
DTA
mada::reitsma()
(bivariate)
Pooled Se/Sp + SROC with confidence/prediction regions
Intervention
meta::metagen()
/
meta::metabin()
Pooled OR/RR, I², Egger's test, leave-one-out
Dual (comparative + single-arm)
metabin
+
metaprop
PRIMARY vs SECONDARY per pre-specified protocol
Load-on-demand: Read
${CLAUDE_SKILL_DIR}/references/phase6_statistical_synthesis.md
for the full R code templates, the dual-approach decision table (comparative vs single-arm), practical cautions (method.tau, HK CI, zero-cell correction), publication-bias test power, sensitivity-analysis menu, and error-handling rules.
Three checks before the pool is written up — each is a Methods sentence, not only a setting. R and detail in the same reference:
  1. Is the event rare? A pooled event rate < 1%, or any zero-event arm, moves the analysis off the inverse-variance default onto Peto / Mantel-Haenszel without a zero-cell correction / GLMM. Inverse-variance methods including DerSimonian-Laird are to be avoided for rare events, and so are 0.5 continuity corrections with them.
  2. Why this model? Fixed vs random is a judgment about whether one common true effect exists — never derived from Cochran's Q or I². "A random-effects model was used because I² was 65%" is a reviewer catch, not a rationale.
  3. Does one study contribute several correlated effect sizes? Multiple outcomes, readers, thresholds, or time points from the same participants need one pre-specified estimate per study, a multivariate model, or robust variance estimation — not independent pooling.
目标:执行元分析并生成可发表的输出结果。
失败模式交叉引用
references/data_integrity_checklist.md
中的DI-6/DI-7/DI-9为一致性审核(CSV ↔ 脚本 ↔ 文字;单源k值;第4阶段前的三重数值一致性核对)。
重要提示:元分析始终使用R语言(包:
meta
metafor
mada
)。 详见
${CLAUDE_SKILL_DIR}/references/r_templates.md
获取完整代码模板。
分析类别主要工具关键输出
DTA
mada::reitsma()
(双变量模型)
合并灵敏度/特异度 + 带置信区间/预测区间的SROC曲线
干预研究
meta::metagen()
/
meta::metabin()
合并OR/RR、I²、Egger检验、逐一排除分析
混合研究(比较研究+单组研究)
metabin
+
metaprop
按预先指定的方案区分主要分析 vs 次要分析
按需加载:阅读
${CLAUDE_SKILL_DIR}/references/phase6_statistical_synthesis.md
获取完整R代码模板、双方法决策表(比较研究vs单组研究)、实用注意事项(method.tau、HK置信区间、零单元格校正)、发表偏倚检验效能、敏感性分析选项及错误处理规则。
写入前的三项检查 — 每项均需在方法部分说明,而非仅设置参数。详情见同一参考文件:
  1. 事件是否罕见? 合并事件率<1%,或存在零事件组时,需将分析从逆方差默认方法切换为Peto / Mantel-Haenszel法,且不进行零单元格校正/GLMM。包括DerSimonian-Laird在内的逆方差方法应避免用于罕见事件,同时也应避免使用0.5连续性校正。
  2. 为何选择该模型? 固定效应vs随机效应是关于是否存在共同真实效应的判断 — 绝不能从Cochran's Q或I²推导。“因I²为65%而使用随机效应模型”会被审稿人抓住,并非合理依据。
  3. 一项研究是否贡献多个相关效应量? 同一受试者的多个结局、评价者、阈值或时间点需预先指定每项研究的一个估计值、使用多变量模型或稳健方差估计 — 而非独立合并。

Phase 6b: Post-Analysis Source Fidelity Audit (MANDATORY)

第6b阶段:分析后原文保真度审核(必须完成)

Goal: Catch numerical hallucinations that survived the forward pipeline (CSV → .R → manuscript).
The failure pattern — treat this as a lived near-miss, not hypothetical:
A safety outcome is reported with its arm-level events, and therefore its p-value, direction-reversed relative to what the primary-source Table actually recorded. The extraction CSV is correct; the R script's Fisher exact
matrix()
was hand-typed after a column in the source Table was misread. Internal consistency checks passed because every downstream artifact (Abstract, Discussion, Table, forest caption) echoed the same wrong number. The reversal was caught only on a second-pass audit with random extraction sampling against the primary paper.
Non-negotiable rules:
  1. No hand-typed numerical matrices when a CSV exists.
    • Use
      read.csv(...)
      + subset / filter. Never copy a 2x2 table from a paper's Table into
      matrix(c(...), ...)
      by eye.
    • If hand entry is truly unavoidable (e.g., text-only extraction), the
      matrix
      ,
      c()
      , or
      data.frame
      line MUST carry a comment citing the exact CSV row + column OR the exact primary-source Table/Page coordinate. Example:
      r
      # source: data_extraction_final.csv row <N> (<first-author> <year>), cols <event_arm1>=0, <event_arm2>=1
      # verified against primary source Table <X>, page <P>
      fisher.test(matrix(c(0, 45, 1, 55), nrow = 2, byrow = FALSE))
  2. Comparative-arm subsets are a separate consensus-log row.
    • When one study's arm-specific values (e.g., one arm of a multi-arm study) are used in a comparative analysis while the full cohort of that study appears elsewhere,
      extraction_consensus_log.md
      must carry an explicit row for the arm-specific values. Pooled totals and arm-specific values MUST NOT share a row.
  3. Random 3-claim back-check before closing Phase 6.
    • After the forest/funnel/subgroup outputs stabilize, randomly sample 3 numerical claims from the Results section of the draft manuscript and trace each back to (a) the R output log and (b) the original paper's Table/Figure.
    • Record the back-check as a small table in
      peer_review_<vN>_internal.md
      :
      Claim (manuscript line)R output file:linePrimary source (paper, Table/Fig, page)Match?
    • A single mismatch is a P0 blocker — do not advance to Phase 7 until resolved.
  4. Revision-introduced numbers must be tagged.
    • Any new number added after v1 — including numbers produced by a new comparative / subgroup / sensitivity script — MUST be wrapped inline as
      [VERIFY-CSV]
      in the manuscript until the Phase 2.5a audit in
      /self-review
      clears it.
  5. Sensitivity analyses must be recomputed on the modified data, not copied.
    • When you add a sensitivity / leave-one-out / erosion / alternative-model analysis, every reported effect size (Cohen's dz/f, AUC, OR, HR, β, sens/spec, ICC) MUST be re-derived from the modified dataset. If a sensitivity-table effect size is identical to the primary analysis to two decimals across ≥4 values, the recomputation almost certainly did not run and the primary values were transcribed — re-run the script on the modified data.
    • The underlying means/SDs/counts will change even when the effect size looks similar; if the effect sizes are byte-identical while the inputs differ, that is the tell. Probability of ≥4 independent values coinciding to 2 decimals by chance is ≈ (0.01)^4 — essentially zero.
    • The failure it catches: a sensitivity analysis reports a block of effect-size values byte-identical to the primary tables while the underlying means/SDs differ — the sensitivity analysis was never actually recomputed. Internal consistency cannot see it.
  6. A "fixed" / "resolved" audit note requires re-run evidence, not a claim.
    • When a prior audit note records a number as
      fixed
      ,
      resolved
      , or
      corrected
      , that status is only valid if it carries the re-run evidence: a timestamp and the relevant stdout / output-file line showing the corrected value, or the commit that changed it. A bare "fixed in v10" with no re-run artifact does NOT clear the finding — re-run the script and attach the output.
    • The forward pipeline can echo a stale value through every artifact while an audit note claims it was fixed (e.g., a major-comparison N still reading the old total after a "fixed" note). The outcome-denominator cross-check (
      /self-review
      Phase 2.5b, the cohort-arithmetic / pool-lock assertions) must pass against the current outputs before any "fixed" status is accepted.
When this phase triggers: every time Phase 6 outputs change (first draft, revision, reviewer- requested re-analysis). Not optional on "minor" re-runs — the precedent reversal above occurred inside a "minor" revision-era re-analysis.
目标:发现正向流程(CSV → .R → 手稿)中未被发现的数值幻觉。
失败模式 — 视为实际发生的近失误,而非假设:
某安全性结局报告了组级事件数,因此其p值与原文表格记录的方向相反。 提取CSV是正确的;R脚本中的Fisher精确检验
matrix()
是手动输入的,因误读了原文表格的一列。内部一致性检查通过,因为所有下游文件(摘要、讨论、表格、森林图标题)都重复了相同的错误数值。仅在二次审核时随机抽取提取内容与原文对照才发现了反转。
不可协商的规则
  1. 存在CSV时不得手动输入数值矩阵
    • 使用
      read.csv(...)
      + 子集/筛选。绝不要手动将论文表格中的2×2表格复制到
      matrix(c(...), ...)
      中。
    • 若确实无法避免手动输入(例如仅文本提取),
      matrix
      c()
      data.frame
      必须添加注释,引用确切的CSV行+列 原文表格/页码坐标。示例:
      r
      # 来源:data_extraction_final.csv第<N>行(<第一作者> <年份>),列<event_arm1>=0,<event_arm2>=1
      # 已对照原文表格<X>,页码<P>验证
      fisher.test(matrix(c(0, 45, 1, 55), nrow = 2, byrow = FALSE))
  2. 比较组子集需单独记录在共识日志中
    • 当某研究的组特异性值(例如多臂研究中的一组)用于比较分析,而该研究的全队列出现在其他地方时,
      extraction_consensus_log.md
      必须为组特异性值添加明确记录行。合并总数和组特异性值不得共享同一行。
  3. 第6阶段结束前随机抽查3项数值声明
    • 森林图/漏斗图/亚组输出稳定后,从草稿手稿的结果部分随机抽取3项数值声明,并分别追溯至(a) R输出日志和(b) 原文表格/图。
    • 将抽查结果记录在
      peer_review_<vN>_internal.md
      的小表格中:
      声明(手稿行号)R输出文件:行号原文来源(论文,表格/图,页码)是否匹配?
    • 任何不匹配均为P0级阻塞问题 — 解决前不得进入第7阶段。
  4. 修订引入的数值必须标记
    • v1版本后添加的任何新数值 — 包括新的比较/亚组/敏感性脚本生成的数值 — 必须在手稿中内联标记为
      [VERIFY-CSV]
      ,直到
      /self-review
      第2.5a阶段审核通过。
  5. 敏感性分析必须基于修改后的数据重新计算,不得复制
    • 添加敏感性/逐一排除/侵蚀/替代模型分析时,所有报告的效应量(Cohen's dz/f、AUC、OR、HR、β、灵敏度/特异度、ICC)必须从修改后的数据集重新推导。若敏感性表格的效应量与主要分析在≥4个值上精确到两位小数完全相同,则几乎可以肯定未重新计算,而是转录了主要分析的值 — 需基于修改后的数据重新运行脚本。
    • 即使效应量看起来相似,基础均值/标准差/计数也会变化;若输入不同但效应量完全相同,这就是明显的信号。≥4个独立值随机精确到两位小数完全相同的概率≈(0.01)^4 — 几乎为零。
    • 该规则针对的失败模式:敏感性分析报告的效应量块与主要表格完全相同,但基础均值/标准差不同 — 敏感性分析实际上从未重新计算。内部一致性无法发现此问题。
  6. “已修复”/“已解决”的审核说明需附重新运行证据,而非仅声明
    • 当前审核说明记录某数值为
      fixed
      resolved
      corrected
      时,仅当附有重新运行证据才有效:时间戳及显示修正值的相关输出/输出文件行,或修改该值的提交记录。仅标注“v10已修复”而无重新运行文件不能清除问题 — 需重新运行脚本并附上输出。
    • 正向流程可能在所有文件中重复陈旧数值,而审核说明声称已修复(例如,主要比较的N值在“已修复”说明后仍显示旧总数)。在接受“已修复”状态前,必须确保结局-分母交叉检查(
      /self-review
      第2.5b阶段,队列算术/研究池锁定断言)针对当前输出通过。
触发时机:每次第6阶段输出更改时(初稿、修订、审稿人要求的重新分析)。即使是“小”修改也不可选 — 上述先例反转发生在“小”修订期间的重新分析中。

Phase 7: GRADE / Certainty of Evidence

第7阶段:GRADE / 证据确定性

Goal: Assess certainty of the body of evidence.
For DTA meta-analysis, apply GRADE-DTA framework:
  1. Risk of bias (from QUADAS-3, or QUADAS-2 for a legacy review)
  2. Indirectness (applicability concerns)
  3. Inconsistency (heterogeneity)
  4. Imprecision (wide CIs, small sample)
  5. Publication bias
For intervention meta-analysis, apply standard GRADE.
Certainty is assessed per outcome, not once for the review. The five domains resolve differently for each outcome — an outcome pooled from 12 studies with narrow CIs and one pooled from 3 with a wide CI do not share a rating, and a single review-level "moderate certainty" sentence tells a reader nothing about the outcome they came for. Rate every outcome carried into the Summary of Findings table, and state the reason for each downgrade (which domain, why) rather than the resulting label alone.
Output: Summary of Findings table — one row per outcome, carrying the pooled estimate with its precision alongside the certainty rating (high / moderate / low / very low).
目标:评估证据体的确定性。
对于DTA元分析,应用GRADE-DTA框架:
  1. 偏倚风险(来自QUADAS-3,旧版综述使用QUADAS-2)
  2. 间接性(适用性担忧)
  3. 不一致性(异质性)
  4. 不精确性(宽置信区间、小样本)
  5. 发表偏倚
对于干预研究元分析,应用标准GRADE框架。
确定性需针对每个结局单独评估,而非针对整个综述统一评估。五个领域在每个结局中的表现不同 — 一项由12项研究合并的结局(窄置信区间)与一项由3项研究合并的结局(宽置信区间)不应共享同一评级,单一综述层面的“中等确定性”语句无法告知读者他们关注的结局情况。对纳入结果总结表的每个结局进行评级,并说明每个降级的原因(哪个领域,为何降级),而非仅说明最终标签。
输出:结果总结表 — 每行对应一个结局,包含合并估计值及其精度,以及确定性评级(高/中等/低/极低)。

Phase 8: Reporting & Manuscript

第8阶段:报告与手稿

Goal: Generate PRISMA-compliant manuscript sections.
Failure-mode cross-ref
references/submission_package_drift.md
— apply the
_build.sh
pattern +
DO_NOT_EDIT_HERE
gate when staging multi-journal submission folders.
  1. Check reporting compliance: Use
    /check-reporting
    with PRISMA-DTA or PRISMA 2020, then run it a second time over the abstract with PRISMA 2020 for Abstracts — 12 items, its own denominator. One run does not cover both.
  2. Write manuscript: Use
    /write-paper
    with meta-analysis type selected
  3. Figures: Use
    /make-figures
    for:
    • PRISMA flow diagram
    • Forest plots (paired for DTA)
    • SROC curve (DTA)
    • Funnel plot
    • RoB summary (traffic light plot)
  4. Tables:
    • Characteristics of included studies
    • 2x2 data per study (DTA)
    • RoB assessment results
    • Summary of findings / GRADE table (one row per outcome — Phase 7)
  5. The items published radiology SR/MA most often drop. Park 2022 (Korean J Radiol; PMID:35213097) scored 24 SR/MAs against PRISMA 2020 and found 24 of 42 items reported by fewer than 80%. The checklist itself lives in
    /check-reporting
    ; what follows is where drafts actually fail, so check these by hand before the compliance run rather than after it:
    PRISMA itemWhat is missingObserved
    20aFor each synthesis, a brief summary of the contributing studies' characteristics and risk of bias — not one global paragraph covering all pools0/24
    27Data availability: which of the extraction forms, extracted data, analysis dataset, and analytic code are public, and where0/24
    24a–cRegistration number, where the protocol can be read, and any amendment — an explicit "not registered" satisfies 24a0/24
    22 / 15Certainty of evidence per outcome, and the method used to assess it9%
    13f / 20dSensitivity analysis: method and result28%
    18Risk of bias per study, shown study-by-study rather than as a pooled proportion32%
    13dRationale for the synthesis model (see Phase 6 check 2)35%
    16bStudies that look eligible but were excluded, cited individually with the reason25%
    Abstract #3, #12Eligibility criteria and registration inside the structured abstract0/24 each
    The abstract items are the cheapest of these and the most reliably forgotten. PRISMA 2020 devotes a separate 12-item instrument to the abstract — item 2 of the main checklist does nothing but defer to it — so a manuscript can satisfy all 42 main-text items and still fail most of the twelve.
    /check-reporting
    carries it as
    PRISMA_2020_Abstracts.md
    ; run it as its own pass and report its score separately, because folding twelve items into a 42-item total is how they stay invisible.
  6. Data availability statement: name what is being shared (extraction template, locked dataset, analysis code, RoB judgments) and where — repository, DOI, or supplementary file. "Available from the corresponding author on reasonable request" satisfies few journals now and no longer satisfies item 27. If a Zenodo DOI is minted post-acceptance,
    references/post_submission_release_ops.md
    covers propagating it back into this statement.
  7. Supplementary & analysis-code pre-submission gate (run before Phase 9 circulation and before portal upload). Presence of the 8-file package (Empirical Lesson 5) is necessary but not sufficient — each item must also be reviewer-ready:
    • De-scaffold: strip internal-QC / tool artifacts before bundling — raw
      /check-reporting
      output ("Assessed by: <tool>", JSON blocks, "READY FOR SUBMISSION" verdicts, action-item lists), search-development planning docs (decision logs, expected-yield estimates,
      [Check on execution]
      placeholders, version-history dev notes), and stale version stamps. Ship a clean PRISMA 2020 checklist (27-item / 42-subitem table only) and an executed-method search-strategy doc, not the working drafts.
    • Blind: supplementary goes to reviewers — remove author names/initials and sibling-project cross-references ("Designed by: <name>", "identical to a sibling review"). Same standard as the blinded manuscript.
    • Cross-consistency with the manuscript: every supplementary number must match the main text — PRISMA counts, pool k/N, the Cochrane/CENTRAL search description, RoB counts. A supplement that says "Cochrane — NOT SEARCHED" while Methods report a confirmatory CENTRAL search is a contradiction reviewers catch.
    • Submitted analysis code must reproduce and be self-contained: run it from a clean copy of the bundle. It must (a) read the bundled locked dataset (not an out-of-bundle path) and write to the working directory, and (b) regenerate every pool reported in the results table. A hard-coded study-id subset that drifts from the manuscript (e.g., a pool computed over k=7 while the manuscript reports k=9) is a P0 — fix and re-run; never ship stale code or stale figures derived from it.
    • Run a supplementary-only review pass — the manuscript self-review/panel does not see the supplement; mirror
      /self-review
      Phase 2.5c–2.5d (reference + cross-reference QC) over the supplementary files.

目标:生成符合PRISMA规范的手稿章节。
失败模式交叉引用
references/submission_package_drift.md
— 准备多期刊提交文件夹时应用
_build.sh
模式 +
DO_NOT_EDIT_HERE
标识。
  1. 检查报告合规性:使用
    /check-reporting
    结合PRISMA-DTA或PRISMA 2020,然后针对摘要单独运行PRISMA 2020摘要版检查 — 12项,有独立的评分标准。一次运行无法覆盖两者。
  2. 撰写手稿:使用
    /write-paper
    并选择元分析类型
  3. 生成图表:使用
    /make-figures
    生成:
    • PRISMA流程图
    • 森林图(DTA为配对图)
    • SROC曲线(DTA)
    • 漏斗图
    • 偏倚风险汇总图(红绿灯图)
  4. 生成表格
    • 纳入研究特征表
    • 每项研究的2×2数据表(DTA)
    • 偏倚风险评估结果表
    • 结果总结/GRADE表(每行对应一个结局 — 第7阶段)
  5. 放射学系统评价/元分析最常遗漏的条目。Park等2022年(Korean J Radiol;PMID:35213097)对24项系统评价/元分析进行PRISMA 2020评分,发现42项条目中有24项的报告率低于80%。检查清单本身在
    /check-reporting
    中;以下是草稿实际容易遗漏的条目,因此在合规性运行前手动检查,而非之后:
    PRISMA条目遗漏内容报告率
    20a针对每项合成分析,简要总结纳入研究的特征和偏倚风险 — 而非一个涵盖所有研究池的全局段落0/24
    27数据可用性:提取表单、提取数据、分析数据集和分析代码中哪些是公开的,以及存放位置0/24
    24a–c注册号、方案获取位置及任何修改 — 明确说明“未注册”符合24a要求0/24
    22 / 15每个结局的证据确定性,以及评估方法9%
    13f / 20d敏感性分析:方法和结果28%
    18每项研究的偏倚风险,逐研究展示而非合并比例32%
    13d合成模型的选择依据(见第6阶段检查2)35%
    16b看似符合纳入标准但被排除的研究,逐一引用并说明理由25%
    摘要 #3, #12结构化摘要中的纳入排除标准和注册信息各0/24
    摘要条目最容易修复但也最容易被遗忘。PRISMA 2020为摘要专门制定了12项独立工具 — 主检查清单的第2项仅指向该工具 — 因此手稿可能满足所有42项正文条目,但仍未通过大部分摘要条目。
    /check-reporting
    包含
    PRISMA_2020_Abstracts.md
    ;需单独运行并单独报告得分,因为将12项条目合并到42项总数中会导致它们被忽略。
  6. 数据可用性声明:说明共享内容(提取模板、锁定数据集、分析代码、偏倚风险判断)及存放位置 — 代码库、DOI或补充文件。“合理请求可向通讯作者获取”现在很少被期刊接受,也不再符合第27项要求。若在接受后生成Zenodo DOI,
    references/post_submission_release_ops.md
    涵盖将其同步回声明的流程。
  7. 补充材料 & 分析代码提交前审核(第9阶段传阅前和门户上传前运行)。8文件包的存在(经验教训5)是必要但不充分条件 — 每个条目还需符合审稿人要求:
    • 移除脚手架:打包前移除内部质量控制/工具文件 — 原始
      /check-reporting
      输出(“评估者:<工具>”、JSON块、“可提交” verdict、行动项列表)、检索开发规划文档(决策日志、预期产出估计、
      [执行时检查]
      占位符、版本历史开发笔记)及陈旧版本标记。提交干净的PRISMA 2020检查清单(仅27项/42子项表格)和最终检索策略文档,而非工作草稿。
    • 盲化:补充材料会提交给审稿人 — 移除作者姓名/首字母及兄弟项目交叉引用(“设计者:<姓名>”、“与兄弟综述相同”)。与盲化手稿标准一致。
    • 与手稿交叉一致性:补充材料中的所有数值必须与正文一致 — PRISMA计数、研究池k/N、Cochrane/CENTRAL检索描述、偏倚风险计数。补充材料说明“未检索Cochrane”而方法部分报告已检索CENTRAL属于矛盾,会被审稿人发现。
    • 提交的分析代码必须可复现且自包含:从干净的包副本运行。必须(a) 读取包中的锁定数据集(而非包外路径)并写入工作目录,且(b) 重新生成结果表中报告的所有研究池。硬编码的研究ID子集与手稿不符(例如,脚本计算k=7而手稿报告k=9)属于P0级问题 — 修复并重新运行;绝不要提交陈旧代码或由此生成的陈旧图表。
    • 单独运行补充材料审核 — 手稿自我审核/小组看不到补充材料;参照
      /self-review
      第2.5c–2.5d阶段(参考文献+交叉引用质量控制)对补充材料进行审核。

Phase 9: Co-author Circulation

第9阶段:共同作者传阅

Goal: Standardized pre-submission circulation of the manuscript to co-authors and senior methodologist / reviewer, with a bounded review window and a controlled attachment scope.
Trigger: Phase 8 is complete, and the draft has cleared Phase 6b source-fidelity audit.
Summary: Reply to the prior-version email thread to preserve
In-Reply-To
continuity (v1 → v2 → v3 tracked in one place). Attach the manuscript body with figures inline and, for v≥2, a change summary — exclude graphical abstract, cover letter, COI forms, and supplementary until the target journal is confirmed. TO = corresponding author + one senior methodologist; CC = remaining co-authors. Set a 7-day deadline (5 business days + weekend). Ask the corresponding author for target-journal preference, reviewer candidates, and cover-letter framing.
Load-on-demand procedural detail (thread continuity, attachment scope rationale, size-to-method table, journal-undetermined framing, response-tracking log):
${CLAUDE_SKILL_DIR}/references/phase9_circulation.md
.
Failure-mode cross-ref
references/review_orchestration.md
RO-1~RO-5 (dual-rating completeness, defensive-tone bias audit, response-matrix numeric tracking, 2nd-reviewer availability blocking).

目标:将手稿标准化提交前传阅给共同作者和资深方法学家/审稿人,设定明确的审核窗口和受控的附件范围。
触发条件:第8阶段完成,且草稿通过第6b阶段原文保真度审核。
总结:回复之前版本的邮件线程以保持
In-Reply-To
连续性(v1 → v2 → v3在同一位置追踪)。附件包含内嵌图表的手稿正文,对于v≥2版本,还需包含变更摘要 — 在确定目标期刊前,排除图形摘要、投稿信、利益冲突表单和补充材料。收件人TO = 通讯作者 + 一名资深方法学家;抄送CC = 其余共同作者。设定7天截止日期(5个工作日+周末)。询问通讯作者目标期刊偏好、审稿人候选人和投稿信框架。
按需加载流程细节(线程连续性、附件范围理由、规模-方法表、未确定期刊框架、回复追踪日志):
${CLAUDE_SKILL_DIR}/references/phase9_circulation.md
失败模式交叉引用
references/review_orchestration.md
RO-1~RO-5(双评级完整性、防御性语气偏差审核、响应矩阵数值追踪、第二名评价者可用性阻塞)。

Phase 10: Self-Audit Recovery (v{N} → v{N+1} sprint)

第10阶段:自我审核修复(v{N} → v{N+1}冲刺)

Goal: When an audit uncovers a structural data or protocol-application error, withdraw the current version, rebuild, and re-circulate with a transparent audit trail. Catching the error yourself before a journal reviewer does is the principal trust-building move in this phase.
Trigger conditions (any one):
#TriggerSource
T1Extraction CSV ↔ primary source disagreement for a cell feeding a pooled/subgroup estimate or reported proportionPhase 6b audit
T2Included/excluded study violates the pre-specified criteria on re-readProtocol review
T3Hand-typed numerical literal in the analysis script traces to a wrong valuePhase 6b audit
T4PROSPERO protocol ↔ delivered analysis disagreement on outcome, subgroup, or eligibilityProtocol ↔ analysis diff
T5Dual-reviewer consensus record ↔ locked dataset disagreement on inclusionConsensus log diff
Non-negotiable rule: if the trigger fires after Phase 9 circulation but before journal submission, withdraw the current version within 24 hours. Reviewer discovery is a strictly worse failure mode than self-withdrawal.
Sprint outline (12 steps): (10.1) audit log at
qc/audit_vN_to_vNplus1.md
→ (10.2) CSV re-verification with
[VERIFY-CSV]
tagging → (10.3) fresh script re-run (fixed seed, logged) → (10.4) manuscript auto-sync (grep for v{N} residue) → (10.5) supplementary regeneration (consensus log, RoB, GRADE/SoF, PRISMA flow) → (10.6) figure regeneration via
/make-figures
→ (10.7) change summary with delta table → (10.8) PROSPERO amendment (application correction, not criteria change) → (10.9) re-circulation in the Phase 9 thread with the "On re-review" framing → (10.10) anti-patterns to avoid (hide-and-submit, "minor revision" reframe, cover-letter-only disclosure) → (10.11) post- submission escalation path → (10.12) post-recovery loop (Phase 9 restart; tighten Phase 6b if a second sprint is needed).
Load-on-demand procedural detail (exact audit-log fields, delta-table template, amendment language template, re-circulation paragraph template, anti-pattern rationale):
${CLAUDE_SKILL_DIR}/references/phase10_recovery.md
.
Failure-mode cross-ref
references/post_submission_release_ops.md
Gate 4 covers reject/revise Zenodo versioning, tag-cleanup gate, and re-target workflow (avoid "new version" misuse on re-target).

目标:当审核发现结构性数据或方案应用错误时,撤回当前版本,重建并重新传阅,同时提供透明的审核追踪。在期刊审稿人发现前自行发现错误是本阶段建立信任的关键举措。
触发条件(满足任意一项)
#触发条件来源
T1提取CSV与原文在合并/亚组估计值或报告比例的单元格上存在不一致第6b阶段审核
T2重新阅读发现纳入/排除研究违反预先指定的标准方案审核
T3分析脚本中手动输入的数值字面量对应错误值第6b阶段审核
T4PROSPERO方案与实际分析在结局、亚组或纳入排除标准上存在不一致方案 ↔ 分析差异
T5双评价者共识记录与锁定数据集在纳入情况上存在不一致共识日志差异
不可协商的规则:若触发条件在第9阶段传阅后但期刊提交前触发,需在24小时内撤回当前版本。审稿人发现错误比自行撤回的失败模式严重得多。
冲刺大纲(12步):(10.1) 在
qc/audit_vN_to_vNplus1.md
记录审核日志 → (10.2) CSV重新验证并标记
[VERIFY-CSV]
→ (10.3) 重新运行脚本(固定种子,记录日志) → (10.4) 手稿自动同步(查找v{N}残留内容) → (10.5) 重新生成补充材料(共识日志、偏倚风险、GRADE/结果总结、PRISMA流程图) → (10.6) 通过
/make-figures
重新生成图表 → (10.7) 带差异表的变更摘要 → (10.8) PROSPERO修改(应用修正,而非标准变更) → (10.9) 在第9阶段线程中以“重新审核后”框架重新传阅 → (10.10) 需避免的反模式(隐藏提交、“小修订”重构、仅投稿信披露) → (10.11) 提交后升级路径 → (10.12) 修复后循环(重启第9阶段;若需第二次冲刺则收紧第6b阶段)。
按需加载流程细节(确切的审核日志字段、差异表模板、修改语言模板、重新传阅段落模板、反模式理由):
${CLAUDE_SKILL_DIR}/references/phase10_recovery.md
失败模式交叉引用
references/post_submission_release_ops.md
第4项涵盖拒稿/修订后的Zenodo版本控制、标签清理和重新投稿流程(避免重新投稿时误用“新版本”)。

Failure Modes (prior MA projects, anonymized)

失败模式(来自过往元分析项目,已匿名化)

Failure patterns observed across three prior MA projects (anonymized). Each topical reference extends the phase it cross-references above — consult alongside phase procedural docs, not in isolation.
DomainPhase spanLoad-on-demand reference
Data integrity (2x2 arm-swap, KM audit, methodology mismatch, PRISMA 5-way drift, single-source k)Phase 3 → 6
references/data_integrity_checklist.md
(DI-1~DI-9)
Review orchestration (2nd-reviewer blocking, dual-rating completeness, defensive-tone audit, response-matrix tracking)Phase 9 circulation (extends
phase9_circulation.md
)
references/review_orchestration.md
(RO-1~RO-5)
Submission package drift (multi-journal folder hygiene,
DO_NOT_EDIT_HERE
gate, build artifact vs master)
Phase 8 → submission
references/submission_package_drift.md
Post-submission release ops (Zenodo DOI timing, tag-cleanup gate, reject-retarget versioning)Submission → Phase 10
references/post_submission_release_ops.md
在三个过往元分析项目中观察到的失败模式(已匿名化)。每个主题参考文件扩展了上述交叉引用的阶段 — 需结合阶段流程文档阅读,而非单独阅读。
领域阶段范围按需加载参考文件
数据完整性(2×2组交换、KM审计、方法学不匹配、PRISMA五重偏差、单源k值)第3 → 6阶段
references/data_integrity_checklist.md
(DI-1~DI-9)
综述编排(第二名评价者阻塞、双评级完整性、防御性语气审核、响应矩阵追踪)第9阶段传阅(扩展自
phase9_circulation.md
references/review_orchestration.md
(RO-1~RO-5)
提交包偏差(多期刊文件夹管理、
DO_NOT_EDIT_HERE
标识、构建文件 vs 主文件)
第8阶段 → 提交
references/submission_package_drift.md
提交后发布操作(Zenodo DOI时机、标签清理、拒稿后重新投稿版本控制)提交 → 第10阶段
references/post_submission_release_ops.md

Automation hooks (invoke at the phase listed)

自动化钩子(在指定阶段调用)

WhenScriptGate
Phase 3f reconciliation (before Phase 5 write-up)
python3 ${CLAUDE_SKILL_DIR}/scripts/check_exclusion_code_validity.py --protocol 0_Protocol/protocol.md --screening 2_Screening/*.tsv --strict
validates each applied exclusion code against the registered eligibility criteria:
CODE_CONTRADICTS_ELIGIBILITY
(a code excludes a design the protocol includes — the bulk study-loss defect no arithmetic/inter-rater gate can see),
CODE_NOT_REGISTERED
(off-protocol code),
CODE_RENUMBERED
(same code, two meanings). Challenge card:
scripts/check_exclusion_code_validity_challenge/
.
Phase 4 kickoff (before first extraction row)
python3 ${CLAUDE_SKILL_DIR}/../../scripts/extraction_consensus_log_init.py --output 2_Data/extraction_consensus_log.md
DI-1: creates standalone consensus log so comparative arm-specific rows are never folded into R-script comments.
Phase 3f reconciliation + every revision touching PRISMA numbers
python3 ${CLAUDE_SKILL_DIR}/../../scripts/prisma_5way_consistency.py --ssot prisma.yaml
DI-6: 5-surface drift check (abstract / main text / flow figure / supplement / CSV) against YAML SSOT. Non-zero exit blocks Phase 5 writeup.
Phase 8 pre-submission + every journal retarget
bash ${CLAUDE_SKILL_DIR}/../../scripts/tag_cleanup_gate.sh
DI-8: fails if
VERIFY-CSV
/
TODO
/
FIXME
/
XXX
survive in
7_Manuscript
,
supplement
,
SUBMISSION
, etc.
Phase 8 on first build per journal (
--record
), then before every re-submission (
--verify
)
python3 ${CLAUDE_SKILL_DIR}/../../scripts/verify_package_integrity.py --record --journal <name>
then
--verify --journal <name>
SPD: checksum-based drift detection between master manuscript and built
SUBMISSION/{journal}/
folder. Journal-editable files (cover letter, response, MANIFEST,
DO_NOT_EDIT_HERE.md
) are auto-excluded.
All four scripts are repo-shipped as of 2026-04 (FOLLOWUPS P10). Non-zero exit = gate failure; resolve before proceeding to the next phase.

时机脚本审核内容
第3f阶段一致性核对(第5阶段撰写前)
python3 ${CLAUDE_SKILL_DIR}/scripts/check_exclusion_code_validity.py --protocol 0_Protocol/protocol.md --screening 2_Screening/*.tsv --strict
验证每个应用的排除代码是否与注册的纳入排除标准一致:
CODE_CONTRADICTS_ELIGIBILITY
(代码排除了方案纳入的设计 — 算术/评价者间审核无法发现的大量研究丢失缺陷)、
CODE_NOT_REGISTERED
(方案外代码)、
CODE_RENUMBERED
(同一代码有两种含义)。挑战案例:
scripts/check_exclusion_code_validity_challenge/
第4阶段启动(首次提取前)
python3 ${CLAUDE_SKILL_DIR}/../../scripts/extraction_consensus_log_init.py --output 2_Data/extraction_consensus_log.md
DI-1:创建独立共识日志,确保比较组特异性行不会被折叠到R脚本注释中。
第3f阶段一致性核对 + 任何涉及PRISMA数值的修订
python3 ${CLAUDE_SKILL_DIR}/../../scripts/prisma_5way_consistency.py --ssot prisma.yaml
DI-6:五重偏差检查(摘要/正文/流程图/补充材料/CSV)与YAML可信源对比。非零退出会阻止进入第5阶段撰写。
第8阶段提交前 + 每次重新投稿到不同期刊
bash ${CLAUDE_SKILL_DIR}/../../scripts/tag_cleanup_gate.sh
DI-8:若
VERIFY-CSV
/
TODO
/
FIXME
/
XXX
残留在
7_Manuscript
supplement
SUBMISSION
等文件夹中则失败。
第8阶段针对每个期刊首次构建时(
--record
),然后每次重新投稿前(
--verify
python3 ${CLAUDE_SKILL_DIR}/../../scripts/verify_package_integrity.py --record --journal <name>
然后
--verify --journal <name>
SPD:基于校验和检测主手稿与构建的
SUBMISSION/{journal}/
文件夹之间的偏差。期刊可编辑文件(投稿信、回复、MANIFEST、
DO_NOT_EDIT_HERE.md
)自动排除。
所有四个脚本自2026年4月起随代码库发布(FOLLOWUPS P10)。非零退出 = 审核失败;解决后才能进入下一阶段。

Empirical Lessons (peer-review cycles)

实证经验(同行评审周期)

Sixteen accumulated SR-MA peer-review / submission lessons (2026-05 and 2026-06) — the drivers behind the Phase 4 extraction-form schema, the Phase 4c QC scripts, and the Phase 8 submission gates. To keep this entry point lean they live load-on-demand in
${CLAUDE_SKILL_DIR}/references/empirical_lessons.md
. Load that file when designing the extraction form (before Phase 4) and before submission (Phase 8) — it covers dual-extractor 2x2 integrity, cohort-overlap clustering, small-k subgroup caution, the supplementary 8-file bar, PROSPERO ID format, AI-disclosure presence, recompute-don't-copy sensitivity analyses, outcome harmonization, heterogeneous-RoB κ, survival-specific concerns, supplement blinding / de-scaffolding, self-contained reproducible analysis scripts, sidecar re-sync, methodological
  • software citations, wide-table PDF rendering, and submission-portal journal-identity checks.

16条积累的系统评价-元分析同行评审/投稿经验(2026年5月和6月) — 指导第4阶段提取表单设计、第4c阶段质量控制脚本和第8阶段提交审核。为保持入口简洁,这些经验按需加载于
${CLAUDE_SKILL_DIR}/references/empirical_lessons.md
在设计提取表单(第4阶段前)和提交前(第8阶段)务必阅读该文件 — 涵盖双提取者2×2完整性、队列重叠聚类、小k值亚组注意事项、补充材料8文件要求、PROSPERO ID格式、AI披露要求、敏感性分析重新计算而非复制、结局标准化、异质性偏倚风险κ值、生存研究特定问题、补充材料盲化/移除脚手架、自包含可复现分析脚本、副文件同步、方法学+软件引用、宽表格PDF渲染及提交门户期刊身份检查。

DTA-Specific Pitfalls (Always Check)

DTA特定陷阱(务必检查)

PitfallProblemSolution
Separate pooling of Se/SpIgnores correlationUse bivariate/HSROC model
Ignoring threshold effectFalse heterogeneityCheck Spearman correlation, SROC plot
Standard funnel plot for DTAInappropriateUse Deeks' funnel plot
I-squared only for heterogeneityDoesn't capture threshold effectUse prediction region on SROC
Missing GRADECommon omission in DTA MAApply GRADE-DTA. If <4 studies, assess each domain narratively and state the limitation explicitly
Partial verification biasInflates sensitivityQUADAS-3 3.2 (target condition assessed in all participants). QUADAS-3 has no Flow & Timing domain — that was QUADAS-2
Differential verification biasDistorts both Se and SpQUADAS-3 3.3 (target condition assessed the same way in all participants)
Unevaluable results excludedBiases accuracy estimatesReport intent-to-diagnose analysis

陷阱问题解决方案
灵敏度/特异度单独合并忽略相关性使用双变量/HSROC模型
忽略阈值效应假异质性检查Spearman相关性、SROC图
DTA使用标准漏斗图不适用使用Deeks漏斗图
仅用I²评估异质性未捕获阈值效应使用SROC图的预测区间
缺少GRADE评估DTA元分析常见遗漏应用GRADE-DTA。若研究数<4,逐一叙述评估每个领域并明确说明局限性
部分验证偏倚高估灵敏度QUADAS-3 3.2(所有受试者均接受目标疾病评估)。QUADAS-3无流程与时间领域 — 该领域属于QUADAS-2
差异验证偏倚同时扭曲灵敏度和特异度QUADAS-3 3.3(所有受试者接受相同的目标疾病评估方式)
排除无法评估的结果导致准确性估计偏倚报告意向诊断分析

Small Study Considerations

小样本研究注意事项

When the number of included studies is small (< 10):
  • Bivariate/HSROC model may not converge -- consider univariate random-effects as fallback
  • Publication bias tests are underpowered -- state this limitation
  • Subgroup/meta-regression analysis not recommended
  • Wide prediction regions expected -- emphasize uncertainty in conclusions
  • Consider narrative synthesis as alternative/complement

当纳入研究数量较少(<10项)时:
  • 双变量/HSROC模型可能不收敛 — 考虑采用单变量随机效应模型作为备选
  • 发表偏倚检验效能不足 — 说明该局限性
  • 不推荐亚组/元回归分析
  • 预测区间较宽 — 在结论中强调不确定性
  • 考虑采用叙述性合成作为替代/补充

Skill Interactions

工具交互

WhenCallPurpose
Need literature search
/search-lit
PubMed/Semantic Scholar search with verified citations
Need statistical code
/analyze-stats
Execute R/Python analysis scripts
Need figures
/make-figures
PRISMA flow, forest plots, SROC, funnel plots
Need reporting check
/check-reporting
PRISMA-DTA / PRISMA 2020 compliance (includes Step 4c registration / amendment timing)
Need manuscript writing
/write-paper
Full IMRAD manuscript generation
Need self-review
/self-review
Pre-submission quality check
Self-audit recovery entrypoint (Phase 10)
/write-paper
Step 7.4a
Recovery branch for polish pipelines that surface structural audit failures
/sync-submission
SR-MA gate
/sync-submission
Before submission, verify supplementary package matches all 8 files in
templates/supplementary_8file_checklist.md
(PRISMA, PROSPERO, search strategy, exclusion list, extraction table, per-study x per-domain RoB, subgroup forests, sensitivity / publication bias). AI Disclosure presence check (cross-link
/peer-review
Phase 2A P8). Cite-list duplicate check via
/verify-refs
Gate 5 (duplicate PMID/DOI).

时机调用命令目的
需要文献检索
/search-lit
检索PubMed/Semantic Scholar并验证引用
需要统计代码
/analyze-stats
执行R/Python分析脚本
需要生成图表
/make-figures
生成PRISMA流程图、森林图、SROC曲线、漏斗图
需要检查报告合规性
/check-reporting
PRISMA-DTA / PRISMA 2020合规性检查(包含第4c阶段注册/修改时机)
需要撰写手稿
/write-paper
生成完整IMRAD结构的手稿
需要自我审核
/self-review
提交前质量检查
自我审核修复入口(第10阶段)
/write-paper
第7.4a步
针对发现结构性审核失败的打磨流程的修复分支
/sync-submission
系统评价-元分析审核
/sync-submission
提交前,验证补充材料包是否包含
templates/supplementary_8file_checklist.md
中的所有8个文件(PRISMA、PROSPERO、检索策略、排除列表、提取表格、逐研究逐领域偏倚风险评估、亚组森林图、敏感性/发表偏倚分析)。AI披露存在性检查(交叉链接
/peer-review
第2A阶段P8)。通过
/verify-refs
第5项检查重复引用(重复PMID/DOI)。

Error Handling

错误处理

  • If study type is ambiguous (DTA vs intervention), ask user to clarify before proceeding.
  • If fewer than 4 studies for DTA, warn that bivariate model may not converge.
  • If data extraction is incomplete (missing 2x2 cells), suggest contacting authors or sensitivity analysis with imputed values.
  • If PROSPERO ID is missing, flag as a limitation but continue.
  • Always remind user: this is a methodological support tool; final decisions rest with the research team and ideally include a biostatistician/methodologist.
  • 若研究类型不明确(DTA vs 干预研究),在继续前请用户澄清。
  • 若DTA研究数少于4项,警告双变量模型可能不收敛。
  • 若数据提取不完整(缺少2×2单元格),建议联系作者或采用填充值进行敏感性分析。
  • 若缺少PROSPERO ID,标记为局限性但继续流程。
  • 始终提醒用户:这是方法学支持工具;最终决策由研究团队做出,理想情况下应包含生物统计学家/方法学家。

Anti-Hallucination

防幻觉

  • Never fabricate variable names, dataset column names, or variable codings. If a variable mapping is uncertain, output
    [VERIFY: variable_name]
    and ask the user to confirm against the data dictionary.
  • Never fabricate statistical results — no invented p-values, effect sizes, confidence intervals, or sample sizes. All numbers must come from executed code output.
  • Never generate references from memory. Use
    /search-lit
    for all citations.
  • If a function, package, or API does not exist or you are unsure, say so explicitly rather than guessing.
  • 绝不要编造变量名、数据集列名或变量编码。若变量映射不确定,输出
    [VERIFY: variable_name]
    并请用户对照数据字典确认。
  • 绝不要编造统计结果 — 不得虚构p值、效应量、置信区间或样本量。所有数值必须来自执行代码的输出。
  • 绝不要凭记忆生成参考文献。所有引用均使用
    /search-lit
    获取。
  • 若不确定函数、包或API是否存在,明确说明而非猜测。