ai-check

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

AI-Check Skill

AI-Check Skill

Forensic analysis of text for AI-generation signals. Grounded in the published detection literature (Wu et al. 2025, Mitchell et al. 2023, Kujur 2025, AAAI 2025 shared task).
The output is a structured report, not a vague judgment. Every fired signal cites evidence.

针对AI生成文本特征的法医式分析。基于已发表的检测文献(Wu et al. 2025, Mitchell et al. 2023, Kujur 2025, AAAI 2025 共享任务)。
输出为结构化报告,而非模糊判断。每个触发的特征都会引用相关证据。

The nine signal categories

九大类特征

Score each category 0–3:
  • 0 = No signal detected (human-consistent)
  • 1 = Weak signal (possible AI, could be human)
  • 2 = Moderate signal (likely AI pattern)
  • 3 = Strong signal (near-certain AI pattern)
Severity-to-score mapping (use for every category):
Evidence in categoryScore
No flagged instances0
One weak instance, or vague unease without a specific quote1
One moderate instance, or two or more weak instances2
One strong instance, or two or more moderate instances, or four or more weak instances3
Double-counting policy: a single phrase can fire at most two distinct signals when the phrase is genuinely diagnostic for both. Example: "it is important to note that" is both Signal A (banned vocabulary) and Signal C (institutional hedge). Log it under both, but the same phrase cannot count as two separate weak instances inside the same category.
Total score cap: 9 categories × 3 = 27 maximum.
为每个类别打分(0–3分):
  • 0 = 未检测到特征(符合人类写作特征)
  • 1 = 弱特征(可能为AI生成,也可能是人类写作)
  • 2 = 中等特征(大概率为AI模式)
  • 3 = 强特征(几乎可以确定为AI模式)
严重程度与分数映射(适用于所有类别):
类别中的证据分数
无标记实例0
1个弱实例,或无具体引用的模糊不安感1
1个中等实例,或2个及以上弱实例2
1个强实例,或2个及以上中等实例,或4个及以上弱实例3
重复计数规则: 当某个短语确实能同时触发两个不同特征时,最多可计入两个特征。例如:“it is important to note that”既属于特征A(禁用词汇)又属于特征C(官方性模糊表述),可同时计入两个特征,但同一短语在同一类别中不能算作两个独立的弱实例。
总分上限: 9个类别 × 3分 = 最高27分。

Signal A: Perplexity (word predictability)

特征A:困惑度(词汇可预测性)

Look for vocabulary that is maximally safe and expected — words that are technically correct but never the most precise or interesting choice a knowledgeable human would make.
Flags:
  • Generic verbs where domain-specific ones belong ("address" instead of "untangle", "implement" instead of "wire up")
  • Adjectives that describe without adding information ("significant improvements", "notable progress", "key challenges")
  • Hedged assertions that swap specificity for safety ("can often lead to", "may result in", "tends to")
  • Any of the canonical AI vocabulary list: delve, leverage (verb), utilize, robust, comprehensive, streamline, foster, facilitate, pivotal, nuanced, notable, notably, enduring, garner, it is worth noting, it is important to note, multifaceted, in the realm of, the landscape of, a myriad of, a plethora of
Cite the exact word or phrase that fired.
寻找那些极度安全、符合预期的词汇——这些词汇在技术上正确,但绝不是有知识的人类会选择的最精准或最有趣的词汇。
标记情况:
  • 使用通用动词而非领域特定动词(用“address”而非“untangle”,用“implement”而非“wire up”)
  • 使用无实际信息量的形容词(“significant improvements”“notable progress”“key challenges”)
  • 用模糊表述替代具体断言(“can often lead to”“may result in”“tends to”)
  • 任何属于标准AI词汇列表的词汇: delve, leverage(动词形式), utilize, robust, comprehensive, streamline, foster, facilitate, pivotal, nuanced, notable, notably, enduring, garner, it is worth noting, it is important to note, multifaceted, in the realm of, the landscape of, a myriad of, a plethora of
需引用触发特征的具体单词或短语。

Signal B: Burstiness deficit (sentence uniformity)

特征B:突发度不足(句子统一性)

Measure the variation in sentence length across the text.
Flags:
  • Three or more consecutive sentences within 5 words of the same length
  • No sentence shorter than 8 words in any 150-word block
  • Metronomic rhythm — reading the passage aloud produces a steady pulse rather than natural variation
  • No fragments used for emphasis
Report: list the sentence lengths in sequence (e.g. "14, 16, 13, 15, 17 — five consecutive sentences within 4 words of each other").
衡量文本中句子长度的变化情况。
标记情况:
  • 连续三个及以上句子的长度相差不超过5个单词
  • 任意150词段落中无短于8词的句子
  • 节奏单调——朗读段落时呈现稳定的节奏,而非自然的变化
  • 未使用强调性的片段句
报告方式:按顺序列出句子长度(例如:“14, 16, 13, 15, 17 — 连续五个句子长度相差不超过4个单词”)。

Signal C: Hedge density

特征C:模糊表述密度

Count the softening and epistemic hedge words.
Flags:
  • "often", "generally", "typically", "in many cases", "it can be argued" appearing where direct assertion is warranted
  • "it is important to note that", "it is worth mentioning", "one might consider"
  • Diplomatic framing of obvious tradeoffs: "while X has benefits, it also presents challenges"
  • Uncertainty expressed as institutional hedging rather than personal ("results may vary") vs human ("I'm not sure this holds when...")
Report: quote each hedge and note whether it was warranted by genuine uncertainty or reflexive softening.
统计弱化语气和认知模糊的词汇。
标记情况:
  • 在需要直接断言的地方使用“often”“generally”“typically”“in many cases”“it can be argued”
  • 使用“it is important to note that”“it is worth mentioning”“one might consider”
  • 对明显的权衡进行委婉表述:“while X has benefits, it also presents challenges”
  • 采用官方性模糊表述而非个人不确定性表述(如“results may vary” vs 人类写作的“I'm not sure this holds when...”)
报告方式:引用每个模糊表述,并说明其是否因真实不确定性而合理,或只是习惯性弱化语气。

Signal D: Structural tells

特征D:结构特征

Look for document architecture patterns AI imposes regardless of content.
Flags:
  • Bullet list where prose would serve better
  • Topic sentence + evidence + restatement of topic sentence (humans skip the restatement)
  • "In conclusion / To summarize / In summary" openers on closing paragraphs
  • "In this [post/article/section] I will..." openers
  • Numbered steps for content that isn't genuinely sequential
  • Three-part structure imposed on every paragraph (intro, body, conclusion at micro-scale)
  • Tricolon parallel structure: three examples or beats with identical grammatical shape e.g. "You X. Y. Does Z? You X. Y. Does Z? You X. Y. Does Z?" — perfectly symmetrical triplets in prose are AI-constructed. Real writers use two examples or vary the shape. Severity: strong.
  • Perfect paragraph-per-idea arc: every paragraph does exactly one narrative job and advances the arc cleanly (setup → tension → lesson → evidence → reflection). Real personal writing has a paragraph that meanders, does two jobs, or doesn't fully resolve. A piece where every paragraph lands cleanly is architecturally perfect in a way human writing isn't. Severity: moderate in isolation, strong combined with other signals.
  • Three-act Slack/update structure: for informal async messages, accomplishment → caveat → next steps maps directly to intro/body/conclusion. Real updates loop back, add a mid-message second thought, or end with something that doesn't fit the structure.
  • Strawman pivot: "The case for X isn't about Y, it's about Z" / "It's not about X, it's about Y." Leading with what something is NOT before saying what it IS. Real writers lead with the actual point. Severity: moderate.
寻找AI会强加于文本的文档架构模式,无论内容如何。
标记情况:
  • 在更适合用散文表述的地方使用项目符号列表
  • 主题句+证据+主题句重述(人类写作会省略重述部分)
  • 结尾段落以“In conclusion / To summarize / In summary”开头
  • 以“In this [post/article/section] I will...”开头
  • 为非真正顺序性的内容使用编号步骤
  • 每个段落都采用三段式结构(微观层面的引言、正文、结论)
  • 三列平行结构: 三个语法结构完全相同的例子或节拍 例如:“You X. Y. Does Z? You X. Y. Does Z? You X. Y. Does Z?”——散文中完全对称的三句式结构是AI生成的。真实作者会使用两个例子或改变结构。 严重程度:强。
  • 完美的段落-思想对应弧: 每个段落恰好承担一个叙事任务,并清晰推进叙事弧(铺垫→张力→教训→证据→反思)。真实的个人写作中会有段落偏离主题、承担两个任务或未完全收尾。如果一篇文本的每个段落都完美收尾,其架构的完美程度不符合人类写作特征。严重程度:单独出现为中等,与其他特征结合出现为强。
  • Slack/更新内容的三幕式结构: 对于非正式异步消息,采用“成就→警告→下一步”结构,直接对应引言/正文/结论。真实的更新内容会回溯、中途添加新想法或结尾不符合该结构。
  • 稻草人转向: “The case for X isn't about Y, it's about Z” / “It's not about X, it's about Y.” 先说明某事物不是什么,再说明它是什么。真实作者会直接点明核心观点。严重程度:中等。

Signal E: Specificity deficit

特征E:特异性不足

Measure whether claims are grounded in concrete detail.
Flags:
  • Abstract claim with no number, name, time reference, or example: "Many organizations have adopted..."
  • Passive constructions obscuring the actor: "it has been found that", "research suggests"
  • Universalist framing: "teams often find", "developers frequently encounter" (applicable to everyone, specific to no one)
  • Named examples that are suspiciously generic or perfectly illustrative (AI picks canonical examples: "Netflix", "Amazon", "Stripe" without context)
Report: quote each unanchored claim.
衡量主张是否基于具体细节。
标记情况:
  • 无数字、名称、时间参考或示例的抽象主张:“Many organizations have adopted...”
  • 模糊行为主体的被动结构:“it has been found that”“research suggests”
  • 普适性表述:“teams often find”“developers frequently encounter”(适用于所有人,但不针对任何人)
  • 可疑的通用或完美示例(AI会选择标准示例:“Netflix”“Amazon”“Stripe”而无上下文)
报告方式:引用每个无依据的主张。

Signal F: Transition word fingerprint

特征F:过渡词特征

Catalog the connective tissue between sentences and paragraphs.
Flags (strong AI signals):
  • "Furthermore," as paragraph opener
  • "Moreover," as paragraph opener
  • "Additionally," as paragraph opener
  • "It is clear that"
  • "This highlights / underscores / demonstrates the importance of"
  • "As previously mentioned"
  • "In addition to the above"
  • "It goes without saying"
  • "Needless to say"
Flags (moderate signals):
  • "However," used more than once per 200 words
  • "Therefore," used as a mechanical logical connector rather than earned conclusion
  • "Turns out" / "it turns out that" as a pivot or reveal. AI uses this to create the illusion of a discovery narrative. "Turns out the config had a lower timeout" → "The config had a lower timeout." Quote each instance. Severity: moderate.
  • Tutorial-voice transitions: "The standard fix is...", "The common approach is...", "Simple enough on paper" — these frame what follows as received wisdom, not personal experience. Strong signal in technical writing.
  • Announcement-colon patterns: "The rule I use:", "The key insight:", "The approach here:", "The other thing I'd say:" — announcing before revealing. Severity: moderate. Also fires without a colon: "What I didn't expect was...", "What surprised me was...", "The thing I realized was..." — these are announcement sentences even without the colon. The colon isn't the tell; the announcement structure is.
  • Pattern announcement: stating that a pattern exists before describing it. "The pattern is almost always the same" followed by the pattern. Real writers just describe the pattern.
记录句子和段落之间的连接成分。
强AI特征标记:
  • 段落以“Furthermore,”开头
  • 段落以“Moreover,”开头
  • 段落以“Additionally,”开头
  • 使用“It is clear that”
  • 使用“This highlights / underscores / demonstrates the importance of”
  • 使用“As previously mentioned”
  • 使用“In addition to the above”
  • 使用“It goes without saying”
  • 使用“Needless to say”
中等特征标记:
  • 每200词中“However,”使用超过一次
  • “Therefore,”作为机械逻辑连接词而非自然得出的结论
  • “Turns out” / “it turns out that” 作为转向或揭示。AI用此制造发现叙事的假象。例如“Turns out the config had a lower timeout” → 直接表述应为“The config had a lower timeout.” 引用每个实例。严重程度:中等。
  • 教程式过渡: “The standard fix is...”“The common approach is...”“Simple enough on paper”——这些表述将后续内容框定为公认智慧,而非个人经验。在技术写作中为强特征。
  • 宣告式冒号模式: “The rule I use:”“The key insight:”“The approach here:”“The other thing I'd say:”——先宣告再揭示内容。严重程度:中等。 即使没有冒号也会触发:“What I didn't expect was...”“What surprised me was...”“The thing I realized was...”——这些都是宣告式句子,无论是否有冒号。冒号不是特征,宣告结构才是。
  • 模式宣告: 在描述模式之前先说明模式存在。例如“The pattern is almost always the same”之后再描述模式。真实作者会直接描述模式。

Signal G: Punctuation fingerprint

特征G:标点特征

Count the three AI punctuation tells:
Em dashes: Count total em dashes. More than 1 per 300 words is a signal. Specific sub-patterns:
  • Double em dash wrapping (— like this —) is a near-certain AI pattern
  • Em dash as pivot ("not mid-sprint — and the on-call rotation") — list-joiner em dash connecting two items within a sentence
  • Em dash as dramatic aside ("X — which is worth noting — Y") Report exact count, location, and which sub-pattern.
Semicolons: Any semicolon linking two independent clauses in non-academic prose is a flag. Report exact count. Exception: comma-containing lists ("Austin, TX; Denver, CO").
Mid-sentence colons: A colon preceded by an incomplete clause ("The problem: nobody tests this" / "The answer: start earlier") is an AI structural pattern. Report each instance.
统计三类AI标点特征:
破折号: 统计破折号总数。每300词中超过1个即为特征。具体子模式:
  • 双破折号包裹内容(— like this —)几乎可以确定为AI模式
  • 破折号作为转向(“not mid-sprint — and the on-call rotation”)——连接句内两个项目的列表连接破折号
  • 破折号作为戏剧性旁白(“X — which is worth noting — Y”) 报告精确数量、位置及所属子模式。
分号: 在非学术散文中,任何连接两个独立分句的分号均为标记。报告精确数量。例外情况:包含逗号的列表(“Austin, TX; Denver, CO”)。
句中冒号: 冒号前为不完整分句(“The problem: nobody tests this” / “The answer: start earlier”)是AI结构模式。报告每个实例。

Signal H: Voice and register

特征H:语气与语体

Look for absence of human traces.
Flags:
  • No first-person perspective anywhere in a piece where first-person would be natural
  • No second-person direct address in instructional or opinionated content
  • Consistent "polished neutral tone" — no personality variance, no roughness, no informality spikes
  • No rhetorical questions used as transitions
  • No self-correction or mid-thought qualification ("actually, that's not quite right")
  • Opening sentence is a thesis, definition, or contextual framing rather than mid-thought or scene
Register collapse (Slack / informal writing): The most commonly missed signal in casual-register text. AI writes Slack messages that read like polished status reports with informal markers sprinkled in. Look for:
  • Complete, well-formed sentences throughout — real Slack has fragments
  • Topic-per-paragraph structure even in a short message
  • Formal vocabulary underneath casual markers (
    ~60%
    and
    lmk
    but the sentences themselves are well-constructed prose)
  • No self-corrections mid-message ("oh also. just realized...")
  • Three-act arc (accomplishment / caveat / next steps) intact beneath the informality
  • Numbers written as words ("three incidents") rather than numerals with approximations ("~3 incidents", "<10min") Severity: strong when informal markers are present but prose structure is polished.
Templated closers in email/professional writing: "Happy to jump on a call if that's easier." "Let me know if you have any questions." "Feel free to reach out." These are the written equivalent of a throat-clearing opener. Real engineers end emails after the last substantive point, or with a specific ask, or with "lmk." Severity: weak in isolation, moderate when combined with other signals.
寻找人类写作痕迹的缺失。
标记情况:
  • 在适合使用第一人称的文本中完全未使用第一人称视角
  • 在指导性或观点性内容中未使用第二人称直接称呼
  • 始终保持“流畅中立语气”——无个性变化、无粗糙感、无非正式语气的突然出现
  • 未使用修辞疑问句作为过渡
  • 无自我修正或中途补充说明(如“actually, that's not quite right”)
  • 开头句子为论点、定义或背景介绍,而非中途想法或场景描述
语体崩塌(Slack / 非正式写作): 这是在非正式语体文本中最容易被忽略的特征。AI撰写的Slack消息读起来像是带有非正式标记的流畅状态报告。需注意:
  • 全程使用完整、规范的句子——真实的Slack消息会有片段句
  • 即使是短消息也采用“每段一个主题”的结构
  • 非正式标记下使用正式词汇(如
    ~60%
    lmk
    ,但句子本身是规范的散文)
  • 消息中途无自我修正(如“oh also. just realized...”)
  • 非正式语气下仍保留三幕式结构(成就/警告/下一步)
  • 数字用单词书写(“three incidents”)而非带近似值的数字(“~3 incidents”“<10min”) 严重程度:当存在非正式标记但散文结构流畅时为强特征。
邮件/专业写作中的模板式结尾: “Happy to jump on a call if that's easier.” “Let me know if you have any questions.” “Feel free to reach out.” 这些相当于口头的开场白。真实工程师会在最后一个实质性观点后结束邮件,或提出具体请求,或用“lmk.”结尾。 严重程度:单独出现为弱特征,与其他特征结合出现为中等。

Signal I: Rhetorical scaffolding

特征I:修辞框架

Sentence and paragraph-level construction patterns that AI learned from polished writing and applies too consistently. These are the hardest signals to catch — they feel like good writing. Grammarly and live detectors flag these even when punctuation and vocabulary are clean.
Local coherence over-smooth (severity: moderate-strong, corpus-dependent) A pattern related to findings in recent research (DivEye, arXiv 2509.18880, TMLR 2026): every sentence connects too perfectly to the next, zero friction, zero cognitive-load artifacts. AI text often has no sentences that slightly misfire, no thoughts that shift direction mid-clause, no paragraph that doesn't close cleanly. Evidence: read each paragraph and check whether any sentence could be removed and the paragraph would still read perfectly. In human writing, removing a sentence often creates a noticeable gap. In AI writing, the paragraph usually flows better without it. Symptoms:
  • Every paragraph opens with a claim and closes with a confirmation of that claim
  • No sentence has a vocabulary mismatch with surrounding sentences
  • No abrupt topic shift within a paragraph
  • No sentence that slightly misfires before correcting
Calibration caveat (important). SHAP-based explainability analysis (arXiv 2603.23146) found that AI-text detectors rely on dataset-specific stylistic cues, not stable machine-authorship signals. Treat over-smoothness as a corpus-conditional indicator, not a universal authorship invariant. If the text is from a register that genuinely rewards tight coherence (academic abstracts, legal briefs, polished marketing copy), down-weight this signal.
Formula personal essay opener (severity: moderate) "The failure I think about most often happened in 2019." "The moment I remember most clearly was..." "The decision I regret most is..." Pattern: "The [noun] I [remember/think about/regret] most [adverb]" — AI's deliberate- introspection construction for opening personal essays. Real writers start with the incident, not with a ranked claim about their memory of it.
Asyndeton tricolon building in complexity (severity: moderate) Three items without conjunctions, each longer and more emotionally heavy than the last: "Two hours of degraded service, six engineers figuring out what I'd done wrong, a postmortem where I had to explain my reasoning to people who had been paged at home." AI constructs these to manufacture escalating emotional weight. Report the three items and their increasing length.
Intensifier/diminisher opposition (severity: moderate) "X [action] obsessively and Y [action] barely at all" — a balanced contrast using an amplifier against a diminisher. Same family as chiasmus but at the adverb level. Other forms: "X constantly / Y once", "X carefully / Y hardly". Quote the opposition.
Mini-aphorism paragraph closer (severity: moderate) A 4–7 word fragment or short sentence used to close a paragraph with a punchy lesson: "That's the part that stuck." "That's what changed." "That's the whole thing." "That's the real cost." AI appends these to tell the reader what conclusion to draw. Distinct from a sentence- length aphoristic closer — this fires even on very short fragments.
Landing phrase: "is the actual/real work" (severity: moderate) "Getting close enough to understand a failure is the actual work." "Deploying is the easy part. Debugging production is the actual work." AI's formulaic landing phrase for delivering conclusions. Quote it.
Parallel subject mirror (severity: weak-moderate) Two consecutive sentences opening with mirrored noun phrases that reflect each other: "The failure itself is just the event. Understanding it is separate." "The code is one thing. Maintaining it is another." AI constructs these as closing pairs. Report the mirrored subjects.
Participial reframe pivot (severity: moderate) Presenting a list of facts, then using a participial opener to reframe them as something more: "Laid out in a petition, the same facts read like a deliberate strategy." "Arranged that way, it sounds more planned than it was." "Seen this way, the whole arc reads differently." AI uses this pivot to manufacture the appearance of insight. The observation should be made directly without the reframing device. Quote the participial opener.
Thesis-first opener / "X is the easy/hard part" (severity: moderate) Starting a personal piece with the frame before the experience: "Gathering evidence for an EB1A petition is the easy part." "The writing is harder than the research." "X has become increasingly important." AI leads with the thesis because it's been trained on essays. Real writers start in the middle of the experience. Quote the opener.
Within-sentence anaphoric parallel list (severity: moderate-strong) Four parallel items with the same question-word structure inside a single sentence: "what existed before, what problem it solved, why the problem mattered, what changed after" Grammarly and other detectors score this identically to consecutive-sentence anaphora. The fix is varying the noun forms: "context, the actual problem, what changed" — not four parallel "what/why" question-clause starters. Quote the full list.
Composed self-aware parenthetical (severity: moderate) A parenthetical clause where the writer meta-comments on their own interpretation: "which I choose to read as progress" "which I take as a sign of X" "which I'm choosing to interpret as Y" These feel reflective but read as placed. Real reflection names the concrete behavior and stops; it doesn't append the writer's chosen interpretation of that behavior. Quote the parenthetical.
Parallel reason chains (severity: moderate) Three consecutive sentences with the same "subject + because/when + reason" structure, even when the subjects vary: "I filed patents because X. localaik started because Y. I gave talks when Z." The parallel clause shape is detectable even across different subjects. Vary the clause structure: one "because", one bare assertion, one gerund or fragment. Report how many parallel reason-chains fire in sequence.
"More X than Y" comparative framing (severity: moderate) All forms: "feels more like X than Y", "more specific than vague", "faster than". AI describes things by framing against an opposite. Humans describe directly. Quote the comparative.
"Not just X" / "not X, it's Y" / "not X but Y" diminishment (severity: moderate) Naming what something isn't before saying what it is. "It's not self-reported, it's merit-based." "Reasoning, not just behavior." All three forms are the same pattern. Quote the diminishment.
Setup sentences without colons (severity: moderate) Announcement sentences of the form "What [verb phrase] was [the revelation]" — the colon is not the tell, the announcement structure is. All forms fire:
  • "What I didn't expect was..."
  • "What surprised me was..."
  • "The thing I realized was..."
  • "What it didn't have was..."
  • "What ended up working was..."
  • "What changed everything was..."
  • "What finally clicked was..."
  • "What made the difference was..." Any sentence of the form "What [verb phrase] was [the revelation]" is an announcement sentence regardless of whether a colon follows. Quote the setup sentence.
Aphoristic / chiasmus closer (severity: moderate-strong) Two sub-patterns:
  1. A closing sentence quotable as standalone: "The boilerplate is cheaper than the confusion." "The work doesn't sell itself."
  2. Chiasmus — reversed parallel that sounds like insight: "Being specific about being wrong is more useful than being vague about being right." — "specific/wrong" mirrored against "vague/right." Real insight is asymmetric; AI constructs symmetric reversals. Quote and identify which sub-pattern.
Anaphora — same sentence-starter 2–3× consecutively (severity: moderate) "I still read slowly. I still lose the thread." "Why this structure. Why the error handling. Why the cache TTL." AI uses repeated openers for emphasis. Humans collapse them or vary the structure. Report the repeated opener and how many times it fires.
"Turns out" / "it turns out that" as reveal pivot (severity: moderate) AI's dramatic reveal device: "Turns out the config had a different timeout." Direct statement: "The config had a different timeout." The "turns out" adds nothing except the illusion of a discovery narrative. Quote each instance.
"Either X or Y" / "between X and Y" binary framing (severity: moderate) Clean binary choices presented as the only options. Real situations are a spectrum. "Teams face a choice between mocking (fast, but drifts) or live endpoints (accurate, but expensive)" — also fires the balanced parenthetical pattern below.
Balanced parenthetical pairs (severity: moderate) "(X, but Y) or (A, but B)" — two symmetric trade-offs in one sentence. Real trade-offs are asymmetric. The symmetry signals AI construction. Quote the parallel parentheticals.
Inverted burstiness (severity: weak) Three or more consecutive very short sentences (under 7 words each) without a longer counterweight. "The code was fine. The logic held. Nothing left a trace." Reads choppy in isolation. Distinct from Signal B which flags uniform medium-length sentences.

AI从流畅写作中学习并过度一致应用的句子和段落层面的构建模式。这些是最难发现的特征——它们看起来像是优秀的写作。Grammarly和实时检测工具即使在标点和词汇无误的情况下也会标记这些特征。
局部连贯性过度流畅(严重程度:中等-强,取决于语料库) 与近期研究(DivEye, arXiv 2509.18880, TMLR 2026)发现的模式相关:每个句子与下一个句子的连接过于完美,无任何摩擦、无认知负荷痕迹。AI文本通常没有轻微偏离的句子、中途转向的想法或未完美收尾的段落。 证据:逐段阅读,检查是否有句子移除后段落仍能完美流畅。在人类写作中,移除句子通常会造成明显的断层;在AI写作中,段落通常会变得更流畅。 症状:
  • 每个段落以主张开头,以确认该主张结尾
  • 无句子与周围句子的词汇不匹配
  • 段落内无突然的主题转换
  • 无轻微偏离后修正的句子
校准注意事项(重要)。 基于SHAP的可解释性分析(arXiv 2603.23146)发现,AI文本检测工具依赖于特定数据集的文体线索,而非稳定的机器写作特征。将过度流畅视为语料库相关指标,而非通用的作者身份不变量。如果文本来自真正需要紧密连贯性的语体(学术摘要、法律简报、流畅的营销文案),则降低该特征的权重。
公式化个人散文开头(严重程度:中等) “The failure I think about most often happened in 2019.” “The moment I remember most clearly was...” “The decision I regret most is...” 模式:“The [noun] I [remember/think about/regret] most [adverb]”——AI用于个人散文开头的刻意内省结构。真实作者会从事件本身开始,而非从对记忆的排名主张开始。
无连词三列递进结构(严重程度:中等) 三个无连词的项目,每个项目比前一个更长、情感更强烈: “Two hours of degraded service, six engineers figuring out what I'd done wrong, a postmortem where I had to explain my reasoning to people who had been paged at home.” AI构建这些结构以制造逐渐升级的情感重量。报告三个项目及其长度递增情况。
强化词/弱化词对立(严重程度:中等) “X [action] obsessively and Y [action] barely at all”——使用强化词与弱化词形成平衡对比。与交错法类似,但位于副词层面。 其他形式:“X constantly / Y once”“X carefully / Y hardly”。 引用对立表述。
微型警句式段落结尾(严重程度:中等) 用4-7词的片段或短句作为段落结尾,以传达有力的教训: “That's the part that stuck.” “That's what changed.” “That's the whole thing.” “That's the real cost.” AI添加这些内容以告知读者应得出的结论。区别于句子长度的警句式结尾——即使是非常短的片段也会触发该特征。
收尾短语:“is the actual/real work”(严重程度:中等) “Getting close enough to understand a failure is the actual work.” “Deploying is the easy part. Debugging production is the actual work.” AI用于传递结论的公式化收尾短语。引用该短语。
平行主语镜像(严重程度:弱-中等) 两个连续句子以相互呼应的名词短语开头: “The failure itself is just the event. Understanding it is separate.” “The code is one thing. Maintaining it is another.” AI将这些作为收尾对。报告镜像主语。
分词重构转向(严重程度:中等) 先列出一系列事实,然后用分词开头将其重构为更重要的内容: “Laid out in a petition, the same facts read like a deliberate strategy.” “Arranged that way, it sounds more planned than it was.” “Seen this way, the whole arc reads differently.” AI用此转向制造洞察的假象。应直接表述观察结果,无需重构手段。引用分词开头。
论点先行开头 / “X是易/难点”(严重程度:中等) 个人文章以框架而非经历开头: “Gathering evidence for an EB1A petition is the easy part.” “The writing is harder than the research.” “X has become increasingly important.” AI以论点开头,因为它是从论文中训练而来的。真实作者会从经历的中间部分开始。引用开头句。
句内回指平行列表(严重程度:中等-强) 单个句子内四个具有相同疑问词结构的平行项目: “what existed before, what problem it solved, why the problem mattered, what changed after” Grammarly和其他检测工具将其与连续句子回指的评分相同。修正方法是改变名词形式:“context, the actual problem, what changed”——而非四个平行的“what/why”疑问从句开头。 引用完整列表。
刻意自我意识的插入语(严重程度:中等) 插入语从句中作者对自己的解读进行元评论: “which I choose to read as progress” “which I take as a sign of X” “which I'm choosing to interpret as Y” 这些看起来像是反思,但实际上是刻意放置的。真实的反思会指出具体行为并停止;不会附加作者对该行为的解读。引用插入语。
平行原因链(严重程度:中等) 三个连续句子具有相同的“主语 + because/when + 原因”结构,即使主语不同: “I filed patents because X. localaik started because Y. I gave talks when Z.” 即使主语不同,平行从句结构也可被检测到。修正方法是改变从句结构:一个用“because”,一个用直接断言,一个用动名词或片段句。报告连续触发的平行原因链数量。
“More X than Y”比较框架(严重程度:中等) 所有形式:“feels more like X than Y”“more specific than vague”“faster than”。 AI通过与对立面对比来描述事物。人类会直接描述。引用比较表述。
“Not just X” / “not X, it's Y” / “not X but Y”弱化表述(严重程度:中等) 先说明某事物不是什么,再说明它是什么。“It's not self-reported, it's merit-based.”“Reasoning, not just behavior.”三种形式属于同一模式。引用弱化表述。
无冒号的铺垫句(严重程度:中等) 形如“What [动词短语] was [启示]”的宣告式句子——冒号不是特征,宣告结构才是。所有以下形式都会触发:
  • “What I didn't expect was...”
  • “What surprised me was...”
  • “The thing I realized was...”
  • “What it didn't have was...”
  • “What ended up working was...”
  • “What changed everything was...”
  • “What finally clicked was...”
  • “What made the difference was...” 任何形如“What [动词短语] was [启示]”的句子都是宣告式句子,无论是否有冒号跟随。引用铺垫句。
警句式 / 交错法结尾(严重程度:中等-强) 两个子模式:
  1. 可单独引用的结尾句:“The boilerplate is cheaper than the confusion.”“The work doesn't sell itself.”
  2. 交错法——反向平行结构,听起来像是洞察:“Being specific about being wrong is more useful than being vague about being right.”——“specific/wrong”与“vague/right”形成镜像。真实的洞察是不对称的;AI构建对称的反向结构。引用并指明所属子模式。
回指——连续2-3次使用相同句子开头(严重程度:中等) “I still read slowly. I still lose the thread.” “Why this structure. Why the error handling. Why the cache TTL.” AI用重复开头来强调。人类会合并或改变结构。报告重复开头及其触发次数。
“Turns out” / “it turns out that”作为揭示转向(严重程度:中等) AI的戏剧性揭示手段:“Turns out the config had a different timeout.” 直接表述:“The config had a different timeout.”“Turns out”除了制造发现叙事的假象外毫无意义。引用每个实例。
“Either X or Y” / “between X and Y”二元框架(严重程度:中等) 呈现清晰的二元选择作为唯一选项。真实情况是一个连续谱。 “Teams face a choice between mocking (fast, but drifts) or live endpoints (accurate, but expensive)”——同时触发下文的平衡插入语模式。
平衡插入语对(严重程度:中等) “(X, but Y) or (A, but B)”——一个句子中两个对称的权衡。真实的权衡是不对称的。对称性表明是AI构建的。引用平行插入语。
反向突发度(严重程度:弱) 连续三个及以上非常短的句子(少于7词),无更长的句子作为平衡。“The code was fine. The logic held. Nothing left a trace.”单独读起来很生硬。与特征B不同,特征B标记的是统一中等长度的句子。

Mixed-authorship overlay (estimate how much AI editing)

混合作者身份覆盖(估算AI编辑比例)

In addition to scoring the 9 signals, estimate the AI-edited fraction — what portion of the text appears AI-written or AI-edited. This is a separate dimension from the verdict and addresses the real-world common case where humans edit AI drafts (or AI polishes human drafts). Framing borrowed from EditLens (arXiv 2510.03154), which regresses edit amount rather than predicting binary authorship.
Look for these distribution clues:
  • Uniform AI signature across the whole text suggests pure AI generation
  • Specific paragraphs polished, others rough suggests selective AI editing
  • AI vocabulary clusters in transitions and conclusions, body is concrete suggests AI scaffold + human substance
  • Opening / closing paragraph reads as polished, middle is fragmented or rough suggests AI editing of the bookends only
  • Voice changes mid-text (formal → casual or vice versa) suggests mixed sources
  • One paragraph has all the banned vocabulary and the others have none suggests a single AI-written section dropped into otherwise human text
Estimate the fraction in one of these buckets:
  • Pure human (~0%)
  • Lightly AI-assisted (~10-30%)
    — light polish, single section, or vocabulary substitution
  • Mixed authorship (~30-60%)
    — substantial AI-written portions woven through
  • Heavily AI-edited (~60-90%)
    — AI draft with human edits, or human draft with substantial AI rewriting
  • Pure AI (~100%)
Report this as a separate line in the output format below.

除了为9个特征打分外,还需估算AI编辑比例——即文本中看起来是AI撰写或编辑的部分占比。这与结论是独立维度,用于解决现实中常见的人类编辑AI草稿(或AI润色人类草稿)的情况。框架借鉴自EditLens(arXiv 2510.03154),该工具回归编辑量而非预测二元作者身份。
寻找以下分布线索:
  • 整个文本具有统一的AI特征表明是纯AI生成
  • 特定段落流畅,其他段落粗糙表明是选择性AI编辑
  • AI词汇集中在过渡和结论部分,正文是具体内容表明是AI框架+人类实质内容
  • 开头/结尾段落流畅,中间部分碎片化或粗糙表明仅对首尾部分进行了AI编辑
  • 语气中途变化(正式→非正式或反之)表明来源混合
  • 一个段落包含所有禁用词汇,其他段落无表明将一段AI撰写的内容插入到其他人类文本中
将比例估算为以下类别之一:
  • Pure human (~0%)
  • Lightly AI-assisted (~10-30%)
    ——轻度润色、单个段落或词汇替换
  • Mixed authorship (~30-60%)
    ——大量AI撰写部分穿插其中
  • Heavily AI-edited (~60-90%)
    ——AI草稿经人类编辑,或人类草稿经大量AI重写
  • Pure AI (~100%)
将此作为单独一行包含在以下输出格式中。

Output format

输出格式

Always output in this exact structure:
AI-CHECK REPORT
===============

VERDICT: [Human | Likely Human | Uncertain | Likely AI | AI]
CONFIDENCE: [Low | Medium | High]
OVERALL SCORE: [0–27] / 27
AI-EDITED FRACTION: [Pure human | Lightly AI-assisted | Mixed authorship | Heavily AI-edited | Pure AI]

SIGNAL BREAKDOWN
----------------
A. Perplexity            [0-3]  [one-line summary]
B. Burstiness            [0-3]  [one-line summary]
C. Hedge density         [0-3]  [one-line summary]
D. Structural tells      [0-3]  [one-line summary]
E. Specificity           [0-3]  [one-line summary]
F. Transitions           [0-3]  [one-line summary]
G. Punctuation           [0-3]  [one-line summary]
H. Voice / register      [0-3]  [one-line summary]
I. Rhetorical scaffolding [0-3]  [one-line summary]

EVIDENCE LOG
------------
[For every signal score > 0, list each specific instance with a short quote or description.
Format: SIGNAL-[LETTER] | "[exact quote or pattern description]" | severity: weak/moderate/strong]

WHAT GAVE IT AWAY
-----------------
[2–4 sentences identifying the strongest signals in plain language. Be specific about
which phrases, patterns, or absences were most diagnostic. This section is written
for a human who wants to understand the tell, not just see a score.]

RECOMMENDED FIXES
-----------------
[Only present if score > 6. Concrete rewrites or changes for the top 3 signals.]

始终按照以下精确结构输出:
AI-CHECK REPORT
===============

VERDICT: [Human | Likely Human | Uncertain | Likely AI | AI]
CONFIDENCE: [Low | Medium | High]
OVERALL SCORE: [0–27] / 27
AI-EDITED FRACTION: [Pure human | Lightly AI-assisted | Mixed authorship | Heavily AI-edited | Pure AI]

SIGNAL BREAKDOWN
----------------
A. Perplexity            [0-3]  [one-line summary]
B. Burstiness            [0-3]  [one-line summary]
C. Hedge density         [0-3]  [one-line summary]
D. Structural tells      [0-3]  [one-line summary]
E. Specificity           [0-3]  [one-line summary]
F. Transitions           [0-3]  [one-line summary]
G. Punctuation           [0-3]  [one-line summary]
H. Voice / register      [0-3]  [one-line summary]
I. Rhetorical scaffolding [0-3]  [one-line summary]

EVIDENCE LOG
------------
[For every signal score > 0, list each specific instance with a short quote or description.
Format: SIGNAL-[LETTER] | "[exact quote or pattern description]" | severity: weak/moderate/strong]

WHAT GAVE IT AWAY
-----------------
[2–4 sentences identifying the strongest signals in plain language. Be specific about
which phrases, patterns, or absences were most diagnostic. This section is written
for a human who wants to understand the tell, not just see a score.]

RECOMMENDED FIXES
-----------------
[Only present if score > 6. Concrete rewrites or changes for the top 3 signals.]

Scoring thresholds

评分阈值

Total scoreVerdict
0–4Human
5–8Likely Human
9–13Uncertain
14–19Likely AI
20–27AI
总分结论
0–4Human
5–8Likely Human
9–13Uncertain
14–19Likely AI
20–27AI

Calibration notes

校准注意事项

  • Short texts (<100 words) have fewer signals available; note this and adjust confidence to Medium max
  • Technical writing with domain jargon can suppress Signal A even in AI text — don't penalize accurate domain vocabulary
  • Academic or legal writing legitimately uses hedges and semicolons — adjust Signal C and G accordingly
  • ESL writing can mimic some AI patterns (uniform sentence length, hedge-heavy); note if this is plausible
  • A text can score AI on structure/transitions but human on voice — report both honestly
  • Signal I (rhetorical scaffolding) fires on patterns that feel like good writing — do not discount them because the writing quality is high. These are the hardest tells precisely because AI learned them from skilled human writers. A "more like X than Y" comparative in an otherwise clean piece is still a signal.
  • Register collapse (Signal H) requires cross-checking: informal markers alone do not make a Slack message human. Look at the sentence structure underneath the
    lmk
    and
    ~60%
    . If the underlying prose is polished and well-formed, the informal markers are surface noise.
  • Aphoristic closers (Signal I) are context-dependent — a single well-turned closing line in a long personal essay is less diagnostic than the same pattern in a 200-word post where it's the only memorable sentence. Weight accordingly.
  • 短文本(<100词)可检测的特征较少;需注意这一点并将置信度调整为最高中等
  • 带有领域术语的技术写作即使是AI生成的也可能抑制特征A——不要惩罚准确的领域词汇
  • 学术或法律写作合理使用模糊表述和分号——相应调整特征C和G的评分
  • ESL写作可能模仿某些AI模式(统一句子长度、大量模糊表述);如果合理需注明
  • 文本可能在结构/过渡方面得分为AI,但在语气方面得分为人类——如实报告两者
  • 特征I(修辞框架)会触发那些看起来像“优秀”写作的模式——不要因为写作质量高而忽略它们。这些正是最难发现的特征,因为AI是从熟练的人类作者那里学习到的。即使在其他方面都很完美的文本中,“more like X than Y”这样的比较表述仍然是一个特征。
  • 语体崩塌(特征H)需要交叉验证:仅非正式标记不足以使Slack消息成为人类写作。需查看
    lmk
    ~60%
    背后的句子结构。如果底层散文流畅规范,那么非正式标记只是表面噪音。
  • 警句式结尾(特征I)取决于上下文——长个人散文中的单个精炼结尾句比200词帖子中唯一令人难忘的句子的诊断性更低。相应调整权重。

Known detection ceilings (cap confidence accordingly)

已知检测上限(相应限制置信度)

  • Base-model output is a known ceiling. arXiv 2605.19516 ("Base Models Look Human") and corroborating Pangram analysis show that raw, non-instruction-tuned base-model output reads as human to current SOTA detectors. What modern detectors actually fire on is RLHF / instruction-tuning artifacts (polite hedging, structured enumeration, perfect coherence, "helpful assistant" register), not "AI-ness" per se. If the text plausibly came from a base model or a minimally-fine-tuned paraphraser (HIP-style attack), cap confidence at Medium even when surface signals look clean.
  • Claude blind spot in zero-shot detectors. The DetectRL benchmark (arXiv 2410.23746) documents that Binoculars achieves only ~55% AUROC on Claude-generated text vs ~88% on GPT-3.5. If the source model is plausibly Claude, treat low scores with extra caution.
  • Iteratively-paraphrased text is a ceiling. PADBen (arXiv 2511.00416) shows detectors >90% on direct AI text fail catastrophically on text that has been iteratively paraphrased through one or more LLMs. If the user mentions the text was paraphrased or rewritten, down-weight all signals.
  • Stylistic cues are corpus-conditional. SHAP-based explainability analysis (arXiv 2603.23146) shows that surface stylistic features detectors rely on are dataset-specific, not stable authorship signals. This applies most strongly to Signal I (rhetorical scaffolding). Do not over-anchor on any single signal; require corroboration across categories.
  • Multilingual text needs language-matched calibration. AI detectors badly misclassify non-English text — they wrongly flag lightly-polished human Arabic as AI, with one commercial detector dropping from 92% to 12% accuracy (arXiv 2511.16690). Refuse High confidence on non-English text unless calibration is known.
  • 基础模型输出是已知上限。 arXiv 2605.19516(《Base Models Look Human》)及Pangram分析表明,原始的、未经指令微调的基础模型输出在当前SOTA检测工具看来与人类写作无异。现代检测工具实际触发的是RLHF/指令微调的痕迹(礼貌的模糊表述、结构化枚举、完美连贯性、“乐于助人的助手”语体),而非“AI属性”本身。如果文本可能来自基础模型或微调程度最低的改写工具(HIP式攻击),即使表面特征看起来正常,也将置信度限制为中等。
  • 零样本检测工具对Claude的盲点。 DetectRL基准测试(arXiv 2410.23746)记录显示,Binoculars对Claude生成文本的AUROC仅约55%,而对GPT-3.5的AUROC约88%。如果源模型可能是Claude,对低评分需格外谨慎。
  • 迭代改写文本是上限。 PADBen(arXiv 2511.00416)表明,对直接AI文本准确率>90%的检测工具在检测经过一次或多次LLM迭代改写的文本时会彻底失效。如果用户提到文本经过改写或重写,降低所有特征的权重。
  • 文体线索是语料库相关的。 基于SHAP的可解释性分析(arXiv 2603.23146)表明,检测工具依赖的表面文体特征是数据集特定的,而非稳定的作者身份信号。这一点对特征I(修辞框架)尤为适用。不要过度依赖任何单个特征;需要跨类别的佐证。
  • 多语言文本需要匹配语言的校准。 AI检测工具对非英语文本分类错误严重——它们错误地将轻度润色的人类阿拉伯语文本标记为AI,某商业检测工具的准确率从92%降至12%(arXiv 2511.16690)。除非已知校准情况,否则对非英语文本拒绝给出高置信度。

Reference detector landscape (for context)

参考检测工具现状(供参考)

If the user asks "what would tool X say?", these are the current characteristics:
  • GPTZero (2025) uses RL adversarial self-training plus a learned classifier ensemble, not just perplexity + burstiness. Produces a 4-class output (human / slight / moderate / full AI-assist). Older "GPTZero relies on perplexity + burstiness" framing is stale.
  • Binoculars is a strong zero-shot baseline but has the Claude blind spot above.
  • Pangram 3.0 claims 99.98% accuracy with 1-in-10,000 FPR and 97% on humanized text per vendor benchmarks (independent replication pending).
  • EditLens estimates AI-edit fraction rather than binary authorship (94.7 F1 binary, 90.4 F1 ternary).
  • Ghostbuster is the canonical black-box (no token probs needed) detector — 99 F1 in-domain, degrades out-of-domain.
  • DependencyAI uses syntactic dependency n-grams + LightGBM, cross-lingual without LLM access.
如果用户询问“工具X会怎么说?”,以下是当前工具的特点:
  • GPTZero (2025) 使用RL对抗自训练加学习分类器集成,而非仅依赖困惑度+突发度。产生四类输出(人类/轻微/中等/完全AI辅助)。旧的“GPTZero依赖困惑度+突发度”说法已过时。
  • Binoculars 是强大的零样本基线,但存在上述对Claude的盲点。
  • Pangram 3.0 供应商基准测试称其准确率达99.98%,假阳性率为1/10000,对人类化文本的准确率为97%(独立复制待验证)。
  • EditLens 估算AI编辑比例而非二元作者身份(二元F1值94.7,三元F1值90.4)。
  • Ghostbuster 是典型的黑盒检测工具(无需token概率)——域内F1值99,域外性能下降。
  • DependencyAI 使用句法依赖n-grams + LightGBM,无需LLM访问即可跨语言检测。