attribution
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseAttribution
营销归因
You help users answer the hardest question in marketing: which of my efforts actually caused this conversion and this revenue? Attribution is where marketers lose the most money — to channels that look good in one dashboard and terrible in another, to "direct" and "branded search" that hide the real source, and to models that quietly encode an opinion as if it were fact.
This skill has two pillars. Know which one the user needs before you dive in:
- (A) Interpretation — choosing an attribution model, picking a measurement approach, and reconciling the conflicting numbers your tools report. This applies to everyone, even with zero engineering.
- (B) Own your attribution (first-party) — instrumenting and stitching attribution yourself when you control the site/app. This is the build track. Use it when the user says "I want to track this myself" or is hitting a conversion that lives on a domain they don't own.
Most requests start with (A). Reach for (B) only when they control the surface and want to build.
Product context: check for and read it if present — business type, sales cycle, and primary conversion drive almost every recommendation here.
.agents/product-marketing.md你要帮用户解答营销领域最棘手的问题:我的哪些营销举措真正促成了这次转化和营收? 归因是营销人员最容易亏损的环节——有的渠道在一个仪表盘中表现出色,在另一个中却一塌糊涂;“直接访问”和“品牌搜索”会掩盖真实来源;有些模型会悄悄将主观观点包装成事实。
本技能分为两大支柱,在深入解答前要先明确用户需要哪一类:
- (A) 解读类——选择归因模型、确定衡量方法,以及调和工具报告的冲突数据。所有用户都适用,即使没有任何工程基础。
- (B) 自有归因部署(第一方)——当用户掌控网站/应用时,自行部署和关联归因系统。这属于自建路线,当用户表示“我想自行追踪”或需要追踪非自有域名上的转化时使用。
大多数请求始于(A)类,只有当用户掌控平台且希望自建时,才使用(B)类。
产品背景:查看文件(若存在)——业务类型、销售周期和核心转化指标几乎决定了所有推荐方案。
.agents/product-marketing.mdBoundaries — what this skill does NOT own
边界——本技能不涵盖的内容
State these up front so you don't rebuild neighboring skills:
- General event tracking, tracking plans, UTM setup, GA4/GTM → analytics. Attribution assumes tracking exists. The line: analytics = "what events and how to fire them"; attribution = "how touches join to conversions and survive to revenue."
- Ad-platform pixels, CAPI, server-side conversion tracking → ads (). Attribution consumes platform-reported numbers and corrects for their bias; it doesn't set up the pixels.
references/conversion-tracking.md - Pipeline stages, lead lifecycle, CRM revenue dashboards → revops. Attribution feeds pipeline data; it doesn't define stages.
- Showing up in / measuring AI search → ai-seo. Attribution names AI traffic as a blind spot only.
提前说明这些边界,避免重复开发相邻技能的内容:
- 通用事件追踪、追踪计划、UTM设置、GA4/GTM → analytics(分析)。归因默认追踪系统已存在。两者的界限:分析=“要追踪哪些事件以及如何触发”;归因=“如何将触点与转化关联并追踪至营收”。
- 广告平台像素、CAPI、服务器端转化追踪 → ads(广告)(详见)。归因会利用平台报告的数据并修正其偏差,但不负责像素设置。
references/conversion-tracking.md - 销售管线阶段、线索生命周期、CRM营收仪表盘 → revops(营收运营)。归因向管线提供数据,但不定义阶段。
- AI搜索的曝光与衡量 → ai-seo(AI搜索引擎优化)。归因仅将AI流量列为盲区。
Pillar A — Interpretation
支柱A——解读类
1. What attribution can and can't tell you
1. 归因能做什么,不能做什么
Set expectations before touching a number:
- Attribution is directional, not truth. It's a model of causality built from incomplete data (cookies expire, sessions fragment, offline touches vanish, people research on one device and buy on another). Treat it as a strong hint, never a verdict.
- Every model is an opinion. "First-touch" says the first ad gets all the credit; "last-touch" says the closing click does. Both are wrong in opposite directions. Choosing a model is choosing whose story to believe — say so out loud.
- The attribution gap is normal. The sum of channel-reported conversions almost always exceeds real conversions, because every platform claims credit for the same sale. Your job is to shrink and explain the gap, not to make the numbers tie out perfectly. They won't.
When a user demands one true number, reframe: "We can get you a defensible, consistent number and a read on which channels are trending up. A single objective truth doesn't exist — here's why, and here's what we use to make decisions anyway."
在接触数据前先设定预期:
- 归因是方向性的,而非绝对真相。它是基于不完整数据构建的因果模型(Cookie过期、会话碎片化、线下触点消失、用户在一台设备上调研却在另一台设备上购买)。要将其视为强有力的提示,而非最终结论。
- 每个模型都是一种主观观点。“首次触点”认为首个广告应获得全部功劳;“末次触点”认为最终点击应获得全部功劳。两者都存在相反方向的错误。选择模型就是选择采信哪一方的说法——要明确告知用户这一点。
- 归因缺口是正常现象。各渠道报告的转化总和几乎总是超过实际转化数,因为每个平台都会将同一笔交易归为自己的功劳。你的任务是缩小并解释这个缺口,而非强行让数据完全匹配——这是不可能的。
当用户要求一个绝对准确的数字时,重新表述:“我们可以为您提供一个合理、一致的数字,以及各渠道的趋势判断。绝对客观的真相并不存在——以下是原因,以及我们用于决策的替代方案。”
2. Attribution models
2. 归因模型
The six standard models and when each one lies:
| Model | Credit rule | Best for | How it lies |
|---|---|---|---|
| First-touch | 100% to the first known touch | Top-of-funnel / demand-gen valuation; short cycles | Ignores everything that closed the deal; over-credits awareness channels |
| Last-touch | 100% to the last touch before conversion | Direct-response, quick e-comm | Over-credits bottom-funnel + branded search/direct; ignores what created demand |
| Last non-direct | 100% to last touch, skipping "direct" | A cheap fix for direct pollution | Still single-touch; just moves the blind spot |
| Linear | Equal credit to every touch | Long, multi-touch journeys where every step matters | Treats a throwaway visit like a demo; flatters high-frequency channels |
| Time-decay | More credit to touches nearer conversion | Longer cycles where recency matters | Under-credits the top of funnel; still an assumption, not a measurement |
| Position-based (U-shaped) | 40% first, 40% last, 20% middle | B2B with clear "created" + "closed" moments | The 40/40/20 split is arbitrary; middle touches get shortchanged |
| Data-driven (algorithmic/Shapley) | Credit from modeled marginal contribution | High-volume accounts with enough conversions | A black box; needs volume; can't see offline/dark touches it was never fed |
Rules of thumb:
- Never report a single model in isolation for a long sales cycle. Show first-touch and last-touch side by side — the truth lives between them, and the gap between them is the insight.
- Data-driven attribution needs volume (Google Ads historically gated it behind ~3,000 ad interactions and ~300 conversions in 30 days; it has since relaxed the minimums and made DDA the default, but low volume still makes it noise dressed as science). Use position-based instead when you're thin.
- The model matters far less than being consistent and pairing it with an out-of-model sanity check (Pillar A §4, self-reported).
For the model math, worked examples of one journey scored six ways, and Shapley explained plainly, see .
references/attribution-models.md六种标准模型及其局限性:
| 模型 | 归因规则 | 适用场景 | 局限性 |
|---|---|---|---|
| First-touch(首次触点) | 100%归因于首个已知触点 | 漏斗顶部/需求生成评估;短销售周期 | 忽略所有促成交易的后续动作;过度归因于认知类渠道 |
| Last-touch(末次触点) | 100%归因于转化前的最后一个触点 | 直接响应型业务、快消电商 | 过度归因于漏斗底部渠道+品牌搜索/直接访问;忽略需求产生的源头 |
| Last non-direct(末次非直接触点) | 100%归因于最后一个非“直接访问”的触点 | 低成本解决直接访问数据污染问题 | 仍属于单触点模型;只是转移了盲区 |
| Linear(线性模型) | 所有触点平均分配功劳 | 长周期、多触点的用户旅程,每一步都至关重要 | 将偶然访问与演示等关键动作同等看待;高估高频率渠道 |
| Time-decay(时间衰减模型) | 越接近转化的触点获得越多功劳 | 较长销售周期,近期触点更重要 | 低估漏斗顶部渠道;仍属于假设,而非实际衡量 |
| Position-based(U型模型) | 首次触点40%、末次触点40%、中间触点20% | 有明确“需求产生”+“交易达成”节点的B2B业务 | 40/40/20的分配比例是主观设定的;中间触点被低估 |
| Data-driven(算法/Shapley模型) | 基于建模的边际贡献分配功劳 | 高转化量的企业,有足够数据支撑 | 黑箱模型;需要大量数据;无法识别未纳入系统的线下/暗社交触点 |
经验法则:
- 对于长销售周期,绝不要单独报告一种模型。要同时展示首次触点和末次触点的数据——真相介于两者之间,而两者的差距就是洞察点。
- 数据驱动归因需要足够的转化量(谷歌广告过去要求30天内约3000次广告互动和300次转化;现已放宽最低要求并将DDA设为默认,但数据量不足时仍会产生无效数据)。数据量不足时,改用位置型模型。
- 模型本身的重要性远低于一致性,且要搭配模型外的合理性校验(支柱A第4节,自报归因)。
模型算法、同一用户旅程的六种评分示例,以及Shapley模型的通俗解释,详见。
references/attribution-models.md3. The three measurement paradigms
3. 三种衡量范式
Models split credit within your tracked data. Paradigms are how you get at causality — increasingly rigorous, increasingly expensive:
| Paradigm | What it is | Answers | Needs | Watch out |
|---|---|---|---|---|
| MTA (multi-touch attribution) | Stitch user-level touches, apply a model | "Which touchpoints appear on converting journeys?" | Clean cross-device user-level tracking | Cookie loss + privacy have gutted user-level data; it silently under-measures |
| MMM (media/marketing mix modeling) | Top-down regression of spend vs. outcomes over time | "What's each channel's aggregate contribution, including offline/brand?" | 2–3 yrs of weekly data, spend variation | Correlational; slow to react; needs real budget swings to learn |
| Incrementality (geo holdout, PSA, ghost ads, on/off) | Controlled experiment: exposed vs. withheld | "Did this channel cause lift I wouldn't have gotten anyway?" | Ability to withhold; enough volume for significance | The gold standard, but you can only test a few things at a time |
How to choose: small budget / short cycle → good UTM + last-non-direct + a self-reported survey beats a fancy model. Mid budget, several channels → MTA for day-to-day + periodic incrementality tests on your biggest line items. Large budget, offline + brand spend → MMM for the portfolio + incrementality to validate MMM's coefficients. Incrementality is the tiebreaker whenever two channels both claim the same conversions.
Decision table by budget × sales cycle × channel count, and how to read a geo-holdout / PSA test (not a stats tutorial), in .
references/measurement-paradigms.md模型是在已追踪的数据内分配功劳,而范式是获取因果关系的方法——严谨性越高,成本也越高:
| 范式 | 定义 | 解答问题 | 所需条件 | 注意事项 |
|---|---|---|---|---|
| MTA(多触点归因) | 关联用户级触点,应用归因模型 | “哪些触点出现在转化旅程中?” | 清晰的跨设备用户级追踪数据 | Cookie丢失和隐私政策大幅削弱了用户级数据;会静默低估数据 |
| MMM(营销组合模型) | 基于时间维度的支出与结果的自上而下回归分析 | “每个渠道的总贡献是多少,包括线下/品牌推广?” | 2-3年的周度数据、支出波动 | 仅体现相关性;反应缓慢;需要真实的预算波动才能学习 |
| Incrementality(增量测试)(地域对照组、PSA、幽灵广告、开关测试) | 对照实验:曝光组 vs 对照组 | “这个渠道是否带来了原本不会产生的增量效果?” | 能够设置对照组;足够的数据量以确保统计显著性 | 黄金标准,但同一时间只能测试少数内容 |
选择指南: 预算少/销售周期短 → 规范的UTM设置+末次非直接触点模型+自报调查,优于复杂模型。预算中等、多渠道 → 日常使用MTA+定期对核心支出项目进行增量测试。预算充足、涉及线下+品牌推广 → 用MMM做整体组合评估+增量测试验证MMM系数。当两个渠道都声称同一笔转化时,增量测试是最终的判定标准。
按预算×销售周期×渠道数量划分的决策表,以及如何解读地域对照组/PSA测试(非统计教程),详见。
references/measurement-paradigms.md4. Self-reported attribution
4. 自报归因
The most underused signal, and often the most honest for long cycles and dark social. A post-conversion "How did you hear about us?" survey catches what tracking structurally cannot: podcasts, word of mouth, Slack communities, a founder's tweet, "a friend told me."
- When it beats tracking: long consideration cycles, high word-of-mouth, brand/community-led, or heavy dark-social (see §5). If a big slice of your journeys are "direct," you have a self-reported-shaped hole.
- Ask at the moment of conversion (signup, first purchase, demo request) — highest recall, before memory fades.
- Wording: open-ended ("How did you first hear about us?") captures dark social; a short pick-list is easier to quantify but pre-biases the answer. Best practice: pick-list of your known channels plus a free-text "other/tell us more."
- Treat it as a triangulation input, not gospel — recall is fuzzy and people credit the memorable touch, not the first. It's the out-of-model check that keeps your tracked models honest.
- On the build side, this is a form field written to your CRM/analytics as a person property — see Pillar B and .
references/first-party-tracking.md
这是最被低估的信号,对于长周期和暗社交场景往往最准确。转化后弹出的“您是如何了解到我们的?”调查可以捕捉追踪系统本质上无法识别的触点:播客、口碑、Slack社区、创始人的推文、“朋友告诉我的”等。
- 何时优于追踪系统: 长决策周期、高口碑传播、品牌/社区驱动业务,或大量暗社交(见第5节)。如果“直接访问”占比很高,说明你的归因存在自报数据可以填补的缺口。
- 在转化时刻提问(注册、首次购买、申请演示)——此时用户记忆最清晰,不会随着时间推移模糊。
- 提问方式: 开放式问题(“您最初是如何了解到我们的?”)可以捕捉暗社交;简短的选项列表更易于量化,但会预先引导答案。最佳实践:列出已知渠道的选项列表加上“其他/请说明”的自由文本框。
- 将其视为三角验证的输入,而非绝对真相——用户记忆模糊,且会将功劳归于印象深刻的触点,而非首个触点。它是让追踪模型保持客观的模型外校验手段。
- 在自建方面,这是一个写入CRM/分析系统的表单字段,作为用户属性——详见支柱B和。
references/first-party-tracking.md
5. Reconciling conflicting sources
5. 调和冲突数据源
The request behind most attribution work: "Google says 50, Meta says 40, GA says 60, my CRM says 35 — who's right?" Nobody is. Here's the framework.
Why each source systematically lies:
| Source | Biased toward | Because |
|---|---|---|
| Ad platforms (Google/Meta/LinkedIn) | Over-counts itself | Claims view-through + click conversions in its own window; every platform counts the same sale; motivated to look good |
| GA / web analytics | Last non-direct click | Loses cross-device, loses cookie-blocked users, dumps the unknown into direct |
| CRM | Whatever the rep typed / the form captured | Human entry, lead-source overwrites, offline deals with no digital trail |
| Self-reported survey | The memorable touch | Recall bias; under-counts boring-but-real touches like retargeting |
How to triangulate:
- Pick one source of truth for the conversion count — usually your CRM or backend (the system where money is real). Everything else explains where those came from, they don't get to redefine how many.
- Never sum across platforms. If Google and Meta both claim a conversion, you have one conversion with two claimants, not two conversions. De-dupe against the source-of-truth total.
- Read directional agreement, not absolute match. If every source says paid search is up and organic is down this quarter, that trend is trustworthy even though no two numbers match.
- Use self-reported as the tiebreaker when platforms fight over the same conversions, and incrementality when the stakes justify a test.
- Expect and budget for the gap. Report "platforms claim N; we can verify M; the delta is over-claiming + view-through + untracked — here's our best allocation."
The output is an honest allocation with confidence levels, not a false reconciliation to the decimal.
大多数归因工作的核心需求:“谷歌显示50,脸书显示40,谷歌分析显示60,我的CRM显示35——哪个是对的?” 没有一个是完全正确的。以下是解决框架。
各数据源存在系统性偏差的原因:
| 数据源 | 偏差方向 | 原因 |
|---|---|---|
| 广告平台(谷歌/脸书/领英) | 高估自身贡献 | 在自身统计窗口内统计浏览归因+点击转化;每个平台都会将同一笔交易归为自己的功劳;有动机让自己看起来表现更好 |
| GA / 网页分析工具 | 偏向末次非直接点击 | 丢失跨设备数据、丢失Cookie拦截用户数据、将未知来源归为直接访问 |
| CRM | 取决于销售输入/表单捕获内容 | 人工输入、线索来源被覆盖、无数字轨迹的线下交易 |
| 自报调查 | 偏向印象深刻的触点 | 记忆偏差;低估再营销等平淡但真实的触点 |
三角验证方法:
- 选择一个转化计数的真实数据源——通常是CRM或后端系统(记录真实交易的系统)。其他所有数据源仅用于解释转化来源,无权重新定义转化数量。
- 绝不跨平台求和。如果谷歌和脸书都声称同一笔转化,那是一笔有两个申报方的转化,而非两笔转化。要基于真实数据源的总数进行去重。
- 关注方向一致性,而非绝对匹配。如果所有数据源都显示本季度付费搜索增长、自然搜索下降,即使数值不完全匹配,这个趋势也是可信的。
- 当平台争夺同一笔转化时,用自报数据作为判定标准;当决策风险较高时,用增量测试。
- 接受并预留归因缺口。报告“平台申报N次转化;我们可验证M次;差值来自过度申报+浏览归因+未追踪触点——以下是我们的最佳分配方案。”
最终输出应是带有置信度的真实分配方案,而非强行让数据精确匹配的虚假调和结果。
6. The blind spots
6. 盲区
Where conversions hide, making real channels look weak:
- Direct — the junk drawer. Bookmarks and typed URLs, yes, but also stripped referrers, app-to-web, dark social, and any touch your tracking dropped. A large direct share is a measurement problem, not a channel.
- Branded search — people who discovered you elsewhere and Googled your name. Last-touch hands the credit to paid/organic branded search; the real driver was whatever made them search. Segment branded vs. non-branded or you'll defund the top of funnel.
- Dark social — sharing that carries no referrer: DMs, Slack/Discord, podcasts, newsletters, screenshots. Structurally invisible to tracking; self-reported is the only way to see it (§4).
- AI traffic — assistants and AI search increasingly influence buyers, then send them via branded search or direct, so the AI touch is invisible in analytics. Name it and hand deeper work to ai-seo.
The through-line: when "direct" and "branded search" dominate, your top of funnel is working and your attribution is hiding it. Say that explicitly — it's the single most common misread in marketing.
转化数据隐藏的地方,导致真实渠道看起来表现不佳:
- 直接访问——相当于“杂物抽屉”。包括书签和手动输入URL,但也包括丢失的引荐来源、应用跳转网页、暗社交,以及任何被追踪系统遗漏的触点。高占比的直接访问是衡量问题,而非渠道本身的问题。
- 品牌搜索——用户通过其他渠道了解你后,搜索你的品牌名称。末次触点模型会将功劳归于付费/自然品牌搜索,但真正的驱动因素是让用户产生搜索行为的源头。要区分品牌搜索和非品牌搜索,否则会削减漏斗顶部的预算。
- 暗社交——无引荐来源的分享:私信、Slack/Discord、播客、新闻通讯、截图。本质上无法被追踪系统识别;只有自报数据才能捕捉到(见第4节)。
- AI流量——助手和AI搜索越来越多地影响买家,然后引导他们通过品牌搜索或直接访问进入网站,导致AI触点在分析工具中不可见。要明确指出这一点,并将深入工作移交**ai-seo(AI搜索引擎优化)**技能。
核心结论:当“直接访问”和“品牌搜索”占主导时,你的漏斗顶部运作良好,但归因系统掩盖了这一点。 要明确告知用户——这是营销中最常见的误判。
7. Business-type fork
7. 业务类型分支
Defaults differ sharply. Summary here; full playbooks in .
references/by-business-type.md- B2B SaaS (long cycle, sales-assisted): journeys span weeks–months and multiple people, so single-touch models mislead badly. Anchor on the CRM as source of truth, use first-touch + position-based side by side, lean hard on self-reported at demo/signup, and treat pipeline/revenue attribution (→ revops) as the real scoreboard. Offline touches (events, sales convos) make MTA weakest and self-reported strongest here.
- Ecommerce / DTC (short cycle, self-serve): fast journeys, high volume, spend concentrated in paid social + search. Anchor on platform ROAS but distrust it (iOS/CAPI inflation), validate with MMM once spend is material and incrementality/geo-holdouts on your biggest channels, and use a post-purchase survey to catch what pixels miss. Last-touch is defensible for quick-turn SKUs; MMM+incrementality is how you allocate the real budget.
默认方案差异很大。以下是总结;完整指南详见。
references/by-business-type.md- B2B SaaS(长周期、销售辅助):用户旅程跨越数周–数月,涉及多人,因此单触点模型极易产生误导。以CRM为真实数据源,同时使用首次触点+位置型模型,在演示/注册环节大力推行自报归因,并将管线/营收归因(→ revops营收运营)视为核心指标。线下触点(活动、销售对话)使得MTA的效果最差,而自报数据的效果最好。
- 电商 / DTC(短周期、自助服务):用户旅程快、转化量高,支出集中在付费社交+搜索。以平台ROAS为参考但保持质疑(iOS/CAPI数据膨胀),当支出达到一定规模后用MMM验证,并对核心渠道进行增量/地域对照组测试,通过购买后调查捕捉像素无法追踪的触点。对于周转快的SKU,末次触点模型是合理的;MMM+增量测试是分配真实预算的方法。
Pillar B — Own your attribution (first-party)
支柱B——自有归因部署(第一方)
Use this when the user controls the site/app and wants to instrument attribution themselves — especially for a conversion that happens on a domain they don't own (a SavvyCal/Calendly/Cal.com booking, a Stripe Checkout page). This pillar is grounded in real production builds; the full runbook with code patterns is in . The essentials:
references/first-party-tracking.md当用户掌控网站/应用且希望自行部署归因系统时使用——尤其是转化发生在非自有域名上的场景(如SavvyCal/Calendly/Cal.com预订、Stripe结账页面)。本支柱基于实际生产环境的构建经验;包含代码模式的完整手册详见。核心要点:
references/first-party-tracking.mdThe identity graph
身份图谱
First-party attribution is one idea: join anonymous browsing to the eventual conversion.
- A visitor arrives anonymously; your analytics tool assigns an anonymous and stamps first-touch properties (
distinct_id,$initial_referrer) on their events.$initial_utm_* - At conversion (signup, booking, purchase) you call with a stable id (email or user UUID). This merges the anonymous history into a known person — first-touch now survives all the way to the conversion.
identify() - Every conversion event can now be broken down by first-touch channel. That's the whole game.
第一方归因的核心思路:将匿名浏览行为与最终转化关联起来。
- 访客匿名访问;你的分析工具分配一个匿名,并在其事件上标记首次触点属性(
distinct_id、$initial_referrer)。$initial_utm_* - 转化时(注册、预订、购买)调用**接口,传入稳定ID(邮箱或用户UUID)。这会合并**匿名历史记录与已知用户——首次触点数据现在可以一直追踪到转化。
identify() - 现在可以按首次触点渠道拆分每个转化事件。核心逻辑就是这样。
Closing the identify()
gap
identify()填补identify()
缺口
identify()The most common first-party failure: nothing ever calls , so conversions never join to browsing history and every customer looks like they appeared from nowhere. (Framing adapted from Tessa Kriesel's PostHog approach.) The fix is to call identify at each real conversion. Audit first — many SaaS apps already identify at signup; don't rebuild what works. Find the specific un-instrumented conversions and close only those.
identify()第一方归因最常见的失败:从未调用,导致转化无法与浏览历史关联,所有客户看起来都像是凭空出现的。(框架改编自Tessa Kriesel的PostHog方法)。解决方案是在每次真实转化时调用identify接口。先审计——许多SaaS应用已在注册时调用identify;不要重复建设。找到具体的未部署转化场景,仅填补这些缺口。
identify()Stitching conversions on a third-party domain
关联第三方域名上的转化
The one case that needs real machinery: a conversion that completes on a domain you don't control (a booking tool, a hosted checkout). You can't run your analytics there, so:
- At click time, a capture-phase link decorator appends the visitor's anonymous to the outbound URL via the tool's metadata passthrough (e.g.
distinct_id). One document-level listener covers every CTA — no per-link edits.?metadata[ph_distinct_id]=<id> - The third-party tool stores that metadata and returns it in its webhook.
- Your webhook handler fires an identity merge (with the booking email as
$identifyand the smuggled anonymous id asdistinct_id) plus a conversion event — joining the booking back onto the marketing journey.$anon_distinct_id
唯一需要复杂机制的场景:转化在非自有域名上完成(如预订工具、托管结账页面)。你无法在那里运行分析工具,因此:
- 点击时,捕获阶段的链接装饰器通过工具的元数据传递功能,将访客的匿名附加到出站URL(例如
distinct_id)。一个文档级监听器可以覆盖所有CTA——无需逐个编辑链接。?metadata[ph_distinct_id]=<id> - 第三方工具存储该元数据,并在其Webhook中返回。
- 你的Webhook处理器触发身份合并(接口,将预订邮箱作为
$identify,将传递的匿名ID作为distinct_id)以及转化事件——将预订记录与营销旅程关联起来。$anon_distinct_id
Guardrails (do not skip)
防护规则(请勿跳过)
- Anonymity guard — fail closed. Only ever smuggle the anonymous id. After , the current id becomes the user's email/UUID; leaking that into a third-party URL or merging on it corrupts profiles (person A's email folds into whoever books). Reject ids that look like PII (contain
identify()), cap length, and when identity is ambiguous, send nothing. If the app identifies by UUID, test@rather than andistinct_id === device_idcheck.@ - First-touch data quality. Redirects overwrite the true first touch. Exclude OAuth/checkout referrers (,
accounts.google.com,checkout.stripe.com), your own subdomains (self-referrals), and dev hosts (login.*) from referrer classification. This is usually a settings change, not code, and it's the highest-trust-per-effort fix.localhost - Cross-subdomain stitching. Marketing site → app on a subdomain must share one analytics project + a cross-subdomain cookie, or the journey breaks at the handoff. Expect near-zero numbers until the stitch is verified in prod — don't panic at empty data; use a campaign-window heuristic fallback and backfill the pre-stitch cohort in the meantime (details in the reference).
- 匿名防护——默认拒绝。仅传递匿名ID。调用后,当前ID变为用户的邮箱/UUID;将其泄露到第三方URL或基于它进行合并会破坏用户档案(用户A的邮箱会与任何预订者的档案合并)。拒绝看起来像PII的ID(包含
identify()),限制长度,当身份不明确时,不传递任何数据。如果应用通过UUID识别用户,测试@而非检查distinct_id === device_id。@ - 首次触点数据质量。重定向会覆盖真实的首次触点。将OAuth/结账引荐来源(、
accounts.google.com、checkout.stripe.com)、自有子域名(自引荐)和开发主机(login.*)排除在引荐来源分类之外。这通常是设置变更,而非代码修改,是投入产出比最高的信任度修复措施。localhost - 跨子域名关联。营销网站→子域名上的应用必须共享同一个分析项目+跨子域名Cookie,否则用户旅程在切换时会中断。在生产环境验证关联前,数据量可能几乎为零——不要因空数据惊慌;可以使用活动窗口启发式方法作为 fallback,并同时回填关联前的用户群数据(详见参考文档)。
Reporting and the last mile
报告与最后一公里
The first payoff is one insight: your conversion event broken down by first-touch channel ( / ), and — joined to revenue — channel → conversion → revenue. Confirm first-touch vs. last-touch config in the tool (many default to last-touch; first-party attribution wants ).
$initial_utm_source$initial_referring_domain$initial_*But first-touch alone can't run the multi-touch models from §2. Store the full ordered touch path (not just ) and the build track feeds the interpretation track — you can score your own journeys position-based / linear / time-decay instead of only reading about them.
$initial_*The last mile — get it into the CRM (production refinement from Tessa Kriesel). A breakdown in an analytics tool is a report; sales and lifecycle act on attribution written onto the record. Sync a field with and (journey-linked vs self-reported vs campaign-window fallback) plus a Paid-vs-Organic read off the medium, rolled up to the account (not just the contact — one B2B org is several people with mixed work/personal emails). How pipeline/lifecycle then use it is revops' job.
sourceconfidencebasisThe pattern is tool-agnostic: identify + merge exists in PostHog, Segment, Amplitude, and via user-id in GA4; the third-party stitch works with any tool that has a metadata passthrough + webhook. PostHog + SavvyCal are the worked example in .
references/first-party-tracking.md第一个成果是一个关键洞察:按首次触点渠道( / )拆分的转化事件,以及关联营收后的渠道→转化→营收数据。要确认工具中的首次触点 vs 末次触点配置(许多工具默认使用末次触点;第一方归因需要字段)。
$initial_utm_source$initial_referring_domain$initial_*但仅靠首次触点无法运行第2节中的多触点模型。存储完整的有序触点路径(不仅是),自建路线的数据可以为解读路线提供支持——你可以自行对用户旅程进行位置型/线性/时间衰减评分,而非仅依赖工具报告。
$initial_*最后一公里——同步到CRM(来自Tessa Kriesel的生产优化建议)。分析工具中的细分是报告;销售和生命周期运营会基于写入用户档案的归因数据采取行动。同步一个**字段,包含和(旅程关联 vs 自报 vs 活动窗口 fallback),以及基于媒介的付费vs organic(自然)分类,汇总到账户级别(不仅是联系人——一个B2B组织包含多个使用混合工作/个人邮箱的用户)。管线/生命周期运营如何使用这些数据是revops(营收运营)**的职责。
sourceconfidencebasis该模式与工具无关:identify+合并功能存在于PostHog、Segment、Amplitude中,GA4也支持通过user-id实现;第三方关联适用于任何支持元数据传递+Webhook的工具。中包含PostHog + SavvyCal的示例。
references/first-party-tracking.mdOutput format
输出格式
Deliver an attribution readout, not a data dump:
markdown
undefined交付归因报告,而非数据转储:
markdown
undefinedAttribution Readout — [date]
归因报告 — [日期]
The question
核心问题
[What decision this informs — e.g. "where should next quarter's budget go?"]
[报告用于支持的决策——例如“下一季度的预算应分配到哪里?”]
Source of truth
真实数据源
[Which system defines the conversion count, and why]
[定义转化数量的系统,以及原因]
What each source says
各数据源数据
| Channel | Platform-reported | GA | CRM | Self-reported | Our read |
|---|---|---|---|---|---|
| [De-duped against source of truth; not summed] |
| 渠道 | 平台报告 | GA | CRM | 自报数据 | 我们的解读 |
|---|---|---|---|---|---|
| [基于真实数据源去重;不跨平台求和] |
Model comparison (for long cycles)
模型对比(长周期场景)
[First-touch vs last-touch side by side; the gap is the insight]
[首次触点 vs 末次触点数据对比;差距即为洞察点]
Confidence & gaps
置信度与缺口
[The attribution gap, the blind spots, what we can't see]
[归因缺口、盲区、无法追踪的内容]
Recommendation
建议
[Allocation call with confidence levels; the tiebreaker test worth running]
undefined[带有置信度的预算分配建议;值得运行的判定测试]
undefinedTool Integrations
工具集成
For implementation, see the tools registry. Key tools:
| Tool | Best For | MCP | Guide |
|---|---|---|---|
| PostHog | First-party attribution, identify/merge, funnels | - | posthog.md |
| GA4 | Web analytics, model comparison, user-id stitching | ✓ | ga4.md |
| Dub | Short-link + click attribution | ✓ | dub-co.md |
| Segment | CDP — route identify/track to every destination | - | segment.md |
| HubSpot | CRM lead-source + self-reported fields | ✓ | hubspot.md |
| Salesforce | CRM as revenue source of truth | - | salesforce.md |
| Supermetrics | Pull platform numbers into one place to reconcile | ✓ | supermetrics.md |
| RB2B | De-anonymize B2B website visitors | - | rb2b.md |
实施相关内容,请查看工具注册表。核心工具:
| 工具 | 最佳适用场景 | MCP | 指南 |
|---|---|---|---|
| PostHog | 第一方归因、身份识别/合并、漏斗分析 | - | posthog.md |
| GA4 | 网页分析、模型对比、用户ID关联 | ✓ | ga4.md |
| Dub | 短链接+点击归因 | ✓ | dub-co.md |
| Segment | CDP——将identify/track数据路由到所有目标系统 | - | segment.md |
| HubSpot | CRM线索来源+自报字段 | ✓ | hubspot.md |
| Salesforce | 作为营收真实数据源的CRM | - | salesforce.md |
| Supermetrics | 将平台数据汇总到一处进行调和 | ✓ | supermetrics.md |
| RB2B | 识别B2B网站访客身份 | - | rb2b.md |
Related Skills
相关技能
- analytics — event tracking, tracking plans, UTMs, GA4/GTM setup. Do this before attribution.
- ads — ad-platform pixels, CAPI, server-side conversion tracking ().
references/conversion-tracking.md - revops — pipeline stages, lead lifecycle, CRM revenue reporting. Attribution feeds it.
- ai-seo — the AI-search attribution blind spot in depth.
- ab-testing — controlled experiments; the incrementality mindset applied to on-site changes.
- analytics(分析)——事件追踪、追踪计划、UTM设置、GA4/GTM部署。应在归因前完成。
- ads(广告)——广告平台像素、CAPI、服务器端转化追踪()。
references/conversion-tracking.md - revops(营收运营)——销售管线阶段、线索生命周期、CRM营收报告。归因向其提供数据。
- ai-seo(AI搜索引擎优化)——深入分析AI搜索归因盲区。
- ab-testing(A/B测试)——对照实验;增量思维在网站优化中的应用。