benchmark-methodology

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Benchmark Methodology

基准分析方法论

Use this skill to turn a scoped competitor set into comparable, defensible scores. Each competitor is assessed on the same nine dimensions, with explicit 1–5 rubrics, then captured in a uniform profile card. Consistency is the point: scores are only useful if the same evidence would earn the same number for any competitor.
使用该技能可将已划定范围的竞品集合转化为具备可比性、可辩护的评分。每个竞品都将基于相同的九个维度,通过明确的1-5分评分标准进行评估,之后被纳入统一的竞品档案卡片。核心在于一致性:只有当相同证据能为任意竞品带来相同分数时,评分才具备参考价值。

When to Activate

激活时机

  • A scoped, tiered competitor set from competitive-platform-analysis is ready to score.
  • Need comparable, evidence-anchored scores across competitors — not gut-feel rankings.
  • Client's strategic tension (the paired axes defining their target white-space) has been established.
  • Preparing to produce profile cards for assembly in competitive-report-structure.
  • 已通过competitive-platform-analysis完成竞品范围划定与分层,可进行评分。
  • 需要基于证据的、具备可比性的竞品评分,而非主观直觉排名。
  • 已明确客户的战略张力(定义其目标空白市场的成对轴)。
  • 准备生成竞品档案卡片,用于后续组装进competitive-report-structure。

Client positioning brief (establish first)

客户定位简报(需先确立)

Before scoring, establish the client's positioning brief. It supplies:
  • Strategic tension — the two axes (e.g., memorability × hireability) whose intersection marks the client's target white-space. Dimension 9 is always the client's named tension; report both poles separately, never averaged.
  • Differentiator — what makes the client's moat. This informs which dimensions matter most for the client's positioning argument.
  • Brand balance — the intended mix of distinct strategic emphases. Strategic recommendations must not break this balance without flagging it.
开始评分前,需先确立客户的定位简报,它将提供:
  • 战略张力——两个轴(例如:记忆度×可雇佣性),其交叉点即为客户的目标空白市场。第9维度始终为客户指定的张力;需分别报告两个极端,绝不取平均值。
  • 差异化优势——客户的核心护城河所在。这将决定哪些维度对客户的定位论证最为重要。
  • 品牌平衡——预期的不同战略侧重点组合。战略建议不得打破这种平衡,除非提前标注说明。

Why these dimensions

为何选择这些维度

The client competes on a specific tension held across two poles, not on service breadth. The dimensions are weighted to reflect that moat. Two dimensions — the tension poles — are scored separately and never averaged together, because the client's strategic question is precisely whether a rival achieves both simultaneously.
客户的竞争点在于横跨两个极端的特定张力,而非服务广度。维度权重的设置正是为了反映这一护城河。其中两个维度——张力的两个极端——需分别评分,绝不取平均值,因为客户的战略问题恰恰在于竞争对手是否能同时达成两个极端的要求。

The nine dimensions (with weights)

九个维度(含权重)

Weights guide synthesis emphasis, not a single blended score (avoid a false composite — see Bias controls). Sum = 100%.
  1. Positioning clarity & distinctiveness (18%) — Is the studio's position sharp, ownable, and instantly legible? Or generic?
  2. Brand voice / verbal distinctiveness (15%) — Does the copy have an ownable register, or is it interchangeable agency-speak?
  3. Visual identity & site craft (15%) — Quality and ownership of the visual system; site as proof-of-craft.
  4. Service offer & packaging (12%) — Productized and legible (named sprints/audits) vs vague. Packaging maturity.
  5. Evidence & credibility (12%) — Named clients, quantified outcomes, case-study depth. Proof beyond assertion.
  6. Enterprise-readiness / commercial maturity (10%) — Signals they can land and hold SaaS/fintech/B2B/enterprise work (process, logos, scale, contracts).
  7. Thought leadership / content presence (8%) — Owned POV: writing, talks, newsletters, frameworks. Depth over volume.
  8. Pricing transparency & engagement model (5%) — Is pricing/engagement legible? Productized vs bespoke vs opaque.
  9. [Client's strategic tension] (5% as a flag; score BOTH poles, report separately) — Read the tension name and axis descriptions from the client's positioning brief. Plot both; the gap is the insight. The client's target quadrant is the single most important finding: who else is already there?
权重用于指导综合分析的侧重点,而非生成单一综合得分(避免虚假的复合得分——详见偏差控制)。权重总和=100%。
  1. 定位清晰度与独特性(18%)——工作室的定位是否清晰、可独占且易于理解?还是过于通用?
  2. 品牌语气/语言独特性(15%)——文案是否具备独特风格,还是千篇一律的代理行话术?
  3. 视觉标识与网站工艺(15%)——视觉系统的质量与独占性;网站作为工艺实力的证明。
  4. 服务内容与包装(12%)——是否产品化且清晰易懂(如命名明确的冲刺项目/审计服务),还是模糊不清?包装成熟度如何?
  5. 可信度证明(12%)——知名客户、量化成果、案例研究深度。除主张外的实质性证据。
  6. 企业就绪度/商业成熟度(10%)——是否具备承接并维护SaaS/金融科技/B2B/企业级业务的信号(流程、客户标识、规模、合同)。
  7. 思想领导力/内容影响力(8%)——专属观点:文章、演讲、通讯、框架。重深度而非数量。
  8. 定价透明度与合作模式(5%)——定价/合作模式是否清晰?是产品化、定制化还是不透明?
  9. [客户战略张力](5%,作为标记;需对两个极端分别评分并报告)——从客户定位简报中读取张力名称及轴描述。对两个极端分别评分并绘图;两者间的差距即为洞察点。客户的目标象限是最重要的发现:还有哪些竞品已占据该象限?

Scoring rubric (1–5, applies to dimensions 1–8)

评分标准(1-5分,适用于第1-8维度)

Anchor every score to observable evidence. Generic descriptors below; adapt the specifics per dimension but keep the level meaning constant.
  • 1 — Absent / generic. No discernible position or craft; indistinguishable from a template. Active liability.
  • 2 — Below par. Some intent but inconsistent, derivative, or unconvincing. Wouldn't survive a side-by-side.
  • 3 — Competent / table-stakes. Solid, professional, unremarkable. Meets expectation, ownable by nobody.
  • 4 — Strong / distinctive. Clearly above peers; a real strength a buyer would notice and cite.
  • 5 — Category-defining. Best-in-class, ownable, hard to imitate. Sets the bar others react to.
每个评分都必须基于可观察的证据。以下为通用描述;需根据各维度调整具体细节,但保持评分等级的含义一致。
  • 1分——缺失/通用:无明确可辨的定位或工艺;与模板无异。属于明显劣势。
  • 2分——低于标准:有一定意图但不一致、模仿性强或缺乏说服力。无法在对比中胜出。
  • 3分——合格/基础要求:扎实、专业但无亮点。符合预期,但无独占性。
  • 4分——优秀/独特:明显优于同行;是买家会注意并提及的真实优势。
  • 5分——定义品类:行业顶尖、可独占、难以模仿。为同行设立标杆。

Tension axes (dimension 9) — score each 1–5

张力轴(第9维度)——各轴分别按1-5分评分

Read the axis labels and their 1/3/5 anchors from the client's positioning brief. Example anchors for a memorability × credibility tension:
  • Memorability — 1: forgotten instantly · 3: recognizable in context · 5: unforgettable, talked-about, distinctively owned.
  • Credibility — 1: feels risky/amateur · 3: safe, competent, unexciting · 5: enterprise-trusted, obvious safe choice.
Plot competitors on the tension 2×2. The client's target quadrant is named in the positioning brief. Who else occupies that quadrant is the single most important finding of the benchmark.
从客户定位简报中读取轴标签及其1/3/5分的锚点示例。例如,针对记忆度×可信度张力的锚点:
  • 记忆度——1分:瞬间被遗忘 · 3分:在特定场景下可识别 · 5分:令人难忘、被广泛讨论、具备独特独占性。
  • 可信度——1分:感觉风险高/业余 · 3分:安全、合格但平淡 · 5分:获企业信任,是明确的安全选择。
将竞品绘制在张力2×2图中。客户的目标象限已在定位简报中注明。哪些竞品已占据该象限是本次基准分析最重要的发现。

How to collect the data

数据收集方式

For each competitor, work the dimensions in this order (cheapest signal first):
  1. Competitor's own site — positioning, voice, offer packaging, pricing posture, named clients, manifesto/POV. Screenshot the homepage + one case study.
  2. Case studies / work — evidence depth, quantified outcomes, client names. Distinguish asserted ("we delivered X") from proven (metrics, named, verifiable).
  3. Review directories — corroborate clients, project size, engagement model → credibility & enterprise-readiness (e.g. Clutch.co or the niche equivalent).
  4. LinkedIn — team size/model, founder narrative, content cadence → thought leadership, model.
  5. Portfolio / craft platforms — craft register (use the showcase native to the niche: design boards, showreels, published samples, etc.).
  6. Content channels — newsletter/talks/writing → thought-leadership depth.
What to record per dimension: the score, one-line justification, and the source link/screenshot that earned it. No score without evidence.
针对每个竞品,按以下顺序分析各维度(优先获取成本最低的信号):
  1. 竞品自有网站——定位、品牌语气、服务包装、定价姿态、知名客户、宣言/观点。截取首页+1个案例研究的截图。
  2. 案例研究/作品——证据深度、量化成果、客户名称。区分主张性内容(“我们交付了X”)与已验证内容(有数据、知名客户、可核实)。
  3. 评论目录——核实客户、项目规模、合作模式 → 可信度与企业就绪度(如Clutch.co或细分领域的类似平台)。
  4. LinkedIn——团队规模/模式、创始人叙事、内容发布频率 → 思想领导力、运营模式。
  5. 作品集/工艺平台——工艺水准(使用细分领域的原生展示平台:设计板、展示视频、已发布样本等)。
  6. 内容渠道——通讯/演讲/文章 → 思想领导力深度。
每个维度需记录的内容:评分、一行理由、支撑该评分的来源链接/截图。无证据则不评分。

Bias controls

偏差控制

  • No single composite score. Report dimension scores and the tension plot separately. A weighted average hides the asymmetry that matters.
  • Asserted vs proven. Downgrade credibility/evidence scores for self-reported claims with no corroboration. Site copy is marketing, not fact.
  • Aesthetic affinity bias. Reviewers may over-score studios whose aesthetic they share and under-score rivals' commercial strength. Score craft and credibility independently; a "boring" site may be winning bigger clients.
  • Recency / flashiness bias. Award-winning, showpiece work dazzles but may lack commercial depth — verify with directories/clients before scoring credibility.
  • Survivorship. The visible, well-marketed studios aren't the whole market; note strong-but-quiet operators found via directories/reviews.
  • Calibrate across the set, not in isolation. Before finalizing, re-read scores side-by-side — a "4" must mean the same thing for every competitor. Adjust outliers.
  • 不生成单一综合得分:分别报告维度评分和张力图。加权平均值会掩盖关键的不对称性。
  • 区分主张与验证:对于无佐证的自我声明,降低其可信度/证明维度的评分。网站文案属于营销内容,而非事实。
  • 审美偏好偏差:评审者可能会给与自身审美一致的工作室高分,而低估竞品的商业实力。需独立评分工艺与可信度;“乏味”的网站可能正在赢得更大的客户。
  • 近期/炫技偏差:获奖的展示性作品令人印象深刻,但可能缺乏商业深度——在评分可信度前需通过目录/客户进行核实。
  • 幸存者偏差:可见的、营销良好的工作室并非全部市场;需注意通过目录/评论发现的实力强劲但低调的运营商。
  • 跨集合校准,而非孤立评分:最终确定前,需并排重新审阅所有评分——“4分”对每个竞品的含义必须一致。调整异常值。

Competitor profile card (output format)

竞品档案卡片(输出格式)

Produce one card per profiled competitor — the atomic unit the report assembles from:
undefined
为每个被分析的竞品生成一张卡片——这是后续报告组装的基本单元:
undefined

<Competitor name>

<Competitor name>

  • Profile / Tier: <positioning stance · specialization · size band> / <Direct | Adjacent | Aspirational>
  • One-liner: <how they position themselves, in their words>
  • Model / size / geography: <solo|micro|boutique> · <region> · <pricing/engagement model>
  • Notable clients / evidence: <named, with proven/asserted tag>
  • Profile / Tier: <positioning stance · specialization · size band> / <Direct | Adjacent | Aspirational>
  • One-liner: <how they position themselves, in their words>
  • Model / size / geography: <solo|micro|boutique> · <region> · <pricing/engagement model>
  • Notable clients / evidence: <named, with proven/asserted tag>

Dimension scores

Dimension scores

DimensionScore (1–5)Justification (1 line)Source
Positioning clarity & distinctiveness
Brand voice / verbal distinctiveness
Visual identity & site craft
Service offer & packaging
Evidence & credibility
Enterprise-readiness / commercial maturity
Thought leadership / content presence
Pricing transparency & engagement model
DimensionScore (1–5)Justification (1 line)Source
Positioning clarity & distinctiveness
Brand voice / verbal distinctiveness
Visual identity & site craft
Service offer & packaging
Evidence & credibility
Enterprise-readiness / commercial maturity
Thought leadership / content presence
Pricing transparency & engagement model

Tension plot

Tension plot

  • [Axis 1 from positioning brief]: <1–5> — <why>
  • [Axis 2 from positioning brief]: <1–5> — <why>
  • Quadrant: <high/high | high-1/low-2 | low-1/high-2 | low/low>
  • [Axis 1 from positioning brief]: <1–5> — <why>
  • [Axis 2 from positioning brief]: <1–5> — <why>
  • Quadrant: <high/high | high-1/low-2 | low-1/high-2 | low/low>

Read for [client]

Read for [client]

  • Strength to learn from: <…>
  • Weakness to exploit / white-space it exposes: <…>
  • Threat to [client]: <…>

Hand the completed cards plus the tension plot to `competitive-report-structure`.
  • Strength to learn from: <…>
  • Weakness to exploit / white-space it exposes: <…>
  • Threat to [client]: <…>

将完成的卡片及张力图提交给`competitive-report-structure`。

Anti-Patterns

反模式

  • Averaging the tension axes. The two poles of the client's strategic tension must be scored and reported separately. Averaging destroys the insight — the gap between poles is the finding.
  • Scoring without evidence. Every score requires a one-line justification and a source link. A score without evidence is an opinion, not a benchmark.
  • Creating a single composite score. Report dimension scores individually. A weighted average hides the asymmetric strengths that matter for positioning.
  • Applying generic rubric anchors without adapting. The 1–5 anchors must be calibrated to the specific dimension and competitor set. The generic descriptions are a starting point, not a fixed standard.
  • Running before the competitor set is scoped. Use competitive-platform-analysis first to produce a tiered, pruned set. Scoring an unscoped list wastes effort on irrelevant competitors.
  • 对张力轴取平均值:客户战略张力的两个极端必须分别评分和报告。取平均值会摧毁洞察价值——两个极端间的差距才是关键发现。
  • 无证据评分:每个评分都需要一行理由和来源链接。无证据的评分只是观点,而非基准分析结果。
  • 生成单一综合得分:需单独报告各维度评分。加权平均值会掩盖对定位至关重要的不对称优势。
  • 直接套用通用评分锚点而不调整:1-5分的锚点必须针对特定维度和竞品集合进行校准。通用描述只是起点,而非固定标准。
  • 未划定竞品范围便开始评分:需先使用competitive-platform-analysis生成分层、精简的竞品集合。对未划定范围的列表进行评分会在无关竞品上浪费精力。

Related Skills

相关技能

  • competitive-platform-analysis
    — the prerequisite; produces the tiered competitor set this skill scores.
  • competitive-report-structure
    — the next step; assembles the scored profile cards into a client-deliverable report.
  • competitive-platform-analysis
    ——前置技能;生成本技能所需的分层竞品集合。
  • competitive-report-structure
    ——后续步骤;将已评分的竞品档案卡片组装成可交付给客户的报告。