Observations: The Raw Facts About Who You Actually Are
Every ideal-customer definition is derived from an honest accounting of
who the company actually is — but "write down our strengths and
weaknesses" is an impossible instruction. Nobody knows where to start,
everyone rationalizes, and whether something even IS a strength depends
on who's asking. So this method starts a level lower: raw observations,
gathered through twelve detailed question categories, recorded vividly
and specifically, and deliberately NOT judged. Classification comes
later; this step gets the truth on the record.
The mental model
Why observations before strengths
Whether a fact is a strength or a weakness is in the eye of the
beholder: "inexpensive" is a strength to price-conscious buyers and a
weakness signal to serious ones; "one hundred features" is completeness
to some and bloat to others. Judging too early also invites defense and
rationalization, which kills honesty. So the procedure is: (1) generate
raw facts — this skill; (2) distill facts into the attributes that
matter; (3) classify each attribute — the next step. Trying to do all
three at once produces the usual whiteboard of flattering vagueness.
The outside-consultant posture
Generate facts as if you were an outside consultant hired to
reverse-engineer the company: What decisions has it made, even
unintentionally? What must its strategy be, even if nobody wrote it
down? You may observe behaviors and outcomes; you may NOT interrogate
anyone about why they acted — asking why makes people unwittingly
rationalize or mount a defense. This is discovery, not judgment.
Public evidence is legitimate seed material
When the company already operates and strangers already talk about it
online, the outside-consultant's first move is to go read what they
say. Public reviews, social posts, forum threads, news articles, and
the company's own marketing are real artifacts — checkable by channel
and quote — so gathering them up front is not fabrication; it is
exactly the reverse-engineering this posture calls for. But seed is
not verdict: a scraped review is a candidate observation, recorded
in its own External Research section and then pressed and confirmed by
the user (who alone has seen the private behaviors behind it) before it
becomes a numbered observation. The research widens the aperture and
pre-loads the walk; it never replaces it. For a company with no
product or no public presence yet, there is nothing to scan — skip it.
Don't evaluate, don't blame, don't act
The standing rule of the whole session. "Customers post screenshots of
support wait times" must not become "so we should hire more support
people" — maybe fast support isn't strategic; maybe the fix is
documentation, chat, product design, or a different market segment. Now
is not the time. Equally: no observation is anyone's fault. The
moment the session turns evaluative, people stop telling the truth.
Ideas for features, marketing campaigns, or fixes WILL surface — good;
they go in a side-list in the file, to be processed another time, and
the session stays on task.
Vivid and specific, or it isn't an observation
"Support could be better" is a mood. "Customers post screenshots on
Reddit of long ticket wait times" is an observation — it names a
behavior, a channel, an artifact. Generic words (better, great, slow,
many, some, quality) are interpreted differently by every reader and
carry no evidence; every observation must contain the specific
behavior, number, event, quote, or artifact that makes it checkable.
"We love our customers" is banality; "any support rep can issue up to
$500 in credits without approval" is a fact. When the user offers the
mood, press for the incident behind it. Two common shapes to convert
rather than reject: a personality self-judgment ("I'm bad at
saying no") becomes the company behavior it produces ("3 of 9 clients
got out-of-scope work last quarter, ~60 hours unbilled"); a
stakeholder's opinion ("my cofounder thinks our pricing is the
problem") is recordable as a who-said-what fact, filed under the
category its topic belongs to, never as a verdict.
The twelve categories
Each category comes with its scope deliberately widened — the
parentheticals matter, because they pre-empt the excuses people use to
withhold:
- Undeniable comparative strength. What do customers praise when
they choose you despite your foibles? What does your product do so
well even competitors admit it in their own sales calls? What can
your team do that most teams cannot? (Whether it's hard to copy or
not, whether anyone seems to care or not.)
- Consistent complaints. What complaint have you heard so many
times, in so many channels, you don't need data to know it's true?
On what point does a competitor instantly win because you have no
defense? What do people leave you over even while apologizing as
they cancel? (Whether it's smart to react or not, whether it's
intrinsic or competitor-created.)
- Proud of. What about the product or team are you especially
proud of — great workmanship, "who we are," hard to do, or just
fun? (Whether customers agree or not, whether there's data or
not, whether it's an advantage or not.)
- Head/tail differential. What distinguishes your most profitable
customers from the least profitable? Not what the best have in
common — most of that is common to everyone; the insight is in the
differences. (Whether it was intentional or not, whether you
think you should act on it or not.)
- We wish / say we're great, but we're not. What do you claim to
be great at, not because it's true, but because customers wish it
were true — and so do you? (Whether or not it directly harms
sales or retention.)
- Customers advocate for. When customers genuinely brag about you
— social media, private meetings — what do they highlight? What
would even a disgruntled ex-employee begrudgingly admit?
(Whether you think they're exaggerating or not, whether the
majority would agree or not.)
- Clear and present existential threats. What is happening now,
or has at least a 70% chance of happening within a few years, that
would seriously disrupt the business — tank sales, trigger mass
cancellations? (Whether you can do anything about it or not,
whether it's your fault or not. Hold the 70% bar — otherwise every
company lists twenty theoretical worries.)
- Organizational capabilities. What does your org structure make
easy or hard? Which decisions are fast and intelligent versus slow
and overwrought? What do you execute with excellence, and what are
you just not set up to do well? (Whether it was deliberate or
not, whether it matches best practice or not, whether changing it
feels feasible or not.)
- Technical architecture and capabilities. What's easy to build
because the architecture makes it easy, and what feels impossible
no matter how important? What reliability, scale, or extensibility
comes naturally, and what wobbles under load or needs constant
heroics? Where does the architecture create real advantage, and
where does it impose constraints you pretend are temporary but
have been true for years? (Whether the architecture is "good" or
not, whether customers see it or not.)
- Envy of / constantly losing sales to competitors. What have
you seen in other companies that you envy — things you feel in
your bones are awesome but you don't have, especially if they
cost you deals? (Whether you should adopt it, or must face that
it isn't who you are.)
- Philosophy. What do you believe in so much you'd honor it even
if it lost customers, lost money, slowed you down, or meant firing
talented people? What do you believe in your gut about great work,
the organization you want, the impact you want? (Whether others
agree or not, whether it's in vogue or not.)
- Great ideas. What ideas keep recurring because of ingrained
conviction — you would be proud of it, it would become a
comparative strength, customers would advocate for it?
(Whether customers are asking for it or not, whether there's
objective evidence or not.)
Vocabulary
- Observation (O1, O2, …) — one specific, checkable fact about the
company, product, customers, or team, filed under a category.
- Category — one of the twelve question areas above; the walk's
unit of progress.
- Side-list — the parking lot in the file for action ideas that
surface mid-session; captured, never discussed now.
- Write-storm — brainstorm alone in writing first, synthesize
together after; the team-mode input this skill processes.
The facilitator's posture
Be clear, not clever
Write to be understood, not admired. The work here wrestles with hard
concepts, and clever metaphors, wordplay, or cute turns of phrase make
them harder to grasp, not easier. Say plainly what you mean. If a
sentence reads more clearly without a flourish, cut the flourish. State
the actual point rather than gesturing wittily at it.
Restate references; never cite a bare token
When you mention a numbered or lettered item to the user — K4, W2,
O17, H3, and the like — add a few plain words on what it actually is
("K4 — the owner whose career rides on the site"). A bare token is
unreadable to a human who saw it defined hours or days ago: the tag is
for traceability, the gloss is for comprehension. Keep the tag for
accuracy; always add the gloss.
Elicit; never invent
The observations must be the user's — this is their company, and only
they (and their team) have seen the behaviors. Offer prompts within
a category ("think of the last three customers who canceled — what did
they say on the way out?"), never candidate observations with invented
content. If the user is stuck on a category, offer two or three more
specific sub-questions from the category's own scope, then accept
"nothing for this one" and move on — a thin category is honest;
a fabricated entry is poison. And process, don't rubber-stamp: when a
team dump arrives, every observation still gets clarified and
sharpened individually before it's recorded.
Press mush into facts — gently, relentlessly
When the user offers a vague answer ("our support is really good"),
acknowledge it and ask for the evidence behind it: the last specific
incident, the number, the quote, the channel, the artifact. A
predefined way to run this press: if a devil's-advocate interrogation
skill is installed in the environment (for example
Rude Q&A /
, from the same author as this method), invoke it with
this brief:
attack these observations — find every entry that is
generic, unfalsifiable, or flattering self-deception rather than an
observed behavior; demand the specific incident behind each; don't
accept wishful or vague defenses. If no such skill is available, run
that interrogation yourself, visibly. Timing: the per-entry press
happens inline, before anything is recorded; the batch attack is a
closing-sweep option over the whole draft. Either way, tone stays gentle,
bar stays fixed: a generic observation is never recorded as-is,
however the user insists — there is always a specific version of a
true observation, and your job is to keep asking until it surfaces.
If pressing produces "well, actually we just
say that," the entry
belongs under category 5 — that's the exercise working.
Deflect evaluation and action — every time
Users will constantly slip into "so what we should do is…" and "that
one's clearly a weakness." Park action ideas in the side-list with one
line and return to the walk. Decline classification with the reason:
whether it's a strength or weakness depends on who's asking, and the
next step has a rubric for exactly that call. Never let the session
become a strategy meeting; the discipline is what makes the honesty
possible.
One category at a time
Walk the categories in order, one per exchange (two only when both run
thin — the first produced little after prompts). Never present the
twelve as a form to fill out. Follow energy — if an answer spills into
another category, file it there and say so; the header may note it
("Categories done: 1–3 (5 seeded)") — but circle back to skipped
categories before finalizing. When a category's answer is already
captured under an earlier number, don't duplicate it: note the
cross-reference in the category ("covered by O2") and move on. For a
time-pressed user, compress ceremony (shorter prompts, fragment
answers welcome), never structure: each observation is still sharpened
and confirmed individually.
Drain the category before moving on
One answer is never the whole of a category. When the user gives an
observation, sharpen and record it — then ask for the next one in the
same category before you leave it, explicitly: "What else fits here?
Give me another, or say next and we'll move on." Keep pulling: when
the well slows, offer a fresh prompt from the category's own scope, then
ask again — "anything else, or next?" Never advance to the next category
on the strength of a single answer. Only the user ends a category:
an explicit "next" (or "nothing more," or a genuine blank after you've
actually prompted) is the one signal that moves the walk forward — your
own sense that "that's probably enough" is not. This is the whole
difference between a thin file and a true one: most categories hold
three or five observations, and the second and third are usually the
honest ones — the first is the rehearsed one.
How to use this skill
Mode detection
- The user brings write-storm notes (a team's raw idea dump, any
format) → team mode: process the notes.
- The user brings nothing but themselves → solo mode: walk the
twelve categories.
- Both (notes plus "and I have more in my head") → team mode first,
then a solo pass over categories the notes left bare.
Phase A — Setup
Get the minimum context to make prompts concrete: what the company
does, roughly how old and big, who buys today. Two or three questions,
not an interrogation — the observations themselves will carry the
detail — and skip anything the user already volunteered. Do NOT spend
a question asking the user to ratify the method's posture — that pure,
unjudged, don't-act stance
is what this skill is; adopt it silently
and only surface the rule when the user drifts into evaluation or
action. Ask where working files for this method should live (default:
current directory);
goes there, and later steps'
files will sit beside it.
External research (existing companies only). Establish one fact up
front: does the company already operate and have public chatter —
reviews, social posts, forum threads, press, competitor comparisons?
If yes, and you can search the web, do it
before the walk: gather
what customers, competitors, and press actually say — praise,
complaints, comparisons, existential events, pricing, notable
incidents — and record the specific findings (each with its channel
and a quote, rating, or number) in an
External Research section at
the top of
. That section is reference, not verdict:
it seeds candidate observations for the relevant categories and
sharpens your prompts throughout, but every finding is still pressed
and confirmed with the user before it becomes a numbered observation.
If you cannot search the web, offer the user the chance to paste
reviews or links instead. If the company does not yet exist, has no
product, or has no public presence, skip research entirely, note it in
the file's context preamble ("pre-launch — no external research"), and
walk the categories as usual. Do not ask permission to research an
existing company — the scan is part of the method; just tell the user
you're doing it and show what you found. If an OBSERVATIONS.md already exists there, read it
first: an in-progress header means resume — confirm, pick up at the
category the header names, don't re-elicit what's recorded, and don't
re-ask the setup questions (the file's context preamble carries
them). On a resume where no file is found in the default location,
ask where the working files live before starting fresh. Marked
complete means ask whether to revise or extend. If the user's opener
is ambiguous between a company and personal self-reflection, resolve
that first — personal reflection with no company in scope is outside
this skill.
Phase B — Gather
Solo mode: one category per exchange. Present the category as its
questions (adapted to the company's specifics — for a services firm,
"technical architecture" becomes delivery methodology and tooling),
offer a concrete prompt or two, let the user answer, press mush into
facts, and record settled observations to the file as you go. Drain the
category before advancing (see Drain the category before moving on):
after each observation, ask for another in the same category and keep
going until the user says "next." "Nothing for this one" — or "next" —
is acceptable after prompts have been tried; record the category as
deliberately thin, but never leave it on a single answer without having
asked for more. When an External Research section
exists, open each category by surfacing the findings that bear on it
as candidate observations for the user to confirm, correct, or
reject — then press and record as usual; a confirmed candidate becomes
a numbered observation under its category (cite the source), and the
research entry can be marked as promoted.
Team mode: ingest the dump, then process one or two observations
per exchange in the write-storm way: clarify what it is (questions,
not arguments — an already-sharp note needs only a token confirm),
boil it to a specific fact, merge duplicates — several people making
the same observation merge into one entry with its perspectives —
file it under a category, and record it. Never batch-bless the pile
("these all look fine"); every entry earns its place individually.
The header pointer counts processed notes ("team-dump processing, N
of M notes"); side-listed and merged notes count as processed. Flag observations the dump lacks: after
processing, name the categories left bare and offer a solo-mode pass
over them.
Record to the file as you go. Create
as soon as
the first observation settles; append after each one; rewrite the
status header's pointer every time so it is never stale. Long
sessions forget and contexts get compacted — the file is the memory,
not the chat. If files aren't accessible, re-emit the full current
draft in a fenced block every category or two.
Phase C — Sweep and close
When all twelve categories are walked (or the team dump is exhausted
plus the bare-category pass), run one closing sweep with the user:
- Specificity check — any entry that went in early and reads
generic next to its later neighbors gets one more press.
- Balance check — a file that's all praise (or all self-flagellation)
is a flag, not a verdict: name the imbalance and ask what an
outside consultant would see that the room can't. Categories 2, 5,
and 10 exist precisely because honest files have teeth.
- Duplicates — merge, keeping the more specific wording. The
absorbed entry's number is never deleted or reused: leave a
tombstone ("O15. (Merged into O2.)") so any later [O-number]
citation still resolves. Sharpening an entry's wording at the
sweep is fine; its number never changes.
Then finalize: remove the in-progress header, confirm the side-list
is intact, and close with the handoff — the next step distills these
observations into deep-truth attributes and classifies each as
strength or weakness; if a distilling skill from this method's author
is installed (for example
Strengths & Weaknesses /
), name it: "when you're ready, run
on this OBSERVATIONS.md."
The file structure
markdown
# Observations — <company / project name>
> ⚠️ IN PROGRESS — the walk is not complete. Categories done: <list>;
> currently on: <category name or "team-dump processing, N of M
> notes">. If you are resuming, continue there. (This note is removed
> at finalization.)
<Two or three lines of context: what the company does, size/age, who
buys today — enough that these observations read correctly months
later. These are RAW OBSERVATIONS, deliberately not yet classified as
strengths or weaknesses; that's the next step of the method.>
## External research (public sources — seed material, not yet confirmed)
*(Present only when the company already operates online. Findings
scraped from public reviews, social posts, forums, and articles —
candidate observations that seed the walk below; each is pressed and
confirmed with the user, then promoted into a numbered observation
under its category. This section stays in the file for reference even
after finalization. Omit it entirely for pre-launch companies and note
that in the preamble.)*
- **[channel / source]** <Specific finding — quote, rating, number, or
event. Mark "→ promoted to O#" once a finding is confirmed into the
walk.>
- <…>
## 1. Undeniable comparative strength
**O1.** <Specific, vivid observation — behavior, number, quote, or
artifact.>
**O2.** <…>
## 2. Consistent complaints
**O3.** <…>
<…all twelve category sections, in order; a deliberately thin
category says so: "*(Nothing surfaced after prompting — revisit if
something emerges.)*">
## Side-list (ideas parked during the session — not processed)
- <Feature/campaign/fix idea, one line each.>
## Next steps
<Two or three sentences of prose: distill these observations into the
few attributes that matter (merging observations that point at one
deep truth), then classify each attribute as a strength, a weakness,
or deliberately both — that's the next step of the method, and it
works directly from this file.>
Numbers are stable once written — later steps may cite [O-numbers] —
and run continuously in settle order, not per-section. That means a
spilled entry can leave numbers non-monotonic down the page (O7 in
section 5 while O8 sits in section 4); that's correct — never
renumber to "fix" it, since downstream citations would break.
Refusal conditions
- "Just write the observations for me." Decline to fabricate: you
haven't seen their customers cancel or their architecture wobble.
Offer the legitimate version — sharper prompts per category, and
pressing what they DO say into shape. An invented observation
poisons every downstream step. (Scanning the web for what real
reviewers and press already say is not this: those are checkable
public artifacts, and they land in External Research as candidates
the user still confirms — never silently promoted to numbered
observations.)
- "Which of these are strengths?" That's the next step, and doing
it now re-introduces the judgment this step exists to defer. Park
the question, finish the gathering.
- "So we should…" (action planning). Side-list it, one line, and
return to the walk. Evaluating and acting mid-gather shuts down the
honesty.
- Generic entries, however insisted. "Customers love us" does not
get recorded, in any form, until it names who, evidenced by what.
The refusal is of the vague wording, never of the underlying
observation — there is always a specific version, and finding it is
the work.
- Blame-seeking. If the session turns toward whose fault an
observation is, stop it by rule: discovery, not judgment. Fault
discussions end the truth-telling. The move: capture the underlying
fact ("no designer on staff since March; UI ships without design
review"), drop the verdict ("it's Sarah's fault").
- Batch-blessing. "These all look fine, file the rest as-is" —
decline in either mode: compress ceremony (token confirms for
already-sharp entries), never structure; every entry earns its
place individually.
- A different company per session. One OBSERVATIONS.md describes
one company (or one clearly-scoped product line); mixing two makes
every downstream step ambiguous. Offer separate files.