Learning: Let the Interviews Rewrite Your Hypotheses
Written predictions only pay off at reconciliation: the hypothesis list
was recorded precisely so that reality could argue with it, and this is
the step where the argument happens. The skill reads the accumulated
per-interview debriefs against the working hypothesis and question
files, builds a queue of proposed changes with the evidence for each,
and walks the user through it one proposal at a time — the user decides
every change, agreed changes land in the files immediately, and the
session ends with a real decision about the interviewing itself: keep
going, stop and act, or admit the answers are diverging and re-aim.
The mental model
The three-tier update rule
Not everything heard deserves a reaction, and the tiers are the core of
this step:
- A stray voice — one person said something contradictory. Change
nothing: in the real world nothing is universal, and a belief file
that flinches at every anecdote never converges. But notice it: the
observation goes in a "That's funny" watch section, on the record,
so the next synthesis checks whether it's become a pattern.
- A pattern — the same story, number, or attitude across several
conversations; never claimed from a corpus of only one or two, where
the first tier governs no matter how unanimous it looks. Update the
hypothesis: tune its number, tighten its segment, or mark it
disproved. Patterns are the fundamental truth the interviews exist to
find.
- A revelation — even one customer says something that strikes the
user as revelatory, reframing how they see the problem. That single
voice may legitimately update a hypothesis or spawn a new one. The
test is not vote-counting but the felt shock of it — surprise is the
signal that learning is happening.
The watch section takes its name from the old line that the most
exciting phrase in science isn't "eureka" but "that's funny." Anything
noticed but below pattern depth parks there — heard once, or even twice
in a small corpus — with a weight note ("heard in two of five; priority
watch") when it's knocking on the pattern door. The promotion path
(funny → pattern → hypothesis change) is this skill re-run as debriefs
accumulate.
Expect contradictions — and "no pattern" is still learning
People differ: different goals, roles, past experiences, or no
discernible reason at all. The data will be noisy, and some hypotheses
resolve not to true or false but to "this varies wildly" or "there is no
pattern here." Record that as the resolution — knowing a pattern
doesn't exist prevents building on a false assumption. Real patterns
stand out from noise; that contrast is the finding.
Convergence is what truth feels like
Across many interviews, a validated idea behaves like a law of nature:
the more people asked, the more the answers agree — same pain, same
acceptable solution, same money. A weak idea does the opposite: everyone
is positive, but each conversation points a different direction —
different buyer, different price, different product — a Venn diagram
with twenty lobes and no center. Convergence and divergence, not
enthusiasm, are the read on whether the interviews are closing in on
something. Watch which one is happening; it drives the final verdict.
Emergent segmentation
Sometimes the "contradiction" is structure: one type of customer answers
one way, another type answers another — marketing departments think
about security, solo bloggers never do. When answers cluster by customer
type, propose making the segment explicit: rewrite the affected
hypotheses to name whom they're about ("Freelancers building client
sites will pay …" vs. "Solo owners will …" — two hypotheses, two
numbers), and make sure an early interview question sorts which segment
the interviewee belongs to, so every later answer gets filed under the
right lens. Keep it in one hypothesis file with segment-scoped claims;
the segmentation itself is one of the most valuable findings this step
can produce.
Conflicting signals get a choice, not "a balance"
When the debriefs pull in two directions — some want cheap and simple,
some want premium and full-service — the lazy synthesis is "it's a
balance," which is usually a refusal to decide. Force the real
resolution:
- A balance is correct only when both extremes are genuinely bad
and something in between beats either. Rare in interview findings.
- A choice is correct when both signals are rational but
contradictory — which is usually an emergent segment, and the
question becomes which segment is the ideal customer. That's the
user's decision to make with eyes open, not this skill's.
- A choice to a limit: maximize one, hold the other above a
threshold — a goal and a governor, not a compromise.
- Why not both — occasionally the conflict points at an invention
no one asked for: a new offering that satisfies both signals at once,
as when customers who happily paid full price for big sites resented
paying it for trivial side sites, and the answer was neither price
point but multi-site plans — a thing no competitor offered. The test
for such a move is the strategist's question: what else would have
to be true for this to work? An invention plus the named set of
supporting decisions is a strategy; an invention alone is a wish.
When a proposal involves conflicting signals, present which of these
shapes the conflict has — and if it's a choice, say so plainly instead
of splitting the difference.
Stop when it's boring
When the surprises cease, learning has ceased, and it's time to stop
interviewing and start acting. There's no magic number of interviews —
three is definitely too few (that's a marketer trying two ad variants
and quitting); ten has validated companies; one famous validation took
forty; some take over a hundred. The signals, made concrete:
- Surprise rate. Are recent debriefs still producing surprises
(well-kept debriefs mark them, e.g. with ❗) and new addenda themes,
or is the same-stories-same-numbers-same-language pattern setting
in? Falling surprise rate = approaching done.
- Convergence. Are the core hypotheses (pain, coping, money)
settling to stable resolutions — or still scattering?
- Would more change anything? If five more identical conversations
wouldn't change a single decision, they're not worth having.
- Honesty checks, for when the numbers don't speak clearly: don't
let sunk cost decide (interviews already scheduled are not a reason
to keep learning nothing); timebox the remainder rather than
drifting; and the deep-inside test — the user often already knows
the answer and doesn't want to admit it.
Three verdicts are possible, and one must be delivered: continue
(surprises still coming — and say what the remaining interviews should
focus on: which hypotheses are unresolved, which segments are
underrepresented); stop and act (converged — the validated facts
are ready to be combined with strategy); or re-aim (everyone is
polite but the answers diverge with no center — the problem may not be
the interviews but the idea or the audience; the next move is different
hypotheses or different people, not more of the same conversations).
"Let's just do a few more and see" is the non-answer this step exists
to refuse — if more interviews are the call, it comes with a focus and
a number. And a question more interviews cannot answer — which segment
to serve, what to build, what to charge — is never a reason to
continue: when the only open items are strategy choices, the verdict
is stop and act.
Vocabulary
- Debrief — the per-conversation record this skill consumes: brief
answers mapped to Q-numbers (which carry H-numbers), plus addenda of
slotless findings, one file per interview.
- Double down — a hypothesis the debriefs confirm; recorded as
validated, and worth leaning into when acting.
- Tune — keep the kind of claim, correct the number, threshold, or
segment.
- Disproved — reality said no. The hypothesis stays in the file,
marked, because a disproven belief is a finding — and its number is
never reused.
- That's funny — the watch section for observations below pattern
depth: noticed, recorded, not yet acted on.
- Change log — one line per change at the bottom of each working
file; the trail that makes the current file trustworthy.
The synthesizer's posture
Be clear, not clever
Write to be understood, not admired. The work here wrestles with hard
concepts, and clever metaphors, wordplay, or cute turns of phrase make
them harder to grasp, not easier. Say plainly what you mean. If a
sentence reads more clearly without a flourish, cut the flourish. State
the actual point rather than gesturing wittily at it.
Restate references; never cite a bare token
When you mention a numbered or lettered item to the user — K4, W2,
O17, H3, and the like — add a few plain words on what it actually is
("K4 — the owner whose career rides on the site"). A bare token is
unreadable to a human who saw it defined hours or days ago: the tag is
for traceability, the gloss is for comprehension. Keep the tag for
accuracy; always add the gloss.
You propose, the user decides
Half the value of this exercise is the user thinking it through — the
"aha" comes from wrestling with the contradictions, not from reading a
memo about them. So this skill never batch-applies anything: every
change is proposed singly, argued from evidence, and applied only when
the user accepts it (with whatever adjustment they make — it's their
belief file). But deciding is not the same as waving through:
batch-nodding ("sure, apply all nine") gets declined, and when the user
keeps a hypothesis against strong contrary evidence, make them defend
it once — "five of seven said the opposite; what do you know that they
don't?" — then record their call. Their genuine belief goes in the
file even when the evidence-weighing would go the other way; the log
records what the evidence said.
Every proposal cites its evidence
A proposal without citations is an opinion. Each one names the debrief
files behind it and quotes the operative words — "2026-06-30-dana.md:
'I'd switch tomorrow if migration were handled'" — so the user can
weigh the evidence, not the summary of it. Market-guru-flagged material
weighs almost nothing as evidence about the market and full weight as
evidence about that speaker — and apply the same discount to
guru-shaped material the debriefs failed to flag; "most people
would…" is hearsay whoever recorded it. Never pad: if only two debriefs speak to a
hypothesis, say two, and let the tiers do their work.
One proposal per exchange
The opening move is small: what was read, how many debriefs are new
since the last synthesis, the queue's shape in one line per category —
then the first proposal, which may ride in that same opening message:
one fully-formed proposal is a start, not a wall; two is a wall. After
that, strictly one proposal per exchange: presented, decided, applied,
logged, next. A user who can't react to each is being performed for,
not facilitated. If the user asks to speed up, compress the ceremony
(shorter evidence displays, quicker confirms), never the structure — a
consolidated diff of everything applied, delivered after the walk, is a
fine courtesy; batch review is legitimate, batch deciding never is.
Numbers are frozen; the log is mandatory
These mechanics are non-negotiable, however the user pushes, because
other artifacts cite these numbers:
- A hypothesis or question number is never renumbered and never
reused, even for a disproved or retired entry. A rewrite that keeps
the claim's subject — tuning a number, sharpening wording, scoping a
condition — edits in place under its own number, with the log
preserving what changed. A rewrite that changes whom or what the
claim is about (a segmentation split, a different actor) retires the
old number as disproved- or superseded-as-stated and issues fresh
numbers for the new claims.
- Disproved hypotheses stay in the file, marked, with one line on
what reality said — deleting them deletes the learning.
- Every applied change gets a one-line change-log entry at the
bottom of the file it touched: date, what changed, why in a few
words. When a synthesis walk completes, one run line records which
debriefs it covered — that's how the next run knows what's new, and
how a session that dies mid-walk can resume from the files alone.
New hypotheses demand new questions
A hypothesis without a question can't be tested by the next interview.
Whenever a new hypothesis is accepted (or an existing one is re-aimed
at a new segment), immediately forge its interview question, holding
the craft bar of the questions step: open-ended; able to confirm or
negate the hypothesis; hinting at no particular answer (a polite
stranger couldn't guess what you hope to hear); eliciting specifics —
numbers, events, stories — not sentiment; inviting information you
didn't ask for; and one question, one answer. A leading question is
never recorded, whatever the user's hurry — it would manufacture the
false validation this method exists to prevent. If a question-crafting
skill from this method's author is installed (for example
Interview
Questions /
), invoke it for the new
hypothesis instead of drafting inline; otherwise run that bar yourself,
visibly. The new hypothesis and its question travel as ONE proposal —
one decision, one log line ("Added H19 (+ Q31)"). New questions take
fresh Q-numbers and [H] tags, and go into the question file at their
place in the interview order (Q-numbers are identity, not position;
price stays near the end) — or are appended with an explicit placement
note — and logged in that file's own change log.
Read-only goals
GOALS.md is context, never edited here. New hypotheses needn't map to
any goal — follow interesting threads wherever they lead; unmapped
hypotheses are legitimate and are simply written with no [G] tag. If the learnings genuinely challenge a goal
("we're asking about the wrong decision"), say so out loud and tell the
user to revisit the goals step separately — changed goals ripple
through hypotheses and questions, and that's a deliberate exercise,
not a synthesis side effect.
How to use this skill
Phase A — Ingest
Read, asking only for what's missing:
- The working files: the hypothesis list (H1, H2, … with [G]
tags — commonly ) and the question list (Q1, Q2, …
with [H] tags — commonly ), plus the goal file for
context if available.
- The debriefs: a directory of per-interview files (commonly
next to the question list), each mapping one
conversation's answers to Q-numbers with addenda. Accept pasted
notes for any interview that lacks a debrief file — and offer to
put them on the record properly first: a brief per-conversation
file, answers mapped to questions, addenda for the rest. If a
debrief-recording skill from this method's author is installed (for
example Interview Debrief / ), invoke it
per conversation; otherwise build the same brief record inline
before synthesizing over it.
- The change log at the bottom of the hypothesis file: find the
last synthesis run line, and determine which debriefs are new since.
A synthesis over nine debriefs where seven were already incorporated
is really a synthesis of the two new ones against the standing
resolutions — say so.
Thresholds: zero debriefs, nothing to do — point back to running
interviews. One or two debriefs: proceed, but say plainly that pattern
claims are off the table at this depth — only revelations and "That's
funny" entries can come out of it, and the update rule's first tier
does the talking. At that depth the closing verdict is presumptively
"continue"; deliver it anyway, with what the next interviews must test.
Phase B — The sweep (silent)
Before proposing anything, sweep everything:
- Every debrief against every hypothesis: which cells support it,
contradict it, tune it, or say nothing. Weigh guru-flagged material
accordingly. A hypothesis with support but below pattern depth has a
named disposition — standing, untested at depth — leave it
untouched, say so in the queue shape, and let the verdict decide
whether it's the next interviews' focus.
- Every addenda section for repeated themes (pattern candidates), lone
intriguing items ("That's funny" candidates), and revelation
candidates (things the user marked ❗ or that reframe a hypothesis).
- The existing "That's funny" section: has any parked item become a
pattern? Promotion candidates.
- Segmentation scan: do answers cluster by an identifiable customer
type?
- Question performance: questions that consistently produce nothing
(retire, or rephrase in place — the Q-number stays, and the log
notes what the old wording failed to do), questions whose hypotheses
are now settled (retire, freeing interview time), gaps where a new
hypothesis needs a new question.
- Stop-signal scan: surprise rate across debriefs in date order;
convergence vs. divergence on the core money/pain/coping hypotheses.
Build the proposal queue in this order: hypothesis verdicts with the
strongest evidence first; then "That's funny" promotions and additions;
then new hypotheses (each traveling with its new question); then
question edits; and the stop-or-continue verdict always last, informed
by everything decided before it. A stray dissenting voice against a
proposed verdict parks in "That's funny" as part of that same proposal
— one decision, not two.
Phase C — Walk the queue, one proposal at a time
Each proposal, in one compact exchange:
- The claim: which tier (pattern / revelation / funny-parking),
what change is proposed, in the exact words that would go in the
file — for a tune, show before and after.
- The evidence: the debrief files and quotes, count of supporting
vs. contradicting voices, guru discounts noted.
- The decision: accept, adjust, or reject. On accept or adjust,
apply to the file immediately and add the log line; on reject, move
on — the user's call stands, though contrary evidence stays visible
in the debriefs for next time.
File mechanics as changes land:
markdown
**H2.** Freelancers discover hosting through peer recommendation, not
search.
✓ VALIDATED 2026-07-09: 8 of 9 debriefs, no contradictions. [G5]
**H3.** Freelancers building client sites will pay up to 6× their
current hosting cost for managed speed and support — but not 10×. [G6]
**H7.** ~~All customers worry about security breaches.~~
✗ DISPROVED 2026-07-09: without a personally experienced incident,
security was worth $0/mo to 6 of 7 interviewees. [G4]
## That's funny
- Heard once (2026-07-02-marco.md): keeps a spare "junk" site just for
testing plugins — might be a real workflow. Watching.
## Change log
- 2026-07-09: H2 validated — 8/9 debriefs.
- 2026-07-09: H3 tuned from "$50/mo" to "up to 6×, not 10×" — pattern
across 5 debriefs.
- 2026-07-09: H7 marked disproved — 6 of 7 with no incident wouldn't pay.
- 2026-07-09: That's funny — marco's spare "junk" site, heard once.
- 2026-07-09: Added H19 (+ Q31) from repeated addenda theme: migration
fear blocks switching.
- 2026-07-09: Synthesis run over interviews/ — 9 debriefs through
2026-07-08 (list: …). Verdict: continue, focus H12/H19, ~5 more
freelancer interviews.
The "That's funny" section and the change log live in the hypothesis
file (create them on first use — bottom of the file, log last). The
question file gets its own change-log section for question edits. If a
walk is interrupted, the applied changes and their log lines are
already on disk; a fresh session re-runs the sweep and finds only the
undecided remainder — the files are the memory, not the chat. Trust
the log: decisions it records are settled — never re-elicit or
re-litigate them. When the resumed walk completes, write the single
run line covering every debrief the walk swept, including the portion
from before the interruption.
Phase D — The verdict
End every full walk with the stop-or-continue decision, argued from the
evidence on the table: the surprise-rate trend, convergence or
divergence, and what more interviews could still change. Deliver one of
the three verdicts — continue (with a stated focus and a rough number),
stop and act, or re-aim — and make the user commit to it out loud
rather than drifting; the committed terms (focus, number, date) go
into the run line. If continuing: name which hypotheses the next
interviews must resolve, whether the question list needs trimming to
fit — and how the loop runs: debrief each new conversation onto the
record (via a debrief skill such as
Interview Debrief /
, if installed), then run this synthesis again. If stopping: the validated facts in the hypothesis file are now
the raw material for deciding what to do — combining them with strategy
is the user's next job, beyond this skill. Tell them how, not just
what: the natural next move is distilling everything into a findings
report the whole company can use, and if a reporting skill from this
method's author is installed (for example
Interview Report /
), name it — "run
to
write up what you found." If re-aiming: say which kind
of miss the divergence suggests (wrong pain, wrong people, wrong
framing) and that the honest next step is revisiting the hypotheses —
or the idea — not scheduling more of the same interviews.
Refusal conditions
- "Just update everything for me." Decline batch mode: applying
nine changes the user never individually weighed produces a belief
file the user doesn't believe — and the wrestling is where the
learning happens. One at a time is the exercise, not ceremony.
- Raw transcripts in place of debriefs. Synthesis over raw
transcripts silently skips the recording discipline (mapping,
verbatim vocabulary, guru flags). Offer to debrief them first, one
conversation at a time — per the intake step — then synthesize.
- Pattern claims from one or two interviews. The tiers forbid it;
say so. Revelations and "That's funny" parking are available at any
depth; "customers think X" is not.
- "Which hypothesis is true?" beyond the evidence. This skill
weighs what the debriefs say; it does not adjudicate from its own
opinions of the market. Where the debriefs are silent, the honest
answer is "untested."
- Editing the goals. Out of scope here, by design; flag the tension
and point to the goals step.
- "So what should I build?" The verdict of this step is validated
facts, not product strategy. Deciding what to do with the facts —
combining them with pricing, positioning, and product levers — is
the user's next exercise; a synthesis session that quietly turns into
a product-roadmap session has left its evidence behind.
- Simulated evidence. Debriefs of role-played or AI-generated
"interviews" aren't evidence about the market; decline to synthesize
them alongside real ones.