Talking to Humans by Giff Constable — Cliff Notes Summary
PULSEKNOWLEDGE LIBRARY
*Talking to Humans* by Giff Constable with Frank Rimalovski is a sub-100-page tactical field guide to running customer-discovery interviews. Its argument: discovery is a learnable skill, and most people fail it by pitching instead of listening. Write hypotheses first, recruit the right segment, listen roughly 80% of the time, and synthesize patterns across interviews — never act on one conversation.
The outcome you should expect from working the book
The honest version of what this book buys you is not a revelation. It is a reduction in wasted quarters. Teams that adopt Constable's discipline stop discovering, six months into a build, that the person they interviewed had no budget authority, no daily contact with the workflow, and no reason to defect from the incumbent. That is the outcome: fewer expensive wrong turns, caught earlier, at the cost of a few weeks of unglamorous conversation.
Concretely, a team that runs the method for one quarter should be able to produce four artifacts it could not produce before. First, a written segment definition specific enough that a stranger could screen candidates from it — not "small business owners" but something with a job title, a company-size band, and a trigger event. Second, a hypothesis sheet with each belief marked confirmed, refuted, or inconclusive, dated, with the interview IDs that moved each verdict. Third, a stakeholder map per target account naming who decides, who uses, who blocks, and who signs. Fourth, a vocabulary list of the customer's own words, captured verbatim, which becomes the raw material for every landing page, cold email, and demo script that follows.
That last artifact is the one teams underrate and the one that pays fastest. Marketing copy written in the customer's vocabulary outperforms copy written in the founder's vocabulary, and the only reliable source of that vocabulary is a recording of someone describing their problem without being prompted with your words first. When a discovery program produces nothing else, it usually still produces this — and it is worth the effort on its own.
The downstream effects reach further than product. Sales discovery calls, onboarding questionnaires, win/loss interviews, churn exit calls, and even investor diligence conversations are all the same motion with different framing: get a human to tell you something true that they had no particular incentive to volunteer. Once a team internalizes the listen-ratio discipline and the hypothesis-first structure, it tends to leak into all of those adjacent workflows without anyone mandating it. A revenue-operations leader who reads this book usually ends up rewriting the discovery-call template, then the win/loss script, then the QBR agenda — because the same failure mode, talking when you should be extracting, is present in all three.

What you should not expect is certainty. Discovery narrows the space of plausible wrong answers. It does not hand you the right one. Constable is explicit about this: customers are excellent at describing their current pain and unreliable at predicting their future behavior. The interview tells you what is broken today. It does not tell you whether they will pay you to fix it. That verdict comes from a test — a letter test, a concierge delivery, a signed pilot — not from a conversation.
What drives that outcome
Four mechanisms do the actual work, and they compound. Remove any one and the program degrades into what most teams already do badly.
Hypothesis-first framing. Writing down what you believe before the conversation is the single highest-leverage habit in the book. It converts an interview from a fishing expedition into an experiment with a verdict. The mechanism is not mystical: unwritten beliefs are unfalsifiable, because a founder with unwritten beliefs will unconsciously steer toward confirmation and then remember the conversation as supportive. Written beliefs can be marked false. A hypothesis sheet for a procurement-software startup might state that purchasing managers lose multiple hours weekly to purchase-order follow-up; that the pain is acute enough to justify a per-seat subscription; that the economic buyer is finance rather than procurement; and that the incumbent is spreadsheets plus email rather than legacy software. Each of those is testable in a twenty-minute conversation, and each maps to a specific question.

Segment precision. Talking to the wrong person is worse than talking to nobody, because it generates false signal that feels like data. The failure is subtle — an interview with an enthusiastic person who is not your buyer produces genuine enthusiasm, genuine detail, and a completely misleading conclusion. Precision in the segment definition is what prevents this, and it has to be written before recruiting starts, because recruiting pressure will otherwise relax the definition one candidate at a time.
The listen ratio. Constable's most-quoted formulation is that every sentence you speak is a sentence they are not speaking. The prescribed ratio is roughly 80% customer, 20% interviewer. This is mechanically true rather than philosophically true: an interview is a fixed-length container, and words you spend are words they cannot. When you catch yourself pitching, the corrective is to stop mid-sentence and ask a question. It feels rude. It is not.
Pattern over anecdote. A single interview is noise. Constable leans on Jakob Nielsen's usability-research finding that a small handful of users surfaces the large majority of usability issues, and argues the same diminishing-returns curve applies to discovery. Somewhere around the fifth to seventh conversation in a segment, you start hearing the same answer phrased three different ways by three different people. That repetition is the signal. Anything heard once is a hypothesis for the next batch, not a conclusion.
The fifth mechanism, which Constable treats as a B2B special case but which deserves equal billing, is stakeholder triangulation. In business software the unit of discovery is the account, not the person. The buyer signs and cares about cost, risk, and procurement fit. The user works in the product daily and cares about speed and not hating their job. The champion sells the deal internally and cares about looking right. These are almost never the same human, and they almost always want different things. Interview only buyers and you build a product users route around. Interview only users and you build something procurement blocks in security review. The correction is a minimum of three conversations per target account, mapped against each other.

Benchmarks and realistic ranges
The book gives working numbers rather than precise ones, and the honest framing is that these are planning heuristics, not measured constants.
Interviews per segment. Five to seven conversations gets you to pattern recognition within one segment. A full discovery sprint against a segment runs larger — on the order of ten to fifteen conversations over two to three weeks — because you want the pattern plus enough surrounding variation to know where its edges are. If you are testing three segments, that is three separate batches, not one blended pool of thirty. Blending segments is the most common way teams destroy their own signal: the pattern in each segment gets averaged into mush.
Recruiting effort. This is the number that surprises people. Recruiting is roughly half the total work, and the funnel math is brutal. To book twenty conversations, plan to source something in the range of eighty to a hundred introductions or outreach attempts. Founders who budget time for twenty interviews and not for the hundred asks that produce them consistently run out of runway on the recruiting step and quietly settle for six conversations with people who were easy to reach — which is to say, people who are not representative.

Channel response rates. The ordering is stable even if the exact percentages vary by market. Warm introductions through a mutual connection convert at a multiple of anything else. The founder's own second-degree network is next. Community channels — the subreddit, the Slack group, the Discord where your segment actually congregates — sit in the middle and work well when you participate rather than broadcast. In-person asks at conferences and meetups convert far better than the same ask sent as an email. Paid research panels are legitimate and fast but cost real money per participant, and are best held in reserve for when warm channels are genuinely exhausted or when you need a segment you have no access to. Cold outreach is the floor: low single-digit response rates on a good day.
Interview length. Twenty to thirty minutes is the right ask and the right actual duration. Asking for an hour depresses acceptance sharply. Asking for fifteen minutes gets acceptance but not depth. Twenty minutes is the number that fits in someone's calendar gap and still allows five to eight real questions with follow-ups.
Question count. Five to eight prepared questions per interview, each mapped to a hypothesis, with room for follow-ups. More than that and you are running a survey out loud, which is the worst of both formats — you get the rigidity of a survey without the sample size.
Debrief latency. Within twenty-four hours, and same-day if you can manage it. Memory for the texture of a conversation — the hesitation, the moment the energy changed, the exact phrase they used — decays fast. The transcript preserves the words; only a fast debrief preserves the read.

Cadence. The book prescribes batch sprints: a batch of interviews, then a full-team debrief, then a re-recruit against updated hypotheses. Teresa Torres's later work on continuous discovery refines this to a steadier rhythm — a couple of conversations per week, permanently, rather than periodic bursts. Both work. The sprint model suits a team that needs to answer a specific question before a specific decision. The continuous model suits a shipping product team that needs a standing connection to users. Most organizations end up running both: continuous background contact, punctuated by focused sprints when a real fork appears.
Buying-committee size. The reason the buyer/user/champion triangulation has aged well rather than badly is that business-software buying committees have grown, not shrunk. Industry research on B2B buying consistently reports committees in the high single digits of stakeholders for meaningful purchases. Three interviews per account is therefore a floor, not a ceiling — it covers the three archetypes, not the full committee.
Risks, edge cases, and failure modes
Confirmation bias wearing a lab coat. The most dangerous failure is a team that runs the process, produces the artifacts, and still only hears what it wanted to hear. Hypothesis sheets do not prevent this by themselves; a team can write hypotheses it has already decided are true and grade generously. The countermeasure is procedural: have someone who did not write the hypothesis grade it, and require a quoted line from the transcript for every "confirmed."

Leading questions. "Would you find it useful if we built X?" is not a question, it is a pitch with a question mark. People are polite, especially to founders who seem earnest. The reliable substitute is past-tense and specific: walk me through the last time this happened; what did you do; what did it cost you; what did you try before. Past behavior is evidence. Future intention is a wish.
Enthusiasm mistaken for demand. Someone can be genuinely excited about your idea and still never buy it. Excitement is cheap. The only expensive signals are the ones with a cost attached — a scheduled follow-up, an introduction to their boss, a willingness to pilot, a pre-payment. Constable's letter test and concierge test exist precisely to convert cheap enthusiasm into an expensive signal early.
Polish suppressing criticism. The sketch-not-polish rule is a real behavioral finding, not an aesthetic preference. A hand-drawn wireframe signals that the work is disposable and invites attack. A pixel-perfect mockup signals that someone spent a weekend on it and triggers a politeness reflex. If you must show a high-fidelity artifact, say out loud that you will throw it away, and mean it — though the sketch is still better.
Interviewing the reachable rather than the relevant. Under recruiting pressure, every team drifts toward whoever answers. Your co-founder's former colleagues, your existing customers, the friendly people in your community — all of them are easy and none of them are a random sample. The specific danger with existing customers is that they already selected into your worldview. They will confirm you. Non-customers and churned customers are the higher-information conversations and the harder ones to book.

N=1 acting. A vivid single interview, especially from a well-known company or a charismatic person, has disproportionate pull in a team debrief. One articulate person's edge case can redirect a roadmap. The rule against acting on a single interview exists because this failure mode is nearly irresistible without an explicit rule.
Wrong stakeholder, right conversation. You can run a technically excellent interview with a person who cannot buy, cannot champion, and does not use the thing. Everything about the conversation feels productive. This is the highest-cost failure in B2B because it is invisible until the deal stalls at a stage nobody interviewed for. Screen for role before you book, not after.
Over-indexing on the loudest pain. Customers describe the pain that is top of mind, which is often the most recent rather than the most costly. Ask about frequency and cost, not just annoyance. A weekly two-hour annoyance is a bigger business than a quarterly catastrophe that everyone remembers vividly.

Discovery as theater. A program that runs interviews but never kills a hypothesis is theater. If six months of discovery has refuted nothing, the process is decorative. A healthy hypothesis sheet has refutations on it.
Compliance and consent edge cases. Recording a conversation requires consent, and in regulated segments — healthcare, financial services, government — even an informal research call can carry constraints on what the participant may discuss. Ask permission to record explicitly, store transcripts where your own policy says they belong, and do not pull confidential details into a deck. Compensation for participants is normal and fine, but disclose it, and be aware that paying changes who volunteers.
Small-market saturation. In a genuinely narrow segment — a few hundred qualified accounts globally — you can burn your addressable market on research calls. In those markets, treat every discovery conversation as also a relationship-opening conversation, do not send a junior interviewer, and do not ask for twenty minutes and take fifty.
A practical rollout plan
Here is a version that works for a team of three to six people starting from nothing.

Week zero — define and commit. Write the segment definition and the hypothesis sheet before anything else. Keep them in one shared document with a version date. Assign two named roles per interview: an interviewer and a note-taker. Constable is firm that one person cannot do both well, and this survives the arrival of automatic transcription — the note-taker's job is not stenography, it is watching for hesitation, energy, and the moment the person's vocabulary diverges from yours. Decide your consent and recording policy now, not in the first call.
Week one — build the recruiting funnel. Treat this as a sales motion, because it is one. Source your list from mutual connections first, your own network second, communities third. Write the ask as a short script: you are working on a project related to their problem area, you are not selling anything, you would value twenty minutes of their perspective. The phrase about not selling is load-bearing — it is the difference between a reply rate that sustains the program and one that kills it. Send the asks in a batch and expect to send four or five for every conversation you book.
Weeks two and three — run the first batch. Five to eight prepared questions, each mapped to a hypothesis. Open with context about them, not about you. Use past-tense prompts. When the conversation goes somewhere unplanned and interesting, follow it — the prepared questions are a floor, not a cage. Debrief the same day, mark each hypothesis, and paste the customer's exact phrasing into a running vocabulary document.

End of week three — synthesize. Cluster the responses into themes. Sticky notes on a wall still work; a shared board or a research repository works too, and modern transcription and clustering tools genuinely accelerate the mechanical part of this step without changing the judgment part. Anything appearing independently in three or more interviews is a pattern. Everything else goes back on the hypothesis sheet for the next batch. Run the four debrief questions as a team: what survived, what was refuted, what surprised us, what changes next sprint.
Week four onward — move to solution testing, then to a standing rhythm. Once a problem pattern is confirmed, climb the fidelity ladder deliberately: napkin sketch, whiteboard, clickable wireframe, low-fidelity prototype, then working product. Most teams skip the first three rungs and learn nothing from the last three, because by the time something is built, nobody wants to hear it is wrong. At each rung, push on what they would actually pay for and actually use daily. Then run a test with a cost attached — the letter test, where you write the launch announcement first and see whether anyone replies asking for it, or the concierge test, where you deliver the service manually before automating any of it.
After the first sprint, shift to a standing cadence of a couple of conversations per week so the connection never goes cold, and reserve full sprints for genuine forks in the road.
Adjacent adoption. The same operating system transfers cleanly to neighboring workflows, and the marginal cost of extending it is low. Win/loss interviews become hypothesis-first rather than post-hoc rationalization. Churn exit calls get a listen ratio and a debrief instead of a survey link. Sales discovery calls get prepared, hypothesis-mapped questions instead of a feature tour — which is why this book has had an unexpected second life in revenue organizations, where discovery-call coaching now borrows its structure wholesale. The core strategy is identical in every case: decide what you believe, ask questions that could prove you wrong, shut up, and grade yourself afterward.
Related questions
How does this compare to *The Mom Test*?
They are complements, not substitutes. Rob Fitzpatrick's *The Mom Test* is sharper on the micro-level — how to phrase a question so that even your mother cannot flatter you. Constable is broader on the operating system: recruiting, structure, stakeholder coverage, synthesis, cadence. Read Fitzpatrick for the sentences, Constable for the program.
Do I still need this if I already run continuous discovery?
Yes, as a primer. Teresa Torres's continuous-discovery work assumes interviewing competence and focuses on cadence and opportunity mapping. Constable supplies the underlying interview skill. Teams that adopt continuous discovery without the interview fundamentals just conduct bad interviews more frequently.
How many interviews before I can trust a conclusion?
Five to seven per segment for pattern recognition, ten to fifteen for a confident sprint verdict. Never one. And "trust" here means trust the problem description, not the purchase prediction — that still requires a test with real cost attached.
Can transcription and AI tools replace the note-taker?
Partly. They handle the words reliably. They do not yet reliably capture hesitation, energy shifts, or the moment someone's body language contradicts their answer — which is a large share of what the second person is there for. Use both.
Is the method different for B2C?
The mechanics are the same; the stakeholder triangulation collapses. In consumer, buyer and user are usually the same person, so three-interviews-per-account becomes one-per-person and the volume shifts upward. Recruiting also gets easier and cheaper, which raises the risk of interviewing whoever is convenient.
FAQ
**What is the core argument of *Talking to Humans*?**
That customer discovery is a learnable skill rather than a personality trait, and that most people fail at it in the same predictable ways — pitching instead of listening, seeking confirmation instead of testing beliefs, and showing polished work that suppresses honest criticism. The book is a tactical field guide for doing it correctly, with scripts, structures, and debrief formats.
Who wrote it and when?
Giff Constable with Frank Rimalovski, self-published in 2014, with a foreword by Steve Blank. It sits deliberately downstream of Blank's customer-development theory: Blank said get out of the building, and this book explains what to actually do once you are outside it.
Who should read it?
Founders and product managers first, but it has spread well beyond that. Revenue-operations leaders, account executives who run discovery calls, customer-success teams doing churn interviews, and anyone building an internal tool for colleagues all get direct value. The core motion — extracting truth from someone with no incentive to volunteer it — is the same everywhere.
How long does it take to read?
Under a hundred pages, so one or two sittings. The brevity is intentional and is part of why accelerators recommend it — it is a book people actually finish. The reading is the small part; the value is in running the first batch of interviews within the following week.
Does it include usable templates?
Yes. Sample conversation scripts, recruiting outreach language, hypothesis-sheet structure, and debrief formats. They are meant to be adapted rather than used verbatim, but they remove the blank-page problem that stops most first-time interview programs before the first call.
Has it aged well since 2014?
The principles have held up entirely — the same advice appears, sharpened, in every subsequent discovery book. What has changed is tooling: remote video calls, automatic transcription, and clustering software have made the mechanical steps faster, and the cadence guidance has shifted from periodic sprints toward continuous contact. Neither change touches the underlying method.
Sources
- https://www.talkingtohumans.com/
- https://giffconstable.com/
- https://steveblank.com/category/customer-development/
- https://www.nngroup.com/articles/why-you-only-need-to-test-with-5-users/
- https://www.nngroup.com/articles/how-many-test-users/
- https://www.momtestbook.com/
- https://www.producttalk.org/continuous-discovery-habits/
- https://www.ycombinator.com/library
- https://entrepreneur.nyu.edu/
- https://hbr.org/2015/09/making-sense-of-customer-discovery
Related on PULSE
- [The Mom Test by Rob Fitzpatrick — Cliff Notes Summary for Sellers](/knowledge/bs0212)
- [Continuous Discovery Habits by Teresa Torres — Cliff Notes Summary](/knowledge/bs0193)
- [The Four Steps to the Epiphany by Steve Blank — Cliff Notes Summary](/knowledge/bs0142)
- [The Challenger Sale by Matthew Dixon and Brent Adamson — Cliff Notes Summary](/knowledge/bs0001)
- [The Advantage by Patrick Lencioni — Cliff Notes Summary for Sales Leaders](/knowledge/bs0318)









