What a good reference call actually sounds like.

The questions that get past the rehearsed answers — and the silences that tell you more than the words. A field guide built from two hundred recorded calls.

01The script everyone uses

The standard angel reference call runs forty-five minutes, follows the same six-question script every IC partner has used since 2012, and yields almost nothing of substance. The founder hands over four customer names. Three are champions. One is the customer the founder wishes they had more of. The call books for next Tuesday at 3pm. The customer says good things, in the order the founder expected. The IC partner takes notes. The call ends. The dossier adds the reference as confirming evidence and the round closes.

This is the modal reference call in our pipeline, and we have seven hundred of them on file. We benchmarked the signal extracted per call across two cohorts: the calls our analysts ran against a structured rubric, and the calls done in the founder-supplied free-form pattern. The structured rubric extracted material new information — defined as a finding that altered the IC's prior — in 38% of calls. The free-form pattern extracted material new information in 7%. The difference is not the interviewer. The difference is the question architecture.

This note is the field guide we use across the eight-agent diligence pipeline when we run reference calls ourselves, and the structure we recommend when an IC partner runs them. Built from two hundred recorded calls, scored by signal yield against IC outcome, refined twice. The script has six sections, asks roughly twenty questions across forty minutes, and reliably produces information the founder would not have volunteered. The discipline is not the questions themselves — half the questions are obvious in hindsight. The discipline is the order, the silences, and the willingness to skip the close.

02Why rehearsed answers happen

Two structural facts make the standard call low-signal.

First, the founder picks the references. The founder picks them precisely because they are the calls the founder would want the IC to make. Selection bias is not avoidable through any technique applied inside the call. It has to be addressed in the ask: which two customers churned in the last six months, and can we speak to them. If the founder cannot or will not produce those names, the structural information has been collected before the first call begins. See Note 05, flag eleven.

Second, the customer being interviewed has a relationship with the founder that survives the call. If the customer flags a real concern in the reference, the founder will hear about it within an hour. Every customer knows this. They titrate their honesty accordingly. The work of the reference call is not extracting opinions the customer hides — most customers will not hide them — it is extracting the information the customer has not been asked for in the form the founder coached them on.

Founders coach the call. They do not coach it in detail; they cannot, because they do not know what the IC will ask. They coach the framing: they're raising, the call is friendly, talk about what's working. The customer arrives at the call ready to talk about what is working. The IC, having no countervailing structure, asks questions that invite the prepared answer. The prepared answer arrives. Both parties leave satisfied. The IC has confirmed a prior they already held.

The job of the architecture below is to ask questions the customer was not coached on, in an order that prevents recovery to the prepared answer once the call begins to drift.

03The question that breaks the rehearsal

Two question forms account for almost all of the high-yield findings in our two-hundred-call corpus. Both work because the customer cannot have prepared an answer in advance — the customer has to think about the question for the first time, on the call, and the thinking is what produces the signal.

The first form is the counterfactual: “What would happen to your workflow if [company] disappeared tomorrow?” The question forces the customer to enumerate the actual substitutes available, the cost of switching, and the residual capability they would retain or lose. Customers who use the product because it is good at one thing answer the question in thirty seconds — they name the substitute, estimate the time, move on. Customers for whom the product is structural answer the question for two minutes and surface dependencies the founder has not mentioned. The shape of the answer is the signal. The content is secondary.

The second form is the moment-of-doubt: “Walk me through the last time you almost stopped using it. What happened?” This is the highest-yielding question in our entire corpus by a wide margin. Every customer has a moment, and most have not been asked about it in any prior conversation with anyone in the founder's ecosystem. The question elicits the actual failure mode of the product, the actual response of the founder's team to the failure, and — most usefully — the actual reason the customer stayed.

The reason the customer stayed is more diagnostic than any other piece of information that surfaces in a reference call. It is the customer's own articulation of the moat. It is uncoached. It is specific. And it is testable against other customers — if three different customers all stayed because of the same property of the product or the team, you have a moat. If three different customers stayed for three different idiosyncratic reasons, you have a product that is good in different ways for different customers, which is a different and weaker form of moat.

Both forms should appear in every reference call. They sit late in the call, after the customer has settled and is no longer scanning for the standard questions.

04What silence tells you

Half of the high-yield information in a reference call surfaces in the pause between the question and the answer.

A customer who is comfortable with a question answers within two seconds. A customer who is being asked something they have not considered answers within seven. A customer who is being asked something they have considered but would rather not address answers within fifteen, and the answer that follows is shorter and more rehearsed than the answers to the other questions in the call.

The fifteen-second pause is the strongest single signal in any reference call. It does not appear in every call. When it appears, it is almost always on a question the founder did not coach the customer on — because the founder did not anticipate the question, or because the founder anticipated it and chose to leave it alone. Either way, the fifteen-second pause locates the conversation the founder did not want to have.

The discipline is against the temptation to fill the pause. Most interviewers, on hearing seven seconds of silence, restate the question or offer an example to clarify. Both ruin the signal. The customer hears the clarification as permission to retreat to a smaller answer. The right response is silence. Wait as long as the customer needs. The answer that arrives after fifteen seconds is almost always longer and more honest than the answer that would have arrived after seven.

We mark every pause longer than ten seconds in the call notes. The marks track with material findings at a roughly two-to-one rate.

05The eight question types, ranked by yield

Across our two-hundred-call corpus we tagged every question into one of eight categories and scored the yield — the rate at which the question produced information that altered the IC's prior. The ranking is durable across sectors and customer sizes. The ranking is also counterintuitive: the questions IC partners default to are concentrated in the bottom three tiers.

  • Moment-of-doubt (92). “Walk me through the last time you almost stopped using it.” Highest yield by a wide margin. Always include. Late in the call.
  • Counterfactual (75). “What would happen if they shut down tomorrow?” Second-highest. Forces enumeration of substitutes. Include in every call.
  • Pain-articulation (55). “What is the problem this solves that nothing else solved?” Mid-yield. Useful when the customer is specific; weak when the customer is generic.
  • Competitive (50). “What else did you consider before choosing them?” Mid-yield. Surfaces the competitive set the founder may not have disclosed.
  • Capability (40). “What's the product like to use?” Low-mid yield. Customers describe what the founder showed them; signal is the gap between the customer's description and the founder's.
  • Origin (30). “How did you find them?” Useful for CAC validation, low yield for product or team signal.
  • Bio (25). “Tell me about your role and team.” Necessary for context, almost no diligence signal.
  • Recommendation-close (22). “Would you recommend them to a peer?” Almost zero signal. Universal yes from any reference the founder volunteered.

A defensible reference call spends roughly 60% of its time on the top two categories and the silences that follow them. Most IC reference calls spend roughly 60% on the bottom three. This is the single biggest improvement available to the average IC reference-call practice, and it requires no new technique. It requires reallocating time.

A practical rule: write the eight categories on a notepad before the call, mark the time spent in each, and review the marks afterwards. The first three calls done this way usually reveal that the interviewer is spending sixteen of forty minutes on bio + capability + close, which between them produce almost no diligence signal. The reallocation happens naturally over the next five calls.

06The three calls every diligence should make

Founders typically volunteer four references. We recommend three additional categories of call that are not on the founder's list. Each surfaces a different kind of signal.

The champion call is the one the founder volunteered. It is useful as confirmation, useful for capability questions, and useful for the moment-of-doubt question if the customer is willing. It is not useful for product–market fit triangulation, because the champion has by definition found fit. Do at most two of these. The yield falls off after the second; the second confirms or contradicts the first, and a third champion who agrees with the first two does not add a new vector.

The detractor call is the one we ask the founder to introduce. The ask is specific: name two customers who have either churned, downgraded, or expressed material dissatisfaction in the last six months, and route us to them. The detractor call is the most valuable single conversation in any reference cycle. It surfaces the actual failure mode of the product, the actual response of the founder's team to the failure, and the customer's revised view of the company in retrospect. Founders who cannot produce a detractor reference flag a problem upstream of the call. Founders who produce one immediately and route us within forty-eight hours pass a meaningful test independent of the call's content.

The orphan call is the one we source ourselves. Pull the customer list from the founder's case-study page, the founder's LinkedIn endorsements, the press release archive, or a regulatory filing. Pick a customer the founder did not volunteer. Cold-outreach them. The orphan call is the only conversation in the cycle that is structurally uncoached — the customer has not been told the call is coming, has no relationship-management reason to be diplomatic, and has the same product-experience distribution as the median customer rather than the selected one. The yield is high. The friction is also high; only one in three orphan-call attempts results in a call within the diligence window.

A complete reference cycle includes both champion and detractor calls, and at least one orphan call. The composition matters more than the count. Six champion calls is worse than two champion + one detractor + one orphan.

07What to capture in notes

We score every reference call on four axes and record them in the dossier:

  • Capability confirmation. Did the customer's description of the product match the founder's. Where it diverged, in what direction. Divergence in the customer's favour — the customer describes more capability than the founder claimed — is rare and important. Divergence the other way is common and load-bearing.
  • Stickiness. The customer's answer to the counterfactual and moment-of-doubt pair. We rate this 1–5 against a rubric. Full enumeration of substitutes and a specific reason for staying scores high; vague affirmation scores low.
  • Renewal posture. What the customer's default action is at the next contract date. Most customers will not volunteer this. The question to ask is “what would make you not renew?” The answer is more useful than any positive endorsement, because it locates the specific change in the product or the relationship that would cost the founder this customer.
  • Founder-team relationship. How the customer describes the support experience and the responsiveness of the founder when things go wrong. This is the single best proxy we have found for how the founder will treat their post-Series A customers when the team has grown and the founder is no longer in the room.

The four scores go into the dossier. They are also the scaffolding for the IC's own conversation with the founder. If stickiness scored a 2 across three calls, the next conversation with the founder should be about retention, framed against the specific reasons customers gave. If founder-team relationship scored a 5 across three calls, the conversation should be about whether the founder has identified the operators who will replace them in that relationship when the company scales.

Run this on your deck

Want the rebuild on your TAM slide?

The four-cut framework — and every other pattern in the research archive — is baked into every Early Capital dossier. Upload a deck and you get the reconciliation back in roughly fifteen minutes: claimed number, rebuilt number, the gap, and where it came from.

Median reconciliation14%of claimed TAM survives a four-cut rebuild · n = 1,142
Continue reading · related notes