
TL;DR: Building an ICP model from closed-won deals is the right instinct, but most companies sample the wrong subset: whichever deals are easiest to pull, usually the most recent ones, or the largest ones, or the ones a founder personally remembers. This creates a biased sample that overweights unusual wins, a personal relationship, an unusually low price, a one-off referral, and underweights the more boring, repeatable deals that actually predict what a scalable motion looks like. The fix is a deliberate, structured sampling method, not just "look at our closed deals," which is vague enough to default to whichever ones are easiest to remember or find.
"Build your ICP from your closed-won deals" is good advice that gets executed poorly more often than it gets executed well. The instinct is right: real closed deals are better evidence than market research. The execution usually goes wrong at the sampling step, because "look at our closed deals" is vague enough that most teams default to whichever deals are easiest to recall or pull, which introduces a specific, predictable bias.
Ask a founder or sales leader to describe the company's best customers, and the answer usually comes from memory, not from a systematic pull of the CRM. Memory is a biased sampling mechanism: people remember unusual, emotionally salient deals far better than routine ones. A deal that closed because of a personal relationship, a deal that closed unusually fast, a deal that involved an interesting story, all stick in memory more than a boring, straightforward deal that closed exactly the way the sales process predicted it would. Building an ICP from memorable deals means building it from outliers, not from the pattern that actually represents a repeatable motion.
A Series A company's most recent closed deals often reflect whatever the current campaign or sales focus happens to be, not a stable, validated pattern. If the last quarter happened to close two deals from a specific vertical because of one well-timed campaign, it's tempting to conclude that vertical is the real ICP, when a longer look back might reveal that vertical represents a small fraction of total closed deals and that the two recent wins were more coincidence than pattern. Sampling only recent deals conflates "what we happened to close lately" with "what we reliably close," which are not the same thing.
Larger deals get discussed in board meetings, celebrated in team updates, and remembered vividly, which makes them disproportionately influential in shaping how a team thinks about its ideal customer, even when those large deals represent a small percentage of total closed-won volume. A company that closes one €80,000 enterprise deal and fifteen €15,000 mid-market deals in the same quarter can easily end up building an ICP around the enterprise deal's profile, simply because it was the more memorable, more discussed win, even though the mid-market deals represent the actual, repeatable engine driving most of the quarter's revenue.
Relationship bias is the most dangerous of the three because it's the least visible to the people involved. A founder who closed an early deal through a personal connection often genuinely believes the deal closed because of product fit and messaging, discounting or simply forgetting how much the existing relationship shortened the sales cycle and lowered the buyer's guard. Building an ICP around these deals implicitly assumes a level of trust and access that a cold or semi-cold sales motion, the kind most demand generation and outbound programs actually have to work with, can't realistically reproduce at scale.
Correcting for these biases requires a deliberate, structured pull rather than a memory-based selection. Pull every closed-won deal from the last twelve months, not just the ones that come to mind first, and tag each one honestly against three factors: how the deal originated, whether cold outbound, inbound content, referral, or existing personal relationship, the deal size relative to the company's average, and how long the sales cycle actually took compared to the typical cycle length. Then look specifically at the deals that closed without a referral or existing relationship, at a size close to the company's median deal, through a cycle length close to the typical duration. That subset, not the full list and not the most memorable deals within it, is the most honest representation of what a repeatable motion actually looks like.
This sampling correction works alongside, not instead of, checking an ICP against real sales call outcomes. purple path's guide to building an ICP model that survives contact with real sales calls covers the ongoing correction loop; the sampling bias covered in this article is what corrupts the very first version of the ICP before that correction loop even has a chance to test it against anything. A biased initial sample means the ongoing testing process is validating the wrong baseline from the start.
purple path's analysis of whether firmographic or behavioral signals should come first in an ICP model assumes, as a starting point, that the underlying sample of deals used to identify those signals is itself representative. A firmographic or behavioral pattern extracted from a biased sample of unusually large, unusually recent, or relationship-driven deals produces a signal that looks precise and predicts nothing reliably, because the underlying data it was drawn from was never a fair representation of the company's actual, repeatable customer base.
Twelve months is a reasonable starting point for most Series A companies, long enough to smooth out short-term campaign effects but recent enough to still reflect the current product and market position. A company with a longer sales cycle or lower deal volume may need to look back further to gather enough data points.
Use whatever history exists, but be explicit that the resulting ICP is a working hypothesis based on a small sample, and prioritize the ongoing correction loop against new deals as they close, rather than treating an early, necessarily thin sample as a settled conclusion.
Not excluded entirely, since they're still real evidence of product-market fit at some level, but they should be tagged and considered separately from deals sourced through repeatable channels, since a referral-driven deal doesn't tell you much about what a cold outbound or demand generation motion can reliably reproduce.
Deal count at a given size band matters more than any single large deal's dollar value when trying to identify a repeatable pattern. A quarter with one very large deal and many mid-sized ones is better represented by the mid-sized cluster, even though the large deal contributes more revenue in that specific quarter.
At minimum annually, and ideally alongside the quarterly correction loop covered in validating the ICP against real sales calls, since both the available data volume and the honest picture of what's actually repeatable shift as a company accumulates more closed deals over time.
Pulling an honest, unbiased sample of your own closed-won deals is worth doing before trusting an ICP that might just reflect whichever wins were easiest to remember. Talk to purple path about building an ICP model from a properly sampled set of your real deals.

Dave leads purple path's content team, getting clients' inbound, outbound, thought leadership, social, and video content running fast, and making sure it actually works. In an AI-saturated content landscape, he's focused on the thing that still wins: content that engages and delivers real value.He's spent his career shaping content marketing strategy for SaaS companies globally, and previously as Head of Content at Minit Process Mining and Senior Copywriter at Exponea. He also built and exited his own company, Elite Language Center, over nearly nine years as CEO. His work has been featured in Forbes, and he's increasingly focused on LLM visibility, making sure content shows up where AI-driven search is heading next (GEO/AEO).