
TL;DR: A HubSpot report can be technically accurate and still be lying, because the numbers it pulls depend on assumptions nobody checks: duplicate contacts inflating reach, first-touch attribution crediting the wrong channel, lifecycle stages that don't match how sales actually works a deal, and form submissions counted as new leads when they're really repeat contacts filling out a different form. None of these show up as an error. They show up as a report that looks clean and is quietly wrong in a specific, traceable way.
A demand gen report pulled straight from HubSpot looks authoritative. Clean charts, exact percentages, a specific number of new contacts this month versus last. The report's precision is exactly what makes its underlying errors so hard to catch, because a wrong number presented with three decimal places still reads as more trustworthy than a right number presented as a rough estimate.
HubSpot faithfully reports whatever it's been told, and it has no way of knowing whether what it's been told reflects reality. A duplicate contact record, a mislabeled lifecycle stage, or an attribution model nobody has reviewed since it was set up two years ago all produce numbers HubSpot will present with total confidence. The platform isn't lying. The data feeding it is, and HubSpot has no mechanism for flagging that on its own.
Duplicate contact records are the most quantifiable of the four, and the most tedious to actually clean up. HubSpot's own deduplication tools catch exact email matches reasonably well, but they miss the messier cases: a contact who signed up once with a work email and again later with a personal one, or a contact whose company changed and created what looks like a new record entirely. A quick way to get a rough sense of the scale: pull the count of contacts associated with your top twenty target accounts and manually check for names that appear twice under slightly different email formats. Finding even a handful across twenty accounts is a strong signal the company-wide number is inflated by more than a trivial amount.
A duplicate contact inflates a vanity number, annoying but not directly costly. A stale attribution model actively misdirects decisions, because it tells a founder or CMO to spend more on the channel it's crediting and less on the one it's ignoring, when the underlying reality might be the reverse. purple path's framework for demystifying marketing attribution in HubSpot covers exactly this risk: an attribution model set up once, with a specific touchpoint window and crediting logic, doesn't automatically stay accurate as the buying journey and campaign mix evolve around it. A model built for a shorter sales cycle two years ago, still running unchanged against a cycle that's since lengthened, systematically undercredits the early-funnel channels that actually started the deal.
Lifecycle stage lies are subtler than duplicates or attribution, because the stages themselves look reasonable in isolation. The test that actually catches this: pull five deals that closed last quarter and manually trace, from the CRM's own activity log, exactly when each one moved from lead to MQL to SQL to opportunity, and compare those dates against what the sales rep on each deal remembers about when the deal actually reached each of those points. A mismatch of a few days here and there is normal. A mismatch of weeks, or stages that were clearly backdated to make a report look better for a specific month, means the lifecycle stage data isn't measuring what the report assumes it measures.
None of this is a hypothetical data-hygiene concern; it's the direct reason demand generation programs get judged on numbers that don't reflect their actual performance. purple path's analysis of why MQL counts don't measure ROI in sales-led B2B SaaS makes a related point from the metric-selection side: even the right metric, tracked against corrupted underlying data, produces a wrong answer that still looks clean on a dashboard. Choosing the right thing to measure and having trustworthy data feeding that measurement are two separate problems, and fixing only one still leaves a report that lies.
A recurring, scheduled check catches most of this before it compounds into a quarter's worth of bad reporting. Once a month, pull the current duplicate contact count using HubSpot's built-in deduplication tool and compare it against the prior month's count, flagging any meaningful jump. Review the attribution model's touchpoint window and crediting rules against the current average sales cycle length, adjusting if the two have drifted apart. Manually trace three to five recently closed deals through their lifecycle stage history, checking for backdating or inconsistent stage jumps. None of these three checks takes more than an hour combined, and running them monthly catches drift while it's still a small, easily corrected problem rather than a quarter's worth of misleading numbers.
A HubSpot partner genuinely accountable for demand generation results should be running this kind of integrity check as a standing practice, not waiting for a client to request it. purple path's guide to choosing a HubSpot demand gen partner covers what a genuinely competent partner relationship looks like; part of that competence is a partner who flags a stale attribution model or a duplicate record problem proactively, rather than reporting the inflated numbers those problems produce without comment because the numbers happen to look favorable.
Finding a lie in the data is only useful if it leads to a fix, not just a footnote. A duplicate record problem gets fixed through a deliberate deduplication project, ideally run once thoroughly rather than repeatedly patched. A stale attribution model gets fixed by revisiting the touchpoint window and crediting logic against the current sales cycle, documented clearly enough that the next person to touch the CRM understands why the model is set up the way it is. A lifecycle stage mismatch gets fixed by sitting down with sales and rebuilding agreement on what each stage actually means, then holding both sides to updating records consistently going forward, since a one-time fix without a change in ongoing behavior just reintroduces the same drift a few months later.
A lighter monthly check, covering duplicates, attribution drift, and a small sample of lifecycle stage traces, catches most issues early. A deeper, more thorough audit is worth running at least once a year, or after any major change to the sales process or product line.
HubSpot's built-in deduplication tool catches many exact matches but misses fuzzier duplicates, like the same person under a work and personal email, or a name change after a company acquisition. A fully clean contact database usually requires some manual review beyond what the automated tool catches on its own.
Both can mislead if the model is stale, just in different directions. First-touch models tend to overcredit awareness-stage channels; last-touch models tend to overcredit the channel active right before a deal closed, regardless of what actually built the relationship earlier in the cycle.
Ideally a RevOps function with clear ownership of data integrity, since marketing has a natural incentive to see reports that reflect well on its own programs, even unintentionally. A neutral party checking the data reduces that risk.
Yes, arguably more, since a smaller company's decisions are more sensitive to a wrong number. A large enterprise can absorb a slightly miscalibrated attribution model without derailing strategy; a Series A company reallocating its entire demand gen budget based on a lying report has much less room for that mistake.
Running this kind of integrity check on your own HubSpot instance is a half-day project that can save a full quarter's worth of budget decisions made on bad information. Talk to purple path about auditing your current reporting for exactly these four issues.

Dave leads purple path's content team, getting clients' inbound, outbound, thought leadership, social, and video content running fast, and making sure it actually works. In an AI-saturated content landscape, he's focused on the thing that still wins: content that engages and delivers real value.He's spent his career shaping content marketing strategy for SaaS companies globally, and previously as Head of Content at Minit Process Mining and Senior Copywriter at Exponea. He also built and exited his own company, Elite Language Center, over nearly nine years as CEO. His work has been featured in Forbes, and he's increasingly focused on LLM visibility, making sure content shows up where AI-driven search is heading next (GEO/AEO).