#The challenge
Targeting was intuition dressed up as data. The account attributes that most strongly predicted deal size were the ones the CRM could not see: the field capturing a category-defining platform attribute was populated on under a fifth of records, the volume metric that best predicted deal size was missing on about two-thirds, and the competitor field was effectively unpopulated. At the same time the company was expanding out of its established vertical into an adjacent one where it had no ICP definition, no persona matrix, and no way to separate a good-fit account from a bad one before a rep spent time on it.
#The approach
Mine closed-won for the signals that actually predict deal size
Rather than asking sales what mattered, we analysed roughly 400 closed-won deals and looked for account attributes that correlated with deal size. The output was a ranked signal set — the attributes that tracked reliably with larger deals, ordered by strength — paired with a fill-rate table showing exactly which of those signals the CRM currently could not see. The pairing is the point: a strong predictor that is 20% populated is worth more enrichment effort than a weak predictor that is 80% populated.
Turn each predictive signal into a specific sourcing recipe
For every high-value-but-empty field we wrote a named path rather than a generic 'enrich the account': job-posting scans for administrator titles tied to the platform in question, technographic detection on the prospect's web properties, press-release monitoring for adoption announcements, and a public filing for the volume metric. Each recipe terminates in three writes to the CRM — the enriched value, a confidence grade, and an evidence URL — so a rep can check the source and an ops lead can audit the model instead of trusting it.
Build the first tiering model from competitors' published customer lists
The initial look-alike universe came from the publicly named customers and testimonial logos of five direct competitors. Each was enriched for employee count, revenue band and tech-stack breadth, then clustered into three tiers: a mid-market sweet spot with enough process pain to buy but not enough scale to build in-house, an enterprise/logo tier valuable mostly as social proof, and a long-tail plus adjacent-vertical tier.
Stress-test the tiering against a real buyer population before shipping it
We ran the tier definitions against a live list of companies drawn from the target vertical. The sweet-spot tier captured only about a quarter of them, which for a vertical-specific list is a failing grade. Three concrete defects: the employee floor was set too high, so genuinely mid-market firms fell into a gap between tiers; the revenue ceiling was set too low, so companies that behaved operationally like the sweet spot were being classified enterprise; and adjacent-vertical companies made up close to a third of the list against a tier definition that barely accounted for them. Widening the bands and folding in the adjacent segment lifted modelled sweet-spot coverage to roughly 35–40% — a re-count of the same list under the revised rules, not a second validation run.
Grade the source data before trusting any tier
About one record in ten carried a data defect that would have mis-tiered it: revenue values off by three orders of magnitude, blank sub-vertical labels, and employee-to-revenue ratios that were physically implausible. Two fixes went into the model — a revenue-per-employee sanity check, and the removal of OR-logic from tier rules, because an OR rule was classifying a company with well over a thousand employees as 'boutique' on the strength of one bad revenue figure.
Stand up the new-vertical audience with a suppression list and a one-way CRM sync
For the adjacent vertical we built an audience from a core list plus two behavioural signals — hiring activity and an agent-run check on AI investment. The CRM connection was deliberately asymmetric: import sync on, export sync off, overwrite never, so nothing could write back to the CRM until the client had reviewed the audience. A suppression list of existing customers was loaded so the audience was net-new by construction rather than by cleanup.
Kill a signal that could not be sourced instead of faking it
A planned technology-footprint signal — detect whether a prospect was already running a competing product — returned nothing across two technographic providers for the entire competitor set. Those products simply do not leave a detectable footprint on a prospect's web properties. We recommended dropping it from v1 rather than shipping a signal with no data behind it, and said so explicitly in the status update.
Layer enablement onto the surviving list before handing it to reps
The vertical audience was narrowed from roughly 700 sourced companies to about 40 qualified accounts. Those 40 then got a research layer: per-account insight notes, persona and title definitions, an explicit split between who owns the buying decision and who is the champion, and draft subject lines and messaging. The full table was held back — only ten rows were run first so the client could judge output quality before credits were spent on the rest.
#Outcomes
A signal-ranked ICP rather than an opinion-ranked one
Account attributes ranked by their observed relationship to deal size, cross-referenced against CRM fill rates, producing an ordered enrichment backlog instead of a wish list.
A tiering model that survived contact with a real market
The first tier definition captured about a quarter of a genuine buyer population; re-counting the same list under the revised, sanity-checked rules put it at roughly 35–40%. The source states that as a range, and it is carried as a range here rather than as its upper bound.
A qualified, net-new audience for a vertical never sold into systematically
A sourced long list of roughly 700 companies was narrowed to about 40 qualified accounts with research, personas and messaging attached, plus a documented pattern for standing up the next segment the same way. Both counts are from LeanScale's own list build, walked through on a recorded review with the client.
A documented dead end
The competitor-footprint signal was proven unsourceable across two providers and dropped from v1 rather than left in the plan as an unbuilt promise.