Most Founders Guess Their ICP. Here's How to Test 100 of Them at Once.

Most founders pick an ICP the way they pick a name for their company.
It feels right. It makes sense on a whiteboard. It survives a few early conversations and gets locked in as strategy. Then the team builds around it. Content gets written for that audience. Ads get targeted at that persona. Partnerships get pursued in that vertical. A whole go-to-market built on a hypothesis nobody actually validated.
Six months later, pipeline is thin. The diagnosis is always murky. Was it the product? The message? The channel? The audience? Nobody knows because nobody tested the assumption underneath all of it.
This is the most common and most expensive mistake in early-stage GTM.
We see it constantly. A founder comes in knowing exactly what they sell. They have built something real. But when you push on the ICP, it is a guess dressed up as strategy. "We sell to mid-market SaaS companies." "Our buyer is a VP of Operations." "We do best with companies that have 50 to 200 employees." Logical but, untested.
One of our clients sells AI voice agents. They came in with eight ICPs. Marketing agencies, home service operators, healthcare practices, insurance outfits. In case you couldn't tell, these are very broad categories that made sense at the category level and told you almost nothing about who to actually email, what to say, or what pain to lead with.
We did not help them pick the right ICP. We built a system to let the market tell them.
Outbound is the fastest market research tool that exists
A paid ad campaign takes weeks to set up, burn budget, and produce signal. A content strategy takes months to build an audience before you know if the audience you built is the one that buys. A partnership takes a quarter to develop and another quarter to find out whether it moves pipeline.
Outbound at scale gives you signal in weeks. Not vanity signal. Real signal. Which audiences reply. Which pain angles convert. Which sub-niches ignore you after a thousand sends. That is a ranked, validated map of your market built from actual buyer behavior, not from your best guess in a strategy doc.
The insight that changed how we built this engine: outbound is not just a pipeline tool. It is the cheapest way to run market research at scale. Most teams use it only to book meetings. We use it to find out which version of the market is real.
The problem with eight ICPs
Eight ICPs sounds like a plan. In practice it is eight different markets, each with a dozen sub-niches underneath it, each with its own vocabulary, its own pain hierarchy, its own title structure, and its own buying process.
"Marketing agencies" is not an ICP. It is a category. Inside it lives the performance marketing shop buying leads for insurance agents. The white-label SEO reseller. The boutique brand consultancy. The agency running paid social for e-commerce brands. The lead generation operation running call centers for mortgage brokers.
Each of those is a different buyer. Different vocabulary. Different pain. Different person who would open the email. Writing one campaign across all of them means writing a campaign that resonates with none of them.
So the first move was expansion. Eight ICPs became roughly 100 sub-niches. Not invented sub-niches. Sub-niches drawn directly from the market structure that already exists.
The ICP Strategy Brief
Before any code runs, each vertical gets a written ICP Strategy Brief. Six sections:
Target Account Definition what these companies ARE vs the look-alikes around them
Buyer Personas exact titles per sub-segment
Pain Map named pains in the ICP's own language
Foot-In-The-Door Use Case the first workflow — the actual offer
Good-Fit Filters what makes a company pass
Disqualifiers what gets them permanently suppressed
Teh brief is not a positioning document. It is the source of truth the engine runs against. Every gate prompt, every research call, every copy prompt traces back to it. The engine cannot invent a segment, a pain, or a title that is not in the brief. If the brief does not say it, the pipeline does not say it.
The brief is also what makes the system transferable. Seven more verticals are waiting in a folder. Each one gets a brief. Each brief feeds the same pipeline. The infrastructure does not change.
Voice of ICP research
Once the brief exists, the pipeline runs a Voice of ICP exercise on each sub-niche.
This is not asking an LLM what the audience cares about. That produces marketing language — polished, plausible, and completely wrong in the way that matters. Real resonance comes from the exact words a buyer uses when complaining to peers, not the words they use in a vendor meeting.
The exercise mines Reddit threads, G2 reviews, and LinkedIn comments. Not for themes. For phrases.
phrases:
- they_say: "quote rate"
we_say: "conversion rate"
quote: "Quote rate is 22% or so. Lead close rate is 1.6%."
source: "reddit.com/r/InsuranceAgent"
role: "independent insurance agent, buys ~100 leads/day"
- they_say: "tire-kickers"
we_say: "unqualified leads"
quote: "Are they actually qualified, or mostly tire-kickers?"
source: "reddit.com/r/CFP"
role: "financial advisor evaluating a lead vendor"
- they_say: "credit or replace bad leads"
we_say: "buyer rejections"
quote: "Do they credit or replace bad/bogus leads, and how easy is that process?"
source: "reddit.com/r/CFP"
role: "financial advisor evaluating a lead vendor"
Every phrase has a source and a role. Not a summary. A quote from a real person doing the real job.
When the email says "quote rate" instead of "conversion rate" to an insurance operator, it is because a real insurance operator used that term when complaining about their vendor. That is the difference between personalization and actual resonance.
The gate and the kill switch
With 100 sub-niches in play, two things become critical: getting the right companies in, and getting the wrong campaigns out.
The gate handles the first problem. A cheap model classifies every company as pass or fail before a contact is touched. Hard disqualifiers — competitors, government entities, offshore operations, software vendors — go to permanent suppression. Not dropped from this run. Gone.
The kill switch handles the second. Every campaign has a threshold. Reach 1,000 sends with zero interested replies and the campaign is killed automatically. Not paused. Killed.
def cell_verdict(sent, interested, meetings, kill_threshold):
if sent < min_read_sends:
return "pending"
if meetings >= 1:
return "scale"
if interested / sent >= 0.005:
return "scale"
if sent >= kill_threshold and interested == 0:
return "kill"
return "iterate"
One important distinction: a dead campaign does not mean a dead audience. A slice only gets called dead after two distinct offers have been killed against it. One bad campaign means a wrong message. Two bad campaigns means a wrong audience. That distinction is what stops you from abandoning a real market because your first message was off.
Every Friday a read-only scorecard posts to Slack. Named repliers. Company. Title. Which pain their email led with. Which campaign. Every Monday the system acts on it — killing what is not working, scaling what is, and refilling the queue with new sub-niches to test.
What the output actually is
After a full cycle, you do not just have pipeline. You have a ranked, evidence-backed map of your market
You know which sub-niches reply. Which ones ignore you at scale. Which pain angles convert by vertical. Which titles open and which titles ghost. Which offers land and which ones get flagged as spam. Not as a guess. As data from thousands of sends across dozens of campaigns.
That is the asset. And it does not stay inside the outbound motion.
How it flows downstream
The marketing team now knows which pain angles are resonating. Content gets written around those angles, not around what felt right in a brainstorm. The LinkedIn posts lead with the exact language that produced replies in outbound, because that language was validated against real buyers before it was published.
Paid ads target the sub-niches that replied, not the broad categories that looked right on a whiteboard. The audience that converted in outbound is the audience you build lookalikes from in paid.
Partnerships get pursued in the verticals where reply rates are highest, because those are the verticals where the product is solving a real problem the market is already feeling.
Referral programs get built around the ICP profiles that produced meetings, not the ones that sounded like a good fit.
Every channel gets smarter because one channel did the work of finding out what is real.
Where this goes
Eight ICPs. One hundred sub-niches. One vertical tested in a quarter.
Seven more briefs in the folder. Each one gets fed into the same pipeline. The gate learns the new vertical. The Voice of ICP research runs. The copy adapts. The kill switch works the same way.
The reply rate went from 6% to 11% in one cycle. Not because one campaign was optimized. Because every week the targeting got more precise, the copy got closer to what the right person actually responds to, and the audiences that were never going to convert got removed from the queue before they wasted another thousand sends.
Most founders guess their ICP. The market already knows the answer. You just need a system fast enough to ask it.
Related articles

Set up 10 female accounts: print $14,000 in 28 days (b2b)
a message from a female account gets opened before the brain even decides to open it in cold outreach

how to use GPT-6 Astra with Grok Bot to create a $3M/yr GTM team
A cold email campaign is a chain of artifacts. a verified lead list. a hypothesis about what to test. a sequence of copy. a campaign built inside a sending tool. a daily read on sending health. each…

Grok Bot Agents: How to Build a 5-Person Sales Team That Never Sleeps (Full Guide)
I'm going to show you the exact 5-agent setup that's been running our lead gen, and the descriptions that go inside each one.