The Lead Scoring Guide That Ends With a Build, Not a Buying Decision
Why the 0 to 100 point model is the wrong shape, how to keep fit and intent apart, what belongs in negative scoring, where deterministic rules still beat a model, and how to define an MQL you can...
A lead scoring model should not be one number out of 100. It should be two scores, fit and intent, kept apart. Deterministic rules handle the facts you already hold. Narrow probability questions handle the free text. An explicit exclusion list handles everyone who must never reach sales. Then you build it.
Search the term and you get eleven vendor pages and a Wikipedia entry. Every vendor page teaches the same four steps, then replaces the fifth with "choose a platform." That is not an editorial accident. The build step is the product they sell. This post is the missing fifth step, and it starts by arguing that the shape those guides teach is wrong before you write a line of it. If you have not met typed decision models yet, start with our plain English explainer.
The short version
- The 0 to 100 points model is a tooling artifact, not a model of buying. It exists because a mid-2000s CRM could add and subtract integers on a record and nothing else.68
- Fit and intent are two axes. Collapsed into one number, a perfect-fit lead with no intent and a poor-fit lead with high intent get the same score and the wrong treatment.6
- Negative scoring should be an exclusion list, not a penalty. A competitor at 80 points is still a competitor.7
- Deterministic facts belong in rules. Rules are free, auditable and instant. A model earns its place on free text and judgment calls only.
- The better shape is several narrow questions, each returning a probability, weighted in your own code. One wide question scored 89.4 percent against 95.0 percent for the same judgment split into five, on identical held-out data.12
- Published MQL to SQL benchmarks disagree by roughly 3x. One puts it at 13 percent, another between 26 and 51 percent depending on channel. Measure your own.1011
- Speed is the best-evidenced finding in inbound sales, and two of the numbers everyone quotes for it cannot be traced to any published study.12
Why lead scoring became a points model in the first place
The practice comes from marketing automation platforms of the mid-2000s. The tooling could do exactly one thing with judgment: add or subtract an integer on a contact record when a field changed or a page loaded. VP in the job title, add 15. Free email address, subtract 10. Pricing page view, add 8.
That was a good idea for the constraint it was built under. It encoded judgment into something the CRM could execute at three in the morning. It was never a model of how people buy. It was the only executable form judgment could take.
Two things followed, and both are still with you. The weights were guesses. Nobody derived "add 15" from closed-won data, and everybody copied the shape from the guide they read. HubSpot's canonical explainer still teaches "award 20 points" arithmetic without a worked derivation.6
And everything collapsed into one integer, because a CRM list view sorts on one column. That constraint causes almost every failure in this topic, and nobody questions it.
Why fit and intent should never be the same number
Every guide teaches the two axes. HubSpot calls them fit and interest. ZoomInfo maps them onto BANT and CHAMP.9 Then almost every implementation adds them together and throws the distinction away.
Fit is whether this is the kind of company you serve. Industry, size, country, the work they described. It is roughly stable. A company that fits today fits in six weeks.
Intent is whether they are trying to buy now. Pricing page views, a stated timeline, a named budget, a demo request. It decays in days.
Add them and you get this. A perfect-fit enterprise account idly reading a blog post scores 55. A poor-fit sole trader who has read your pricing page four times today and asked for a call scores 55. Same row, same priority, completely different correct actions. The first needs a named account plan and no sales call. The second needs a reply in ten minutes, quite possibly so you can disqualify them quickly.
There is a quieter problem too. You averaged a stable quantity with a decaying one, so when the combined number moves you cannot tell which half moved. Score decay becomes impossible to reason about.
Keep two numbers. Route on the pair.
The four cells are not four priority levels. They are four different jobs. Only one of them justifies interrupting someone.
What belongs in negative lead scoring
Negative scoring is the part every guide mentions in one sentence and nobody specifies. Here is the list, from a real B2B inbound form.
Competitors. Match the email domain against a maintained competitor list. Also catch the message pattern: "we do something similar," "researching the market," a request for your process with no project attached.
Job applicants. They arrive constantly. Look for CV, resume, vacancy, portfolio, internship, "looking for a role," "open positions."
Students and researchers. Academic email domains, plus dissertation, thesis, "for my course," "a few questions for my research."
Vendors pitching you. The tell is that the message is about their service, not your work. "I noticed your website," "we help agencies like yours," partnership, guest post, backlink.
Free-email consumer enquiries on a B2B form. A personal address alone means nothing. A personal address plus a personal purchase is not a business lead.
Existing customers arriving through the new-business form. They need support routing, not a sales sequence.
The rule that matters more than the list: negative scoring is an exclusion, not a penalty. Subtracting 30 points from a competitor still leaves a competitor, one who climbs back up with a few pricing page visits. HubSpot's own tool keeps exclusion lists separate from the score for this reason.7 Positive scoring is graded. Negative scoring is binary.
Where rules still beat a model
This is the honest part, and it is missing from the vendor guides and from most of the AI commentary.
Some inputs are deterministic facts you already hold. Country from the form. The company size band already in your CRM. Whether this email belongs to an existing customer. Which page the form sat on. Whether consent was given.
Do not send any of that to a model. Asking a model "is this a UK company" when you have a country field is worse in every dimension. Slower, a cost per lead, and it can be wrong about a fact you are certain of.
Rules are free. They are auditable, so you can explain any outcome to the person who is annoyed about it. They are instant, which matters more than it sounds once you read the response-time evidence below. And you can test them with no training data.
The model earns its place on one class of input: the free-text field. "Tell us about your project" is where budget, timeline, seniority and the actual problem live, and no rule reads it. That is where the judgment calls sit, the ones a rule would need a hundred branches to approximate and still get wrong.
The practical split is a gate. Rules run first and cheaply. They exclude, they set fit from facts, and they decide whether the lead is worth a judgment call at all. Only what survives reaches the model.
Ninety second check
Is your job decision shaped?
Five questions. One no is enough to make this the wrong tool, which is worth finding out before you wire anything up.
Can you write down every possible answer before you run it?
Is the output a decision rather than something a person will read?
Does it happen often enough that doing it by hand hurts?
If a call is wrong, can you undo it cheaply?
Can you live without a written reason for each call?
Answer the five above and you get a straight verdict here.
The better shape: narrow questions, your weights on top
Replace the single 0 to 100 score with a set of narrow questions, each returning a probability, composed in your own code.
For fit, from the free text alone: does this describe a business rather than a personal purchase, is the writer plausibly the decision maker, does the described work match something you actually do, is there a stated timeline, is there a stated budget or range.
For intent: does the message ask for a specific next step, does it name a deadline, does it describe an active problem rather than a future plan.
Your code multiplies each probability by a weight you chose and sums them into the two axis scores. Three practical benefits follow.
You change weights without retraining anything. The weights live in your code. Decide next quarter that a stated timeline matters twice as much, change one number, redeploy. No training run, no vendor in the loop.
You can explain a rejection. "Fit 0.31, because the described work is not something we do and no decision maker is identified" is a sentence a salesperson can argue with. "Score 42" is not. That matters the first time someone says the model got one wrong, which will be within a week.
Each question can be tested in isolation. Take 200 leads from your history where you know the outcome and score each question separately. A broken question shows up as a broken question, not a vaguely disappointing composite.
Decomposition is also measurably better, not just tidier. An independent write-up on 20 September ran the same judgment two ways on identical held-out data. Asked as one wide question, the typed model scored 89.4 percent. Split into five narrow questions and recombined in code, it scored 95.0 percent.12 Different domain, so do not carry the numbers across. Carry the finding: how you split the question moved the answer more than 5 points, further than most model choices move anything. The same test had a frontier model moving the opposite way, which is why you run this on your own labeled leads.
The wiring, the form handler, the typed call, the CRM write, is covered in the implementation post. This post is about what to ask. That one is about how to run it.
How to define an MQL you can defend
An MQL is a promise to the person who picks up the phone: leads above this line are worth your next hour. Break that promise twice and the model is dead no matter how accurate it is.
A defensible definition has three parts.
A threshold on both axes, not one. "Fit above 0.7 and intent above 0.6" is a definition. "Score above 70" is a number.
The action it triggers. An MQL that means "appears in a weekly report" is a label, not an MQL. If clearing the line does not change who gets contacted and when, delete the line.
A review cadence with the number it is reviewed against. Monthly, against the share of leads above the line that became a real conversation.
Now the uncomfortable part. You cannot borrow a threshold from a benchmark, because the benchmarks do not agree. First Page Sage puts MQL to SQL conversion around 13 percent in B2B SaaS from aggregated client data.10 Its own directly tracked B2B SaaS funnel reports 26 percent for PPC and 51 percent for SEO.11 Same nominal metric, same publisher, roughly 3x apart, and no way to tell which one is measuring something else.
So measure your own. Take every lead that cleared your line in the last 90 days and ask what fraction became a real sales conversation. It takes an afternoon and beats every published benchmark in this topic.
Then set the threshold on capacity, not statistics. If you can handle 40 real conversations a month, put the line roughly where 40 leads a month clear it, and move it when capacity moves. That is the lead scoring best practice nobody writes down, because it is not a product feature.
Try it
Where would you draw the line?
4,000 judgments, each returned with a confidence. Move the two lines and watch how much work gets done without you, and what it costs you in wrong calls.
2,679
acted on automatically
67 percent of the run, with about 160 expected to be wrong.
1,104
queued for a human
28 percent of the run. This is the pile that decides whether the whole thing saves you time.
222
left alone
Too uncertain to be worth anyone's attention this round.
The lesson is in the second box. Push the accept line high enough to make the error count comfortable and the review queue grows until a person is doing the job again. The threshold is a business decision about how much a wrong call costs you, and it belongs in your code, not in the model.
What the evidence on responding fast actually says
The reason this has to run fast rather than in a nightly batch is old, well-sampled and consistent.
James Oldroyd's Lead Response Management study, run at MIT Sloan with InsideSales.com in 2007, covered three years of data across six companies, more than 15,000 leads and over 100,000 call attempts. Contacting a web lead within 5 minutes rather than 30 raised the odds of making contact by 100x and the odds of qualifying by 21x.1 Those are contact and qualification odds, not a revenue multiple, and they get quoted as one constantly.
The strongest single citation available is the 2011 Harvard Business Review article by Oldroyd, McElheran and Elkington, "The Short Life of Online Sales Leads," which audited 2,241 US companies with test web leads. Average first response was 42 hours. Only 37 percent responded within an hour, and 23 percent never responded at all. Responding within the first hour made a firm about 7 times more likely to qualify the lead than responding in the second hour, and about 60 times more likely than waiting a day or more.2 That is 2011 data, and you should say so when you quote it.
Drift's Lead Response Report in 2017 submitted real forms to 433 B2B SaaS companies. 7 percent responded within 5 minutes, and 55 percent had not responded within five business days.3 It is widely misdated to 2021, which tells you how carefully it is being cited.
XANT's 2021 audit is the most recent large-sample dataset, covering 5.7 million leads across 400-plus companies. Conversion was 8 times higher inside 5 minutes than after 6, under 1 percent of calls happened within 5 minutes, and 77 percent of leads got no response at all.4
Velocify's 2013 analysis of roughly 3.5 million leads across 400-plus companies found that calling within a minute lifted conversion by 391 percent, that 50 percent of leads never get a second call attempt, and that 93 percent of converted leads were reached by the sixth attempt.5 Vendor-published and dated, but the sample and method are stated, which is more than most of this topic manages.
Fifteen years of separate studies, the same finding, and response times that barely moved. That is the case for scoring inline rather than in a batch, and the latency side has its own post.
The two numbers we could not trace
Two statistics dominate this topic, and we could not find a primary source for either.
The first is "78 percent of buyers purchase from the first responder," almost always attributed to a Lead Connect survey. We went looking for the report. No published study, no stated sample, no method, no date on anything primary. The number circulates entirely through pages citing other pages citing it. The variant "35 to 50 percent of sales go to the vendor that responds first," attributed to InsideSales, has the same problem.
The second is "48 percent of salespeople never make a single follow-up attempt," attributed to the National Sales Executive Association. We could not establish that the organization exists. No register entry, no site, no publication history.
We are not repeating either as color. For the follow-up point, Velocify's 50 percent of leads never get a second call does the same job with a stated sample of about 3.5 million leads.5
We made this same point on 20 September in the 2026 lead stack post. It is worth making twice. A scoring model built to chase a statistic that does not exist is optimized for nothing.
What the model is actually for
Not ranking leads for a report. Not a dashboard. Not a number in a monthly deck.
A lead scoring model has one job: deciding who gets contacted first, while they are still interested. Judge every criterion against that. If a field changes the score but never changes who gets called in the next ten minutes, cut it. If the model produces a number nobody acts on before the lead cools, you built a report with extra steps.
That test also tells you when you are done. When the output changes what happens next, automatically, inside the response window, stop building. The vendor guides cannot end here. Yours can.
If you want this built against your real form and CRM rather than described, that is lead generation and conversion work. It starts with your last 200 leads.
Frequently asked questions
What is a lead scoring model?
A set of criteria that turns what you know about an inbound lead into a decision about how fast and by whom they get contacted. Traditionally it produces one number between 0 and 100. A better version produces two, fit and intent, plus an exclusion flag.
What should a lead scoring model actually score?
Two separate things. Fit is whether this is the kind of company you serve, built mostly from facts you already hold. Intent is whether they are trying to buy now, built from behavior and from what they wrote. Scoring them together destroys the distinction that decides the action.
What is negative lead scoring?
Identifying leads that should never reach a salesperson regardless of how good they otherwise look. Done properly it is an exclusion flag rather than a points penalty, because a competitor who scores badly can still climb back over the threshold by visiting your pricing page.
What belongs in a negative scoring list?
Competitors matched by email domain, job applicants, students and researchers, vendors pitching you a service, free-email consumer enquiries on a business form, and existing customers who arrived through the new-business form. Each is cheap to detect and expensive to miss.
What is the difference between demographic and behavioral scoring?
Demographic scoring uses attributes of the person and company: job title, industry, company size, country. Behavioral scoring uses what they did: pages viewed, forms submitted, time on the pricing page. Demographic scoring maps to fit, behavioral scoring maps to intent, and they belong on separate axes.
What score should make a lead an MQL?
There is no transferable answer, and the published benchmarks make that obvious. One aggregate puts B2B SaaS MQL to SQL conversion at 13 percent, another puts B2B at 40 percent across 939 companies. Set the threshold on how many real conversations you can handle each month, then check monthly what share of leads above the line became one.
Should I use rules or a model for lead scoring?
Both, split by input. Deterministic facts you already hold belong in rules, because rules are instant, free and auditable. Free-text fields and judgment calls belong in a model, because no rule reads a paragraph. Sending a known country field to a model is slower, costs more and can be wrong about something you are certain of.
How many lead scoring criteria should I have?
Fewer than you think, and each one should change an action. Remove a criterion and check whether any lead changes priority. If none does, it is decoration. Most working models run on a handful of rules plus a small set of narrow questions against the free text.
How do you know if your lead scoring model is working?
Take every lead that cleared your MQL threshold in the last 90 days and calculate the share that became a real sales conversation. Do the same for leads that fell just below the line. If the two numbers are close, the threshold is not separating anything.
Is the 78 percent first responder statistic real?
We could not find a primary source for it. It is attributed to a Lead Connect survey with no published report, no sample and no method. The same applies to the claim that 48 percent of salespeople never follow up, attributed to a National Sales Executive Association we could not show exists. Use Velocify's 50 percent of leads never get a second call instead.
- Dr James Oldroyd, Lead Response Management study, MIT Sloan with InsideSales.com, 2007. Three years of data, six companies, 15,000+ leads, 100,000+ call attempts. leadresponsemanagement.org
- James B. Oldroyd, Kristina McElheran and David Elkington, "The Short Life of Online Sales Leads," Harvard Business Review 89, no. 3, March 2011. Audit of 2,241 US companies. hbr.org
- Drift, "Lead Response Report," 2017. Real form submissions to 433 B2B SaaS companies. drift.com
- XANT (formerly InsideSales.com), "Lead Response Management 2021." 5.7 million leads across 400+ companies. insidesales.com
- Velocify, "The Ultimate Contact Strategy," 2013. Approximately 3.5 million leads across 400+ companies. Study PDF hosted on Salesforce AppExchange. appexchange.salesforce.com
- HubSpot, "How to Set Up Lead Scoring," marketing blog. blog.hubspot.com
- HubSpot Knowledge Base, "Understand the lead scoring tool," including exclusion lists, score decay and thresholds. knowledge.hubspot.com
- Wikipedia, "Lead scoring," explicit and implicit data, rule-based and predictive methodologies. en.wikipedia.org
- ZoomInfo, "Automated lead qualification," BANT and CHAMP mapped to scoring inputs. pipeline.zoominfo.com
- First Page Sage, "MQL to SQL Conversion Rate by Industry," data 2019 to 2025, agency-aggregated client data. firstpagesage.com
- First Page Sage, "B2B SaaS Funnel Conversion Benchmarks," MQL to SQL by acquisition channel, updated June 2025. firstpagesage.com
- XenoSpectrum, "Jev, TypeSafe and the BERT classifier: what decomposition does to accuracy," 20 September 2026. xenospectrum.com
- TypeSafe AI, "Introducing System One models and Jev," 15 September 2026. typesafe.ai
- Warmly, "AI lead scoring," vendor framework including score decay and failure modes. warmly.ai
Sources checked on 21 September 2026. Where a figure could not be traced to a primary study, it is named above rather than repeated.
Leads sitting in an inbox?
Tell us how enquiries reach you today. We wire scoring and routing into the form itself, so the good ones reach a human while the visitor is still on the page.
- Scoring and routing at form submit, not overnight
- Confidence thresholds so nothing is auto-binned silently
- Scoped estimate within 48 hours
Want to discuss non-tech founders for your business?
Start a project and we'll talk through where you are, what's working, and the highest-leverage moves for the next 90 days.



