Why don’t our reps trust the lead scoring?
Because the score is a number they cannot interrogate, built mostly from behaviour that correlates with engagement rather than with buying, and rarely re-derived when the account changes underneath it. Reps find its blind spots faster than anyone measures them. Trust is not restored by retuning the weights — it is restored by making the score inspectable and keeping it current.
The distrust is usually correct
It is tempting to treat rep scepticism as a change-management problem. It rarely is. Reps talk to the accounts every day, which means they receive feedback on the score's accuracy faster and in more detail than the team that built it. When they say it is wrong, they are typically reporting a real observation through the only channel available to them.
Three structural causes, in rough order of how much damage they do.
1. It measures engagement, not buying
Most scores are assembled from what the marketing stack can observe: emails opened, pages visited, content downloaded, webinars attended. These correlate with engagement. They correlate far more weakly with purchasing, and in some segments they correlate with people who will never buy — students, competitors, consultants, the curious.
Meanwhile the things that genuinely predict buying are frequently invisible to the model: a relevant hire, a reorganisation, a budget cycle, a technology being removed, a champion arriving from a company where your product was already in use. A rep knows those matter. The score does not include them. The gap between the two is exactly the gap in trust.
2. It is an assertion, not an inference
Most scores are computed on a schedule and then stored. Between runs, the number persists as a fact about the account — even after the account has changed underneath it. The champion left in April and the score still says 84 in June.
A score should be an inference: a current conclusion the system re-derives when new evidence arrives. Most are assertions: a value someone computed once, which the system has no mechanism to notice has become wrong. This is the same defect that makes the CRM diverge from reality, and it has the same cause.
A useful test: pick an account that changed materially last month. Ask when its score last moved, and what moved it. If the answer is "the Tuesday batch," the score is a historical record wearing the costume of a current assessment.
3. It cannot be interrogated
A rep looking at an 84 cannot ask why. There is no breakdown, no dated list of contributing evidence, no way to see that sixty of those points came from one person downloading three PDFs in 2024. Without inspectability, the rep has exactly two options: accept the number on faith, or ignore it. Most experienced reps ignore it, and they are making the more defensible choice.
Inspectability also matters for the team that owns the model. A score nobody can decompose is a score nobody can debug, which is why these systems tend to drift for years without anyone being able to prove they have.
Why retuning the weights does not work
The standard response to distrust is a recalibration project: revisit the weights, drop some behaviours, add others, relaunch with a communication plan. It buys a quarter, occasionally two.
It fails because it treats the output as the problem. The three causes above are properties of the architecture — what evidence is available, whether conclusions are re-derived, whether reasoning is visible. New weights change none of them. The recalibrated score is a different number with the same three defects, and reps rediscover them on roughly the schedule you would predict.
What restores trust
- Show the evidence. Alongside any score, the dated observations that produced it. A rep who can see the reasoning can correct it — and will.
- Re-derive on new evidence. The conclusion should update when the world does, not when the batch runs.
- Widen the inputs beyond your own funnel. Hiring, structure, technology, timing. What predicts buying mostly happens outside your marketing stack.
- Separate the two jobs. Use the score to rank a list. Use a specific, dated observation to justify a conversation. One number cannot do both.
- Close the loop. Feed outcomes back so the model learns which evidence actually predicted movement, rather than which evidence was easy to collect.
The first four layers of the GTM Architecture Audit — data, state, signal and decision — are the ones that determine whether a score can be trustworthy at all. A scoring model is only ever as good as the state model underneath it.
Would a predictive or AI-based model fix it?
It changes how the number is produced, not the three things causing the distrust. A predictive model trained on historical conversions still outputs an uninterrogable number, still learns from engagement behaviour if that is what the data records, and still goes stale between runs. It can be more accurate and no more trusted, which is a genuinely common outcome.
Our reps ignore the score but complain when we remove it. Why?
Because the score is doing a political job rather than an analytical one — it justifies territory decisions and pipeline reviews. That is worth knowing, because it means the fix is not purely technical. A score nobody acts on but everybody cites is a symptom of a decision layer that has no other way to explain itself.
How often should a score be recalculated?
The question assumes a batch model. The better target is that a score is re-derived whenever new evidence arrives about that account, so it is a current inference rather than a periodic snapshot. If that is not achievable, the interval should at least be shorter than the rate at which your accounts meaningfully change.
Is scoring worth keeping at all?
Yes, for prioritising a large list — that is what compression is good for. It should not be the thing handed to a rep as a reason to act. Prioritisation and justification are different jobs, and most implementations fail because one number is asked to do both.