The obvious way to score a vendor
Take a capability checklist, mark each item yes or no against a vendor, and score the percentage confirmed. It's the instinctive way to turn a qualitative evaluation into a number a steering committee can compare across shortlisted products. It's also quietly broken, and we found the break in our own internal tooling before we found it in anyone else's.
What actually happened
We run an internal vendor-ranking model for evaluation work: a confidence-weighted score across a capability checklist, meant to reflect not just what a vendor can do, but how thoroughly that was actually verified. Under an earlier, simpler version of the formula, a vendor whose capabilities had only been lightly checked (a handful of items, all confirmed "yes") scored a clean 100%. Two other vendors in the same comparison, each far more thoroughly researched across the full checklist, had each honestly turned up one "partial" answer somewhere in that deeper coverage. Both scored below the thinly-checked vendor.
The formula wasn't rewarding the best product. It was rewarding whoever had been asked the fewest questions.
Nothing about this was a data error or a one-off fluke. It's the structurally correct output of "score = percentage of checked items marked yes," applied to two vendors with genuinely different amounts of scrutiny behind their number. A vendor checked against 5 of 40 possible capabilities, all yes, will always beat a vendor checked against all 40 with one honest gap: the formula has no way to know the first vendor simply hasn't been asked the other 35 questions yet.
The fix
The real problem was treating "not yet checked" and "confirmed working" as interchangeable, since an unchecked item simply doesn't enter a plain percentage-of-yes calculation at all. The fix weights the score by how much of the full capability list was actually verified, with a coverage floor underneath it, so a vendor can't reach a top score on a handful of lucky confirmations. An unresearched capability now costs something in the score, deliberately less than a confirmed "no" would, since "we haven't checked yet" and "we checked and it doesn't do this" are genuinely different findings and shouldn't be scored as if they were the same failure, but it's no longer free.
Why this matters if you're the one reading a vendor comparison
This isn't really a story about a scoring formula. It's a story about what to ask when someone hands you a vendor comparison with a percentage attached to each name: what proportion of the full requirement set does that percentage actually cover? A 100% score against 6 checked capabilities and a 92% score against 40 are not comparable numbers, even though one looks better at a glance, and a comparison that doesn't disclose its own coverage is the exact shape of comparison that would have quietly favoured the wrong vendor in our own case before we caught it.
A question worth asking of any vendor scorecard you're handed: "what percentage of our actual requirements does this score reflect coverage of, not just confirmation of?" If the answer isn't readily available, the score is measuring confidence in incomplete information, not the vendor.
Key takeaways
- A plain "percentage confirmed yes" score structurally rewards whoever's been researched the least: it has no mechanism to penalise an unasked question.
- The fix isn't more scrutiny alone, it's weighting the score by how much of the full requirement set was actually covered, with a floor under how little coverage can still produce a high score.
- An unresearched capability and a confirmed failure are different findings and should cost a vendor differently, but "unresearched" should never cost nothing.
- Any vendor scorecard handed to you is worth one direct question: what's the coverage behind this number, not just the confirmation rate.
