Skip to main content
v2026.11,610 entries · CC-BY 4.0

Vendor Scorecards: What Metrics Actually Belong on One

A practical vendor scorecard framework that separates metrics a facility can calculate from its own PO, receiving, and AP records from metrics that only exist because the vendor reported them.

Ask about Vendor Scorecards: What Metrics Actually Belong on One

Answers are drawn from this guide and the rest of the CASRAI corpus, with a link to every source.

Answers are AI-generated from CASRAI’s own published pages and can be wrong, so check the linked sources before relying on one; your question is logged without personal data — never sold, never used to train a third-party model — to show us what CASRAI is missing, so please do not type personal or confidential details. How we use this

Written and maintained by CASRAI Editorial Board

Last updated

Most vendor scorecard templates treat every metric on the page as equally trustworthy. They aren’t. Some numbers on a scorecard come straight out of your own purchase orders, receiving logs, and accounts-payable records — you can recalculate them yourself, from your own data, at any time. Others exist only because the vendor reported them to you, in a quarterly business review or a self-published dashboard, and there is no independent record on your side to check them against.

That distinction matters more than which metrics you pick. A scorecard built entirely from numbers you can verify is smaller than the generic “track everything” checklists suggest, but every entry on it means something. A scorecard padded with vendor-reported figures you can’t check looks more comprehensive and tells you less.

The Two Kinds of Vendor Metric

Facility-sourced metrics come from data you already generate in the normal course of ordering and receiving — your purchase orders, your receiving/put-away records, your accounts-payable three-way match, your own issue log. You don’t need the vendor’s cooperation to calculate them, and you don’t need to trust their math.

Vendor-sourced metrics exist only in the vendor’s own systems — fill rate measured at their distribution center, internally categorized root causes for a service failure, an “average response time” figure calculated from their ticketing tool. You can ask for these, and for a high-volume or clinically critical supplier it’s often worth asking, but they belong on the scorecard labeled as what they are: reported, not measured.

Metrics You Can Calculate From Your Own Records

These five are the practical core of a scorecard a facility can run without depending on a vendor to hand over anything.

On-Time Delivery Rate

Deliveries received by the requested date, divided by total deliveries, over a rolling period. The detail that gets lost in generic templates: measure against the date you asked for on the PO, not against the date the vendor confirmed after the fact. If a vendor pushes back a promised date and then hits that revised date, that’s a real operational event worth knowing about — but scoring it as “on time” against their own moved goalpost quietly inflates the number. Track both dates if your system allows it: on-time-to-request and on-time-to-confirmation are two different metrics, and the gap between them is itself useful information about how often commitments slip after the order is placed.

Order Accuracy Rate

Line items received exactly as ordered — correct item, correct quantity, correct lot or expiry within acceptable range — divided by total line items received. This comes directly from your own receiving discrepancy log; every facility that receives goods against a PO is already generating the raw data, whether or not anyone is currently rolling it up into a rate.

Backorder Frequency

Line items backordered divided by total line items ordered, over the same rolling period. Pull it from your own open-PO or unfulfilled-line report rather than a vendor-supplied fill-rate figure, because a vendor’s fill rate is typically calculated against their stocking targets, not against what you actually ordered and when you needed it — the two numbers can diverge without either party being wrong by their own definition.

Invoice-Accuracy Rate

Invoices that match the PO/contract price and quantity without requiring a correction, divided by total invoices, drawn from your AP three-way-match exceptions. This is one of the more reliably self-sourced metrics on any scorecard, since a match or mismatch is a binary fact your own AP system already records for every invoice that passes through it — and it’s a useful check on whether contracted pricing (see chargebacks and contract pricing) is actually being applied at the invoice level, not just agreed to on paper.

Responsiveness to Issues (the self-trackable part)

This one needs a specific caution: “responsiveness” is almost always presented as a vendor-reported average response time, and that number is not independently checkable unless you also keep your own log. The fix is simple — timestamp when you reported an issue (an email, a ticket, a phone call logged in your own system) and timestamp when it was actually resolved to your satisfaction, not when the vendor marked it closed. That gap, measured entirely from your own records, is the responsiveness metric worth having. It will often run longer than whatever average the vendor quotes in a business review, and the difference is informative on its own.

Metrics That Require the Vendor’s Own Data

These aren’t worthless, but they need a different label on the scorecard and a different level of trust.

  • Fill rate at the vendor’s distribution center. You see what didn’t ship to you; the vendor sees why — a raw-material shortage, an allocation decision, a demand spike from another customer. Their fill-rate number is real from their side, but you’re trusting their accounting of it.
  • Root-cause categorization. When a vendor reports that a shipping delay was “carrier-caused” versus “inventory-caused,” that categorization happens entirely inside their own process, with no visibility from your side into how consistently it’s applied.
  • Quality or defect rates below your own detection threshold. Your receiving inspection catches what your receiving inspection is built to catch. A vendor’s internal QA process may catch near-misses or systemic issues your own team never sees — useful context, but not something you can verify without an audit of their process (see what a supplier audit actually covers).
  • Vendor-reported “average response time.” As above — ask for the underlying ticket-level data if this metric matters enough to act on, rather than accepting the summary average at face value.

A Minimal Scorecard That Doesn’t Require Vendor Cooperation

The five facility-sourced metrics above are enough to run a functioning scorecard for every active vendor, including ones too small or too new to have a formal QBR relationship with. None of them require the vendor to send you anything:

Metric Source Typical cadence
On-time delivery rate Your PO + receiving records Monthly rolling
Order accuracy rate Your receiving discrepancy log Monthly rolling
Backorder frequency Your open-PO report Monthly rolling
Invoice-accuracy rate Your AP three-way match Monthly rolling
Responsiveness to issues Your own issue/ticket log Per-incident, reviewed quarterly

If a facility is only going to build one version of a scorecard, this is the one to build first — it works the same way for a single-item niche supplier as it does for a full-line distributor, and it doesn’t fall apart the moment a vendor stops sending a QBR deck.

When to Add Vendor-Reported Metrics Anyway

For high-volume or clinically critical suppliers, it’s still worth requesting the vendor’s own fill-rate and root-cause data — it surfaces problems upstream of what your own receiving data can see. The discipline that keeps a scorecard honest is presentation, not exclusion: keep vendor-reported figures in a clearly separate section or column, don’t blend them into the same average as your facility-sourced numbers, and note the reporting period and source each time. A scorecard that silently mixes “we measured” and “they told us” numbers into one blended score has quietly stopped being a measurement tool.

How Often to Score, and What Actually Triggers a Conversation

Monthly rolling calculation for high-volume or critical suppliers is usually enough to catch a developing trend before it becomes a stockout; quarterly is reasonable for low-volume or non-critical vendors, where month-to-month noise in a small number of orders isn’t meaningful. What matters more than the cadence is having a pre-agreed threshold that triggers an actual conversation — a sustained drop in on-time delivery or a rising backorder rate over two consecutive periods, rather than a single bad month, which is often just normal variance in a small sample. Reserve escalation past the account rep (see when to escalate a vendor service failure) for a pattern the scorecard actually shows, not a single incident.

Related Reading

Frequently Asked Questions

What’s a reasonable benchmark for on-time delivery rate?

There’s no single defensible industry-wide benchmark, because “on time” is defined differently across scorecards (against request date versus confirmed date, as above) and because acceptable variance differs by product criticality. The more useful exercise is establishing your own baseline for each vendor over several months and watching the trend, rather than importing a benchmark number from outside your own ordering pattern.

Should we track fill rate ourselves, or rely on the vendor’s number?

Track backorder frequency from your own open-PO records instead of relying solely on the vendor’s fill-rate figure — the two aren’t measuring the same thing, since a vendor’s fill rate is typically calculated against their own stocking targets rather than your specific order pattern. Use the vendor’s fill-rate number as supplementary context, not as your primary metric.

Do small or occasional-use vendors need a full scorecard?

Not a full one. The five facility-sourced metrics scale down fine — even a handful of orders a year still generates a real on-time and accuracy record — but the quarterly-conversation cadence and vendor-reported section are usually only worth the overhead for vendors supplying critical or high-volume items.

What if we don’t have dedicated procurement or receiving software?

All five facility-sourced metrics can be tracked from a basic spreadsheet fed by PO copies, receiving slips, and AP records — the metric doesn’t require special software, it requires that someone consistently logs the requested date, the received date, and any discrepancy at the point of receiving. That habit matters more than the tooling.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Regulatory Radar

Stop finding out after the fact

$29/month, cancel anytime. Daily digest updates from our analysis, a dashboard holding the same items, and a cited assistant for everything they raise.

  • Federal Register, Federal Register+, Grants.gov, Regulations.gov, NSF News, UKRI, plus CASRAI’s own published content.
  • 44,322 indexed passages, and every answer cites the ones it drew on.