Pull up the sales call scorecard your team uses and look at the weights. Objection handling, 30 percent. Discovery, 25. Opening, 15. Somebody chose those numbers. Ask where they came from and the honest answer, almost every time, is that they felt about right.
We build the dialer, CRM, and coaching tools that sales teams run their day on, so we see a lot of scorecards. The useful ones and the decorative ones look nearly identical on the page. What separates them is whether anyone ever ran two checks: do two managers score the same call the same way, and does a high score have anything to do with closing.
Almost nobody runs either one. This post is about how, and what you do with the answers. The template is the easy part.
What is a sales call scorecard actually for?
A sales call scorecard is a rubric that turns a recorded call into numbers, so coaching stops being one manager's memory of a call they half listened to while eating lunch.
It has two jobs. The first is telling an individual rep what to change on Monday. The second is telling the floor which behaviors actually move revenue, so you know what to train.
Most templates are built for the first job and quietly fail the second. That failure is invisible, because a scorecard produces confident looking numbers either way.
Worth saying plainly: a scorecard is a measuring instrument. Anywhere else that people build measuring instruments, they test them before trusting the output. Sales floors skip that step and go straight to the spreadsheet.
Why are the weights on most sales call scorecard templates invented?
Read through the scorecard templates on page one of any search and you will find weightings. Opening 15 percent, value 25, objection handling 30. They read as reasonable. None of them cite anything.
A weight is a claim. Saying objection handling is worth 30 percent and the opening is worth 15 says that a point of objection handling is worth twice as much revenue as a point of opening. That is a testable claim about your business, and it is almost never tested.
Here is the cheap version of the test. Pull 40 calls that closed and 40 that did not, from the last two quarters. Have someone score them without knowing the outcome. Then compare the category averages between the two groups.
The blind part matters. A reviewer who already knows the deal closed will score the call higher, and you will have measured your reviewer instead of the call.
What usually comes back is uneven. One or two categories separate won calls from lost calls by a wide margin. Two or three barely move. A category that does not separate at all is not earning its place on the form, and it certainly should not carry a weight.
Two cautions before you rewrite anything based on one pass:
- Separation is not proof of cause. Lead source alone can explain a lot of it. Calls from inbound demo requests go better and score better, and the rep may have had nothing to do with either.
- Eighty calls is a small sample. Treat the first result as a hypothesis, then repeat it next quarter. A pattern that shows up twice is worth acting on.
Would two of your managers score the same call the same way?
This is the check that decides whether any of the numbers mean anything, and it takes an afternoon.
Rating instruments are judged on whether independent raters agree. The standard measure is a chance corrected agreement score called kappa, and the thresholds most fields quote come from a 1977 Biometrics paper by Landis and Koch: under 0.20 is slight agreement, 0.21 to 0.40 is fair, 0.41 to 0.60 is moderate, and 0.61 to 0.80 is substantial. You can read the original paper on observer agreement if you want the maths.
You do not need the statistics to start. Run the drill.
- Pick 10 recorded calls of the same type.
- Two managers score all 10 independently, with no conversation between them.
- Put the scores side by side, category by category.
The first time a team does this, the usual finding is gaps of three or more points on a 1 to 10 scale, on the same call, in the same category. Sometimes the two managers disagree about whether a behavior happened at all.
The 1 to 10 scale is a big part of it. No human being can reliably tell a 6 from a 7 on rapport. Give people ten options and you get ten different readings of the middle.
The fix is to shrink the scale and sharpen the question. Compare these two items:
- "Rate rapport, 1 to 10."
- "Did the rep confirm a specific next step with a date and a time? Yes or no."
Two managers watching the same recording will agree on the second one nearly every time. They will never agree on the first. The second item is also the one a rep can do something about tomorrow.
Agreement drifts, so rerun the drill every quarter, and any time you change the form or add a reviewer.
How many calls before a rep's score means anything?
More than you are grading.
The common advice is three calls per rep per week. Over a month that is 12 calls, spread across call types with completely different shapes, and one rough call moves a 12 call average a long way.
Reps work this out fast. Once a team decides the score is noisy, the review becomes theatre, and the scorecard stops changing behavior no matter how well designed it is.
Three rules that keep the numbers honest:
- Set a minimum sample before a score enters a performance conversation. Pick the number in advance and hold to it. Coaching on a single call is fine, as long as you call it coaching on a single call.
- Keep call types in separate buckets. A cold outbound dial and a booked discovery call have nothing in common structurally. Averaging them produces a number about your call mix.
- Report the trend across a window, not this week's figure. Four week rolling averages per category are boring and much harder to argue with.
If you cannot reach a decent sample by hand, that is the real argument for scoring calls automatically. The advantage is coverage, not insight.
Which calls can you even score on an outbound desk?
The template posts skip this, because most are written for teams whose calls are all booked meetings. An outbound desk is a different animal.
On a cold list, the large majority of dials never reach a person at all. Voicemail, no answer, a number that moved two years ago. Our write up on cold call connect rate goes through how little of a dialing day becomes a conversation.
So the scoreable population is a small and biased slice of the day: the people who picked up and stayed on the line. Scores from that slice tell you about conversations. They tell you nothing about the rest of the hours, and a rep with a thin week of conversations is often looking at a list problem rather than a skill problem.
There is a second layer on a predictive campaign, where software decides things before the rep ever speaks. Answering machine detection has to judge whether a live human picked up, it works against a hard decision deadline, and it defaults to calling the line a machine when it cannot decide in time. A call classified that way never reaches a rep and never gets scored. When a rep's conversation count looks wrong, check the campaign before you check the rep.
Keep those two things apart on purpose. Behavior scores describe conversations. Connect rate, contact rate, and list quality are operational numbers, and mixing them into one composite score hides both.
One more practical limit: you cannot score a call you did not record, and recording rules differ by state and by who is on the line. We covered the shape of that in our piece on sales call recording laws. Treat it as background for a conversation with your own counsel, not as legal advice.
What goes on a sales call scorecard that holds up?
Seven build rules, in the order they matter:
- Every item is observable in the recording. If two people can watch the same 30 seconds and disagree about whether it happened, it is a trait, and traits belong in a development plan rather than on a scorecard.
- Three points, not ten. Did it, partly did it, did not do it. Agreement climbs immediately.
- Six or seven items at most. A 25 item form gets filled in from memory by item nine.
- One scorecard per call type. Cold dial, discovery, follow up, renewal. Separate forms, separate averages.
- Weight only what you have checked. Until the closed won versus closed lost pass tells you something, weight everything equally and say out loud that the weights are provisional.
- Write the anchors, not the labels. Each level gets a sentence describing what it looks like on a call. "Discovery: 2" means nothing a month later.
- Every reviewed call ends with one change and a timestamp. "At 3:40 the prospect raised price and you moved on" beats a paragraph of category commentary.
A workable starting set for a cold outbound call, all binary or three point:
- Gave a specific reason for the call inside the first 20 seconds.
- Asked at least one open question before describing the product.
- Took the first brush off without arguing. Our post on cold call objections covers the one you never argue with.
- Confirmed a next step with a date and a time, or ended the call cleanly.
- Logged an accurate outcome. A scorecard sitting on top of sloppy dispositions is measuring fiction, which is why we wrote about building a call disposition list you can trust.
Notice what is missing from that list: talk ratio.
The 43 to 57 split that shows up on nearly every scorecard template traces back to Gong's research labs, which analysed roughly 25,000 recorded business to business calls on its own platform. That is a real finding about that population, which skewed heavily toward longer discovery and demo conversations. Dropping the same number onto a 40 second cold call applies a benchmark to a call shape it was never measured on. If you want talk ratio on the form, measure your own distribution first and check it against your own outcomes.
Does AI call scoring fix any of this?
It fixes one thing properly, and it is the biggest one. Coverage. Scoring every call instead of three removes the sample size problem, which is what makes most manual programs collapse by month three. Automatic scoring also gets feedback to a rep the same afternoon instead of nine days later, when nobody remembers the call.
What it does not fix is the rubric. A model scoring against your categories inherits your categories and your invented weights, and it will produce them faster and more consistently than your managers did. Consistent and wrong is still wrong.
So point the agreement drill at the software too. Score 20 calls both ways, with a calibrated human reviewer and the automatic score, and compare them category by category before that number touches anyone's pay or a performance plan. If the two disagree, the useful question is which one you can defend to the rep.
Where automated scoring earns its place is speed and reach. Live guidance goes further again, putting the next question on screen during the call rather than after it, which is the part of AI sales coaching that changes the next 30 seconds instead of the next call.
What is the smallest version of this that works?
Four steps, roughly a month of light work:
- This week. Ten recorded calls, two managers, independent scores, side by side comparison. Look at how often they land within one point.
- Next week. Rewrite the form down to six observable items on a three point scale, with written anchors. Keep one form per call type.
- The week after. Score 40 closed and 40 lost calls blind. See which items separate them. Set weights from that, and write down the date you did it.
- Every quarter. Rerun the agreement drill and the separation pass. Both drift.
A team that does this ends up with a shorter sales call scorecard, fewer arguments about scores, and a defensible answer when a rep asks why a 6 was not a 7. That last one is worth more than any template.
SellifyGPT keeps the dialer, the CRM, the recordings, and the coaching on one record, so a call, its disposition, and its score are not scattered across three tools that disagree about what happened. If you want to try the drills above on your own calls this week, you can start a free trial and cancel before it ends.
See it on your own calls.
SellifyGPT puts the dialer, CRM, and an AI coach in one place. 14-day free trial, cancel before it ends.
Start free