The short answer: do not look for a call to coach. Look for a skill that is costing money across a lot of calls, then open the two calls that show it best. That inverts the cost. Finding the pattern is something software can do while nobody is watching; deciding what to say about it is the part that needs a person, and it takes about ten minutes once you know what you are looking at.

The gap is an attention problem

Ask a sales manager whether coaching matters and they will say it is the most valuable thing they do. Ask them when they last did it properly and the answer is usually a specific, slightly embarrassed date. Nothing about that is a failure of intent. It is what happens when the work costs an hour and the calendar has fifteen minutes.

The hour goes on the search, not the coaching. Scrubbing a recording to find the moment a discovery call went shallow, or the point where a rep answered a price objection by discounting, is genuinely slow, and it is slow every single time. So the ritual survives (a weekly one to one, a pipeline walk, some encouragement) and the specific, uncomfortable, useful conversation does not happen.

What "coaching" usually means in practice. A manager listens to a call they happened to join, forms an impression, and gives feedback about that call. It is not wrong, it is just a sample of one, chosen by accident, about a rep who may be strong at exactly the thing that call tested.

Start from the pattern, not the call

A rep is not uniformly good or bad. They are usually strong at two things and weak at one, and the weak one shows up in the same place every time. Scored consistently across every call a rep makes, that becomes visible in about a week, and it points at what to talk about before anybody opens a recording.

The EmpireOS team skill heat map: every rep as a row and discovery, objection handling, closing, pitch and multi-threading as columns, each cell a score out of one hundred, with a team average row underneath and the lowest team skill named below it.
Every rep against every skill, with a team average underneath. The value is not the individual cells; it is the columns. A column where the whole team is weak is a training problem, not six coaching problems.

Two things fall out of a view like that, and they need different responses.

A weak column is a team problem. If closing is the lowest score across six reps, no amount of individual coaching fixes it efficiently. That is a play that needs rewriting, or a role-play session, or a piece of collateral that does not exist yet.

A weak cell is a person problem. One rep well below the team on objection handling is a specific conversation with a specific person, and now you know what it is about before you book it.

What a graded call actually gives a manager

Once the pattern has told you where to look, the recording stops being an hour of audio and becomes a page. What is worth having on that page is narrow: the summary, the objections that were raised and how they were handled, the next steps that were agreed, and the transcript underneath so you can read the exact words rather than a paraphrase of them.

A recorded call opened in EmpireOS: an AI summary with the key topics, the objections raised, the sentiment and the agreed next steps, with the full timestamped transcript beside it.
The same call as a page: what it was about, what was objected to, what was agreed, and the transcript beside it. Reading this takes about two minutes, which is the difference between coaching happening and not.

The transcript matters more than the summary. A summary tells you a price objection came up; the transcript tells you the rep said "I can probably do something on that" nine seconds later, which is the actual coaching moment and is not the kind of thing a summary preserves.

Coaching is not telling someone they are weak at discovery. It is showing them the forty seconds where the call turned, and asking what they would do differently.

Tie it to the deal, or it stays an opinion

A skill score on its own is a performance review. It becomes coaching when it is connected to money. If deals where a second stakeholder is never engaged close at a materially lower rate than the ones where they are, and one rep is bottom of the team on multi-threading, that is not a soft skill note, it is a revenue conversation with a number attached.

That connection is what deal intelligence is for: reading a team's own closed deals and reporting which behaviors separate the ones that closed from the ones that did not, with the difference in close rate and the number of deals it was measured across. We wrote about one of those patterns in detail in multi-threading, the deal-killer nobody measures.

The one thing to do this week

Pick the lowest column, not the lowest cell. Team-wide weakness is cheaper to fix and pays six times. Book thirty minutes, open two calls that show it, and spend the time on what to say differently rather than on establishing that there is a problem. Then look again in three weeks and see whether the column moved, which is the only evidence any of this is working.

What this does not do

A score is a starting point, not a verdict. It is computed from what happened on calls, so it inherits whatever is uneven about which calls got recorded and which deals a rep was given. Treat it as the thing that tells you where to look, and treat the recording as the thing that tells you what is true. Any system that asks you to act on the score without showing you the evidence underneath is asking for more trust than it has earned; that is the argument we make about standalone call intelligence generally.

Questions people ask about this

Do reps have to know their calls are being scored?

Yes, and they should be able to see the scores themselves. Coaching that depends on a rep not knowing how the judgment was made is not coaching. It also fails practically: a rep who can see the factor that is dragging their score down can work on it between conversations, which is where most improvement actually happens.

What if the scoring is wrong about a call?

Then the manager overrides it, on specific grounds, having read the transcript. Every score in EmpireOS shows the factors that produced it precisely so it can be argued with. A number a rep cannot disagree with is a number they will learn to ignore.

How many calls do you need before a pattern means anything?

Enough that one bad Tuesday does not move it. In practice a couple of weeks of a rep's normal call volume is where a column starts to be worth acting on. Before that, look at it and do not schedule anything on it. This is also why the team average matters: it gives you something to compare a cell against other than your own memory.

Can this replace one to ones?

No, and it is not meant to. What it replaces is the hour of searching before the one to one, and the vagueness inside it. The conversation itself is the part that only a person can have.