Measurement That Earns Its Keep
Most service management reporting has no shortage of numbers. What it lacks is a number anybody decided, in advance, to act upon — which is why nobody can prove the improvement worked once the counting rules have quietly moved.
Series: The Improvement Engine, Post #2
Measurement That Earns Its Keep
I sat in a service review where a team showed incident volumes falling since their improvement programme began. I asked what the volume had been the quarter before, and how it had been counted then. Nobody could tell me. The counting rule had changed partway through and the old figures had gone with it. The improvement may have been real, but it was unprovable, which in front of a sceptical finance director amounts to the same thing.
Improvement without a baseline is opinion with a project plan. The difficulty is rarely a shortage of data — most organisations are drowning in it — but a shortage of measures whose purpose somebody settled before collecting them.
Nobody Decided What the Number Was For
The dashboard has no hierarchy because the choosing step never happened
Ask how the measures on a service report were chosen and the honest answer, in most organisations, is that the tool offered them. That is not a hierarchy of measurement but an inventory of what was convenient to extract.
The older vocabulary was better than the practice that replaced it. A critical success factor is a condition that must hold for an objective to be met. A KPI is a metric chosen to manage the thing — in my usage, the small subset you agreed in advance would change a decision. A metric is any measurement you can take. ITIL 4 keeps the ITIL v3 cascade under new names — practice success factors with key metrics — and gives measurement its own general management practice besides: measurement and reporting. The usual explanation for why packs only ever grow — filing costs nothing — is wrong. A crowded pack is protective. With forty numbers on it, one always supports the conclusion the presenter walked in with, so the pack can never be lost, only re-narrated. Past a small set, measurement stops being prudence and starts buying immunity from contradiction.
- Start from the decision, not the data — name who acts, on what, at what threshold, then find a measure that serves them.
- The ceiling is low for a reason — a number nobody claims is still defended when it moves, so each measure past the deciding few buys an argument rather than an answer.
- Every number wants an evaluation point — a date on which a named person looks at it and records a conclusion.
My preference is the quickest audit available: point at a number and ask what would be done differently if it moved sharply either way. If the answer is nothing, it is not a KPI. It is furniture, and furniture is what an unlosable pack is made of.
The Baseline You Did Not Take
The one question routinely answered with an adjective
ITIL 4’s continual improvement model runs to seven steps, six of them questions, and the second — where are we now? — is the one answered with an adjective. By the time an improvement is funded everybody agrees the current state is bad, so stopping to quantify it feels like delay dressed as diligence.
A baseline is not a reading but a period with its counting rules attached. Definitions drift in every tool: whether resolved means the fault was fixed or the service restored, whether the clock pauses awaiting a customer, how a reopened ticket is counted. A month is the minimum and a quarter is better, and it must contain one full instance of any cyclical driver, or your baseline gets beaten by month-end rather than by your improvement.
- Record the counting rule beside the figure — an undocumented rule guarantees a later argument about whether the comparison was ever fair.
- Know the ordinary range before claiming a shift — statistical process control has had the vocabulary since Shewhart, and real improvement moves outside the band the measure normally wanders in.
- Size the comparison window to the driver, not the reporting date — long enough to contain a full turn of whatever normally moves the figure, which is rarely the same length as the month somebody wants to report in.
The baseline is the cheapest part of the exercise and the only part that cannot be recovered afterwards. That leads somewhere unpopular: an improvement whose before-state was never measured should not start. There is one honest exception, and impatience is not it — where the item carries a regulatory finding or a safety exposure, waiting a cycle to measure is itself the negligence. Everywhere else, carve the exception rather than soften the rule: act now, and reconstruct the baseline from history in parallel, before the action lands rather than after.
Numbers That Decide and Numbers That Decorate
Some numbers move when the service improves, others when somebody adjusts the clock
Ticket volume is the classic offender. It describes demand and is useful for staffing, but volume falls when failures reduce and just as neatly when users give up and ask the colleague two desks over. Recurrence carries the signal instead: it moves when a cause has gone and sits still when only the symptom was cleared.
The workable definition is incidents in the period linked to an existing problem or known error record, over the total. That depends on linkage being maintained, which in most estates is patchy. ISO/IEC 20000-1:2018 clause 9.1 requires an organisation to determine what it monitors and measures, the methods, and when results are analysed and evaluated; ISO 9001:2015 asks the same at 9.1.1. A pack full of figures with no stated evaluation point is not merely unhelpful. It is evidence that the determination those clauses ask for was never made.
- Never print a measure without its second number — a figure survives contact with self-interest only when something beside it shows how it was obtained; any single number can be produced, a pair is much harder to arrange.
- A mean conceals its own tail — publish the median beside the 95th percentile, because the average of a password reset and a week-long escalation describes no real ticket.
- Publish the linkage coverage beside the recurrence rate — without it you are reporting the state of your problem records, not the state of the service.
Mean time to restore does not belong on the front page of an improvement report, and I would argue that fairly hard. It cannot tell a service that got better from a portfolio that got tidier: retire a badly behaved queue or adjust a clock rule and it improves without a single user noticing. Print the count of queues and clock rules changed beside it, or leave it out.
What the Failure Actually Costs
Sizing a recurring fault from records you already hold
Evidence has a second half: what the failure costs, and so whether it deserves capacity somebody else was promised. Almost no improvement case attempts it, which is why so many are argued on indignation and lost to whoever is more senior.
Three figures you already hold will do it. Take the frequency of the failure across the baseline period, not from memory. Take the handling effort per occurrence from your own resolution data, remembering that elapsed time is not effort. State the consequence in the audience’s own unit: orders delayed, claims reworked, appointments rebooked. Do not multiply by an invented hourly rate; that one unsupported step invalidates the two honest numbers underneath.
- The sizing is biased towards failures you can already see — every figure comes from your own tooling, so the workaround an operations team built years ago and never logged sizes at nothing; say so, and go and ask once.
- Use the audience’s own unit — a finance director will argue with your cost model all afternoon and will not argue with the number of claims reworked.
- Show the arithmetic — a sized problem that survives a sceptic recomputing it outranks a larger one nobody checks.
The failures that get funded are rarely the most serious ones. They are the ones whose seriousness was expressed in a unit the budget holder already thinks in — including, uncomfortably, when the sizing was incomplete.
When the Number Will Not Arrive in Time
The most honest line in a report is sometimes the absence of a figure
The measures that prove an improvement worked lag by construction: recurrence across a quarter, an audit finding that survives the next visit. What fills the gap is not a leading indicator in the proper sense; that term belongs to measures with a demonstrated predictive link. It is input and compliance measurement: time from evidence to agreed action, the share of actions carrying a named owner and a date. A register with perfect owner-and-date compliance and a ninety-day time to first activity tells you exactly where the mechanism is broken.
Some effects will not produce a number at all on the timescale of the report. Reduce the likelihood of a rare and expensive failure and success is indistinguishable from good fortune for a long while. The temptation is to produce a figure anyway — a modelled saving, an estimated avoidance. Resist it: pull that thread and every other number in the pack becomes suspect, including the honest ones.
- Input measures steer, outcome measures prove — neither should appear in the costume of the other.
- Small numbers do not make percentages — three incidents falling to one is ordinary variation, and calling it a two-thirds reduction will not survive the next month.
- Say “not yet measurable” out loud — record the intended effect, why it cannot be read yet, and the date you will go back and look.
Publish the steering measures with their caveat beside them: a function reporting no throughput is the easiest line in the pack to cut. Then refuse the request that follows. Owner-and-date compliance is the easiest figure in the pack to satisfy without anything actually moving, which is exactly why somebody will propose putting it in an objective — and improvement measures are defined, collected and reported by the same people they judge, the one condition under which a target stops being a reading and becomes a negotiation. Senior people spend their working lives being sold to, and they remember when somebody declines to.
Where I Would Start
Begin by taking numbers away, not adding them
If your improvement reporting is a pile of numbers nobody acts on, the remedy begins with subtraction rather than addition.
- Take the baseline before the next action — one full reporting cycle with the counting rules recorded beside it, reconstructed from history if you have already started.
- Ask each number who acts on it — where no owner can be found, that silence is the answer and the number goes.
- Add recurrence rate with its linkage coverage beside it — and where that coverage proves poor, fixing the linkage discipline is your first corrective action.
- Publish the median and the 95th percentile in place of the mean — the reaction in the room will tell you how much the average had been hiding.
- Size one recurring failure properly — frequency, effort and consequence in the audience’s own units, on one page, with no invented pound figure on it.
None of this needs a measurement framework or a maturity model. It needs one decision taken before anything is collected — what each number is for and who acts on it — then the discipline to discard whatever fails that test.
Every failure named in this post is one failure. The counting rule that moved partway through the programme, the baseline rule nobody wrote down, the restore time that improved because a queue was retired, the hourly rate invented once a case was wanted, the two-thirds reduction built from three incidents, the modelled saving for an effect that has not arrived — in each, the rule of the contest was settled after somebody already knew the answer. Measurement is not the difficult part; almost anyone can count. Fixing the terms while nobody yet knows who wins is, because it is the only moment they can be set for a reason other than the result they produce. A measure earns its keep at the moment it is defined, not the moment it is read.
Sizing a failure honestly also creates the difficulty that follows it. A dozen problems, each properly evidenced, still arrive at one team holding one fortnight, and something has to decide which of them gets it.
Hopefully this has been useful to you and I wish you well on your ITSM journey…