Viewing Performance Numbers: Why NHS League Tables Isn’t Working
We often use numbers to make sense of performance. In the NHS, this frequently means comparing organisations: waiting times, cancer outcomes, emergency department performance, productivity, patient experience and so on. The logic is straightforward: measure performance, rank organisations, identify the best and worst, and learn from the difference.
But what if the difference between organisations isn’t simply a difference in how well they are doing the same thing? What if they are operating as different systems? Imagine we measure the performance of 100 organisations and plot the results. We might expect a normal, bell-shaped distribution: most organisations somewhere around the middle, with a few performing particularly well and a few performing poorly. The natural conclusion is: Some organisations are better than others.
But what if the data don’t form one smooth curve? What if there are two distinct peaks?
That might suggest we aren’t looking at one population of organisations with varying levels of performance. We might be looking at two different types of organisations, behaviours or operating models.
A classic example comes from experiments involving rats learning to navigate mazes.
Researchers observed that some rats appeared to solve the maze differently. One group seemed to learn individual routes and movements. Another appeared to develop a more abstract understanding of the maze itself. It would be easy to describe the second group simply as having “higher ability.” but that may miss the more interesting phenomenon.
Perhaps these rats aren’t simply better at the same thing. Perhaps they have developed a different way of solving the problem. One is a difference in degree: “This rat is better than that rat.” The other is a difference in kind: “These rats may be solving the problem differently.”
Now consider the NHS league table
Suppose one NHS organisation has a significantly better waiting-time performance than another.
The league table tells us who is ahead but it doesn’t necessarily tell us why. Perhaps Organisation A has genuinely found a better way of working. Excellent. We should learn from it. But perhaps A and B are dealing with fundamentally different circumstances: different populations, workforce markets, levels of deprivation, patient flows, geography, relationships with primary care, community capacity, demand patterns or historical configurations. Or perhaps they have developed fundamentally different ways of organising care. A ranking compresses all of that complexity into a number. We then risk making the classic mistake: Turning a difference in kind into a difference in degree.
The more systemic question is: What is producing these different patterns of performance?
The danger of the league table
League tables aren’t necessarily useless. They can reveal variation, challenge complacency and make poor performance visible. The problem is when we mistake measurement for explanation. A league table cannot, by itself, tell us whether that variation is caused by leadership, resources, behaviour, context, relationships, incentives, history, demand or a fundamentally different way of organising the system.
These nuances are useful because the intervention for failing organisations changes depending on the explanation. If Organisation B is simply less efficient, perhaps it needs to adopt Organisation A’s practice, but if Organisation B is operating within a different system, copying Organisation A may achieve very little.
Alongside asking the question of who is performing best, we need to start asking more broadly: “Why are these patterns of performance emerging, and what different systems, relationships and ways of working might be producing them?”
That shifts us from ranking systems to understanding systems.
