You are being asked to cut agency spend. Here is what followed in the trusts that did.
Across three financial years to FY 2022-23, every £100 of agency reduction was followed by about £2.70 of pay-bill-equivalent sickness absence in the six months after each year closed. That is a valuation of lost time, not cash out of the door, and not a forecast for your organisation. It is what comparable trusts recorded the last time this decision was made at scale, and it is very hard to put in a board paper without an independent source.
NO PREPARATION NEEDED · NO DATA REQUIRED FROM YOU
NHS provider agency spend, 2024/25, down from £3.46bn in 2022/23 (NHS England, Financial performance update, 2025)
The largest share of trust-years for which any single failure mechanism is the dominant one, on organisations held out of the analysis until after it was finished. 21.6% on the panel it was developed from. Counted one row per organisation instead, a later check on the unit, the largest share is 25.6%. There is no typical case.
Median gap between staff who say they intend to leave and staff who actually go the following year (N=203 NHS trusts)
There is no typical failing trust.
No single workforce failure mechanism is the dominant one for more than 26% of organisations, however we count it: 21.6% of trust-years on the panel we developed the measure from, 24.6% on a separate group of trusts that was set aside before the work began and not looked at until after it was finished, and 25.6% counted one row per organisation rather than one per trust-year. That last count was a later check on the unit, not one of the gates we fixed in advance. The three pressures we measure are close to independent of each other: mean absolute correlation 0.114, strongest pair 0.217. A trust can be severe on one and unremarkable on the other two.
Which is why a national solution designed for the average lands so differently in each organisation, and why the sector figure above tells you almost nothing about your own position. The only way to know is to be measured.
On a £100 million annual pay bill, a two point reduction corresponds to about £54,000 of pay-bill-equivalent absence over the following six months. That is a valuation of lost time, not cash leaving the organisation, and a historical observation rather than a forecast for your organisation or a demonstrated cause.
See the full cost table, by sector, with confidence labels →
Three questions most workforce reports cannot answer
Not because the reporting is poor. Because the answers need a comparison across organisations that no single trust can build from its own data.
Which pressure is actually the binding one?
Pressure Displacement, Departure Lag and Normalised Fragility are close to independent of each other in our panel: mean absolute correlation 0.114, with the strongest pair at 0.217. A trust can be severe on one and unremarkable on the other two. Treating them as one problem funds the wrong one.
Will the fix hold?
Trusts that get out of the heaviest agency quartile are still out two years later 77.8% of the time, against 91.1% for trusts that were never in it. Most escapes hold. Relapse is 13.3 percentage points more likely than in trusts that were never there, which is the number that belongs in a plan.
What does the board need to stop doing?
Every consultancy arrives with things to start. Four things worth stopping came out of our own testing, and they are immediately below.
Why organisations engage us when they already have workforce dashboards
Your workforce information team is almost certainly better at your data than we are. They have the establishment, the rosters, the exit interviews and ten years of context we will never see. This is not a comparison of skill. There are four things that are unavailable from inside any single organisation, however good the team is.
Peer data is not the scarce thing. The cross-silo read is.
Peer comparison is not the scarce thing. NHS England's Model Health System gives every trust benchmarked productivity and quality data against its peer group, free, and the NHS Benchmarking Network runs a member programme on top of that. An internal team can already tell you whether your sickness rate is unusual for your type of organisation. What none of them does is read agency share, the staff survey and sickness together as one structural position, because the three sit in three separate published sources that no standard dashboard joins. Our panel does exactly that, across between 178 and 244 NHS organisations depending on the measure, over 2021 to 2024.
The measures that matter most sit between directorates.
The three pressures we measure are each assembled from data that lives in different places. Agency share sits in finance. Intention to leave sits in the staff survey. Sickness sits in the workforce return. Nobody owns the space between them, so no dashboard reports it, and the gap between what staff say they will do and what they actually do is invisible to every system that holds only one half of it.
Internal findings are contestable in a way external findings are not.
This is the uncomfortable one. When the workforce team reports that a saving target is undeliverable, that finding arrives from the function with the most to lose from the target. It may be completely correct and it will still be discounted. An independent source with no NHS contracts and nothing to sell that depends on the answer changes how the same evidence is received in the room. We are not better analysts than your team. We are differently positioned in your politics.
Somebody has already done the work of finding out what does not work.
Doing this properly means testing whether your explanations survive contact with data. We ran seventy-nine pre-specified tests. Fifty failed. Geographic isolation does not explain agency dependence. Local labour market competition does not explain it. Housing cost turned out to be a proxy for region and we retracted it the same day we found it. Your team could establish all of that. It would take them the better part of a year, and the deliverable would be a list of things that are not true.
The honest summary: your team knows your organisation. We know what your organisation looks like from outside it, and which of the things everyone assumes are true actually survived testing. Every report ends with a hypothesis your team can test against internal data we were never given. That is the intended use of it, not a limitation.
Four things the evidence says you should stop doing
Every consultancy arrives with a list of things to start. These come from our own testing, not from convention, and each one is a thing to stop.
Across FY 2020-21 to FY 2022-23, lower agency spend in a financial year and higher permanent-staff sickness absence in the six months after it appeared together. We have not established that one caused the other, which is exactly why the question is worth answering before the decision rather than after it.
By the time turnover moves, the decision to leave was taken months earlier.
Distress that lasts long enough stops triggering the alarm.
If the baseline is already broken, nothing will ever look acutely worse.
Each of these came out of the same seventy-nine-test battery, in which fifty hypotheses did not survive. That is why they are short.
We publish what failed
In August 2026 we ran seventy-nine pre-specified tests across every construct we use. Twelve passed. Fifty failed. Two were underpowered and fifteen could not be scored on public data. The significance threshold was set at p < 0.001 in advance, deliberately stricter than convention, because a battery this size would produce about four passes by chance alone at the usual threshold.
One of our four original measures failed its gates: Capability Collapse by Cluster. It is named, with its numbers, on the evidence page.
pre-specified tests, every result recorded and available in full on request
hypotheses that failed, and were kept in the record
NHS trusts in the validated panel, 2021 to 2024
One starting point. Follow-on only where the evidence justifies it.
The Workforce Structural Review is the immediate entry product. It ends with one recorded outcome: Stop, Investigate, or Redesign candidate. Redesign is separately scoped only where the evidence and entry conditions justify it. Diagonal comes later where repeated measurement has a clear use.
01Workforce Structural Review
A board-ready review of three independent workforce measures, peer comparison, watch-list history, financial framing, and a live decision session.
Workforce Redesign Programme
Conditional follow-on where the Structural Review produces a Redesign-candidate outcome and the entry conditions are met.
Diagonal Monitoring
The later monitoring layer where a baseline exists and repeated measurement has a clear use. Client-platform access is not currently available.
A full review of a large acute trust, using only 2023 published data.
No engagement or commission from the trust, and no internal data request of any kind. Their Departure Lag rose while their Normalised Fragility flags fell: two independent measures of the same workforce, moving in opposite directions in the same year. No single dashboard shows both, because each lives in a different silo. We report the three axes separately for exactly this reason, and this trust is why.
Built without a single internal data request
Peer benchmark behind every finding, built the same way across every organisation in it
A testable hypothesis the trust can check against its own data inside a month
From first conversation to a decision you can defend
| Step | What happens | Timing |
|---|---|---|
| Discovery call | You describe the workforce decision you are least confident about. We say whether published data can speak to it. Free, no material required from you, and as long as the question needs. | Week 0 |
| Structural review | We build your three-axis position from published data and benchmark it against up to 203 peers. | Ten working days from agreement to handover |
| Board session | We present what is elevated, what it means for the decision in front of you, and what the evidence says not to do. | Handover and decision session |
| Redesign, if you want it | If the review says the structure has to change, we design and sequence the intervention. | Month 2 onward |
| Monitoring, if you want it | Where repeated measurement is useful, Diagonal is the later monitoring layer. Client-platform access is not currently available. | Ongoing |
Independent, and structurally so
Independent
No NHS contracts, no supplier relationships, no agency partnerships, and no third-party product we earn commission on. Nothing in our commercial model depends on your conclusion.
Public data
Every finding is built from NHS Staff Survey, NHS England Digital workforce statistics, trust annual accounts and NHS Organisation Data Service records. You can reproduce our sources; we will tell you exactly where each figure came from.
Pre-specified
Gates are written down before the data is acquired. When a result fails its gate, it is published as a failure. One of our four original measures, Capability Collapse by Cluster, is on this site as a documented failure.
Frontline
The founder holds CIPD Level 5 and an MSc in Leadership and Human Resource Management, and works as a Healthcare Assistant in a UK healthcare setting. The measures were built by someone who has worked the rota they describe.
What is the workforce decision you are least confident about?
That is the whole agenda for the first call. As long or as short as the question needs, no preparation, no data from you. If published NHS data cannot speak usefully to your question, we will tell you on the call and there is no second conversation to sit through.