BLOG
Risk Matrices and Heat Maps Useful, Dangerous, or Both
Published 6 August 2026 ยท By P Larner
Risk Management#FAIR#risk scoring#quantification#risk matrices#heat maps
Risk Matrices and Heat Maps: Useful, Dangerous, or Both?
Illustration for: Risk Matrices and Heat Maps: Useful, Dangerous, or Both?
The 5x5 risk matrix is the most widely used decision tool in security, and one of the most criticised. Both camps are right. This article sets out what matrices genuinely do well, the specific ways they mislead, and the rules that keep them honest.
Why the 5x5 dominates
The matrix won for practical reasons, not analytical ones. It needs no data beyond expert judgement, so you can populate it in a workshop before lunch. It fits on one slide. Every auditor, regulator and framework recognises it, because ISO 27005 describes consequence and likelihood scales, NIST guidance uses them, and most GRC tools ship a 5x5 by default. And it gives a committee a shared vocabulary. โThatโs a redโ is a sentence everyone in the room understands, whether or not they should.
There is also an honest version of the argument for it. Most organisations do not have the loss data, the modelling skill or the time to quantify every risk. A matrix is a structured way of writing down judgement, and structured judgement beats unstructured judgement.
Annotated 5x5 risk matrix with green, amber and red cells, three example risks plotted, and a speech bubble reading thatโs a red
A clean 5x5 matrix, which is a shared vocabulary for relative priority rather than a measurement instrument.
The genuine value is communication, not calculation
Used properly, a matrix is a communication device. It forces a conversation about two questions, how likely and how bad, that would otherwise stay vague. It makes relative priority visible, so the board can see that supplier compromise sits above laptop theft without reading forty pages. It gives a consistent shape to risk reporting across departments that otherwise describe risk in incompatible ways. And movement on the matrix over time, as in โthis was red in January, it is amber now, here is whyโ, is a legitimate and useful story to tell.
The trouble starts when the communication device is mistaken for a measurement instrument.
Known failure modes
These are not theoretical. Tony Coxโs 2008 paper โWhatโs Wrong with Risk Matrices?โ formalised several, and anyone who has sat through enough risk committees has watched the rest.
Range compression
A โlikelyโ bucket covering everything from 20 per cent to 90 per cent annual probability treats a 1-in-5 event and a near-certainty as identical. Two risks with expected losses differing by an order of magnitude can land in the same cell, and a smaller risk can outrank a larger one purely because of where the bucket boundaries fall.
Ordinal arithmetic
Multiplying a likelihood score of 4 by an impact score of 3 to get โ12โ is not mathematics, it is numerology. The labels are ordinal, so they encode order, not distance. A likelihood of 4 is not twice a likelihood of 2, so the products are not comparable, yet risk registers routinely rank on them and even sum them into โtotal risk exposureโ figures. Those numbers mean nothing.
Colour-driven decisions
The colour scheme is a design choice, but it drives real spending. Shift a threshold and yesterdayโs amber becomes todayโs red, and budget follows the paint rather than the exposure. I have watched a committee spend 40 minutes on a red rated on a worst-case reading and wave through an amber that was quietly a bigger expected loss.
Everything lands in amber
Assessors avoid the extremes. Scoring something 1 invites the question โwhy is it on the registerโ, and scoring it 25 invites the question โwhy has it not been fixedโ, so scores herd into the middle. A register where 70 per cent of entries sit in three central cells is ranking almost nothing.
Worst case or typical case, chosen inconsistently
One assessor scores impact on the plausible outcome, another on the catastrophic tail. Both are defensible, and mixing them in one register is not.
Risk matrix with most dots crowded into three central amber cells and a magnified callout showing two very different risks sharing one cell
Amber herding in practice, where 14 of 20 risks share three cells and two same-cell risks differ by an order of magnitude in expected loss.
Rules that keep a matrix honest
You do not have to abandon the matrix to fix most of this. You have to be strict about how it is used.
- Put numbers behind every label, and publish them. โLikelyโ means 25 to 50 per cent in the next 12 months. โMajorโ means ยฃ1m to ยฃ5m of loss, or a defined operational or safety outcome. If assessors cannot see the ranges, each is scoring against a private scale.
- Fix the time horizon. Likelihood over what period? A risk that is unlikely this year is near-certain over a decade. Pick one horizon, as 12 months is the usual choice, state it on the matrix, and reject scores that quietly use another.
- Decide the impact convention. Score the most plausible bad outcome, not the theoretical worst, and if the tail matters, record it as a separate note. Write the convention down.
- Never do arithmetic on the scores. Use the cell for triage and ordering conversations. Do not sum, average or trend the products. If someone wants โtotal exposureโ, that is a request for quantification, not a bigger spreadsheet formula.
- Score inherent and residual with the control named. โAmber because of the EDR rolloutโ is checkable. โAmberโ is a mood.
- Calibrate assessors occasionally. Score five sample scenarios independently and compare. The spread is usually humbling, and half an hour of argument about it improves every subsequent assessment.
- Interrogate the amber middle. Anything sitting in the central cells for more than two review cycles should be forced to a decision, which means treat it, accept it formally, or re-score it with evidence.
When to reach for quantitative methods instead
Some decisions deserve better than buckets. Reach for quantification when real money hinges on the answer, whether sizing cyber insurance cover, choosing between a ยฃ300k and an ยฃ800k control investment, or reporting a credible loss figure to a board that will act on it. The FAIR model, calibrated estimation as described in Hubbard and Seiersenโs How to Measure Anything in Cybersecurity Risk, and simple Monte Carlo simulation in a spreadsheet are all within reach of a normal security team. Estimating a 5 to 15 per cent annual likelihood of a ยฃ2m to ยฃ8m loss and simulating it takes an afternoon, and it survives scrutiny in a way that โ4 x 4 = 16โ does not.
Loss exceedance curve showing the probability of exceeding a given annual loss, with points marked at two million and eight million pounds
A loss exceedance curve gives a distribution a CFO can price, rather than a single ordinal score.
The pragmatic position is a hybrid, with the matrix for the broad register and for communication, and quantification for the handful of risks where the decision is expensive. That is proportionate, and it is defensible to an auditor and a CFO alike.
A short test for your own register
Pull up your register and check four things. Do the likelihood and impact labels have published numeric ranges? Is there a stated time horizon? Is anyone summing or averaging the scores? What fraction of risks sit in the middle three cells? If the answers are no, no, yes and most of them, the matrix is decorating your decisions rather than informing them. Fix the scales first, because it is a week of work and it rescues the tool you already have.
Run your GRC programme in your own network.
RaptorGRC Community Edition is free โ every module, offline licence activation, nothing phones home.
Register / Download Contact us