BLOG
Data Classification Schemes That People Actually Apply
Published 5 October 2026 · By P Larner
Security in Practice#handling rules#data classification#data loss prevention#labelling#retention
Data Classification Schemes That People Actually Apply
Illustration for: Data Classification Schemes That People Actually Apply
What classification is actually for
Classification is not an end in itself. It exists so that decisions about handling can be made without a conversation. That is the entire value proposition, and it explains why most schemes fail.
If a label does not change what somebody may do with the document, the label is decoration. The test of a scheme is whether a person holding a labelled item knows, without asking, whether they may email it externally, put it in a shared drive, take it home or send it to a supplier.
Which means the handling rules are the product and the labels are just the index. Schemes that spend their design effort on the definitions and then leave handling to a table nobody reads have built the index and skipped the book.
Four labels beat seven
The strongest predictor of whether a scheme gets applied is how many levels it has. Every level you add multiplies the decisions a person has to make and the boundary cases they have to resolve.
Four is the sweet spot for most organisations. Three works if your business is simple. Five is defensible in regulated environments with genuinely distinct handling regimes. Beyond that, people stop distinguishing and default to the middle, which is exactly the herding behaviour that ruins risk matrices.
A workable four level scheme looks like this, and the names matter less than the boundaries.
Diagram of a four level classification scheme with handling rules at each level
The label is the index and the handling rules are the product. If two levels share the same handling, one of them is unnecessary.
The useful discipline is that every level must differ from the one below it in at least one handling rule that people will notice. If Internal and Confidential can both be emailed internally, stored in the same place and shared with the same suppliers, then in practice you have three levels and a naming convention.
Name them for the boundary, not the feeling
“Highly Confidential” and “Confidential” are almost impossible to tell apart in the moment. “Confidential” and “Client Confidential” are easy, because the second one names the boundary it must not cross. Wherever you can, use names that describe who the item must stay inside rather than how serious it feels.
The rule people can follow in three seconds
Here is the test I apply to any scheme. Give somebody a document they have never seen and ask them to classify it. If it takes longer than a few seconds, or if two competent people disagree, the scheme will not survive contact with a busy inbox.
Achieving that speed means writing the decision as a question rather than a definition. Compare these two.
The definition version says that Confidential information is information whose unauthorised disclosure would cause significant harm to the organisation, its clients or its employees. Every word is defensible and nobody can apply it quickly.
The question version says, would we be comfortable with this appearing in a competitor’s inbox? If not, it is at least Confidential. That is answerable immediately and it is right often enough.
Write both. Put the question in the guidance people read and the definition in the policy the auditor reads.
Default to the middle, deliberately
Every scheme needs a defined default for unlabelled material, and most schemes leave this to chance. The choice matters more than the levels.
Defaulting everything to the highest level produces immediate over-classification, and over-classification is not the safe option people assume. It trains staff to ignore labels, it obstructs legitimate work, and it makes genuinely sensitive material indistinguishable from routine material. Defence and government environments know this failure well, and it is why over-classification is treated there as a security problem rather than an excess of caution.
Defaulting to the lowest level is worse for obvious reasons.
The workable answer is to default to your Internal level, apply it automatically wherever your tooling allows, and require positive action to move something up or down. Most material genuinely is internal, so the default is usually correct, and the exceptions are where you want the human attention.
Automate the label, not the judgement
Classification labelling tools have become good enough that manual labelling of everything is no longer a sensible ask. The realistic split is that tooling applies the default and enforces the handling, while people make the exceptions.
Three things are worth automating early. The default label on every new document and message. The handling enforcement, so that the label actually blocks or warns on external sharing rather than merely describing what should happen. And the reporting, so somebody can see how much of the estate carries each label and whether that distribution is plausible.
That last one is the most useful and the least implemented. A scheme where ninety per cent of documents are Confidential is not a scheme with a lot of sensitive data, it is a scheme people have stopped thinking about.
Watch what the distribution tells you
Healthy distributions are lopsided towards the middle and thin at the top. If your top level holds more than a small percentage of the estate, either the boundary is drawn wrong or the culture has defaulted upward. Both are fixable, and neither is visible without the report.
Where classification meets everything else
A classification scheme is not a standalone control. It is the input several other disciplines are waiting for.
Discipline | What it needs from classification |
|---|---|
Access control | Which groups may reach which material, expressed once rather than per system |
Data loss prevention | The signal that decides whether to block, warn or allow |
Retention and deletion | Different clocks for different sensitivities |
Supplier assurance | Which categories may leave the organisation and to whom |
Incident severity | Whether a lost laptop is a nuisance or a notifiable event |
Breach assessment | The first input to the risk to individuals question under UK GDPR |
That last row is the one that turns classification from housekeeping into a legal capability. An organisation that can say immediately what was in a compromised share is in a very different position from one that has to go and look, with a 72 hour clock running.
The honest counterarguments
Users cannot be expected to classify anything
There is real evidence behind this. Manual classification at the point of creation has a poor track record, people are busy, and the incentive to think carefully about a label is close to zero. The response is not to abandon the scheme but to reduce what humans are asked to do. Automate the default, ask for a decision only when material moves outward, and accept that a good default plus a small number of correct exceptions beats a perfect scheme nobody applies.
Classification is a control that mostly generates work
Sometimes true, and it is worth being honest about the cost. A scheme that produces labels but changes no handling, blocks no sharing and informs no decision is pure overhead. The distinguishing question is whether anything downstream consumes the label. If nothing does, fix that before adding levels.
Modern tooling makes labels obsolete
The argument goes that content inspection can identify sensitive data on the fly, so a human applied label is redundant. Content inspection is genuinely good at recognisable formats such as card numbers and national insurance numbers, and genuinely poor at recognising a strategy document, a legal position or an unannounced acquisition. Those are exactly the items where the classification decision matters most, and only a person knows.
Common mistakes
Designing the scheme before the handling rules. If you cannot write down what changes at each level, you are not ready to name the levels.
Copying a government scheme wholesale. OFFICIAL, SECRET and TOP SECRET carry decades of accumulated handling infrastructure. Borrowing the words without the infrastructure produces labels with no consequences.
Classifying documents and forgetting the containers. A Confidential document in an unrestricted shared drive is not protected by its label. The handling rules have to reach the storage location, not just the file.
Ignoring the legacy estate. The scheme applies to new material immediately and to the twenty years of existing material never. Decide deliberately whether you are labelling the past, and if not, say so, because otherwise the gap gets discovered during an incident.
No review of the boundaries. Business changes, and material that was Confidential in 2022 may be routine now. A scheme nobody has revisited will drift towards over-classification with time, because moving something up is easy and nobody ever moves anything down.
Where to start
Do not start with the policy. Take the ten kinds of information your organisation actually handles, write them on a list, and sort them into piles by how they should be handled rather than how sensitive they feel. Whatever number of piles you end up with, that is your number of levels, and the sorting rules you used are your definitions.
Then write the handling table first, on one page, and check that each level differs from its neighbour in something a person would notice. Only after that has survived a conversation with the people who will apply it should anybody write the policy, because the policy is the last artefact rather than the first, and treating it as the deliverable is how organisations end up with a beautifully written scheme and an unclassified estate.
Run your GRC programme in your own network.
RaptorGRC Community Edition is free: every module, offline licence activation, nothing phones home.
Register / Download Contact us