RaptorGRC — Offline GRC

BLOG

Frontier AI What the Term Actually Means

Published 5 August 2026 · By P Larner

Security in Practice#frontier models#artificial intelligence#EU AI Act#AI Security Institute#regulation

Frontier AI: What the Term Actually Means

Illustration for: Frontier AI, What the Term Actually Means

Illustration for: Frontier AI, What the Term Actually Means

A board member asked me last year whether the company was exposed to frontier AI. It was a fair question and nobody in the room could answer it, because three of us had three different definitions in our heads. One meant the chatbot in the service desk. One meant the research labs in California. One meant robots. The risk register said nothing, which was the only honest position available at the time.

Frontier AI is not a marketing phrase, or at least it did not start as one. It is a category with thresholds, regulators and reporting duties attached, and the distinction between a frontier model and the ordinary machine learning already scattered through your estate changes what you have to do about it. This post is about where the line sits and who drew it.

Where the term came from

The phrase entered serious use in 2023, in a paper on frontier AI regulation written jointly by researchers from several of the labs building the models, and it was adopted almost immediately by the UK government for its Frontier AI Taskforce ahead of the Bletchley Park summit that November. The working definition then was blunt. A frontier model is a highly capable general purpose model that could possess dangerous capabilities sufficient to pose severe risks to public safety.

Two words in that definition do the work

General purpose means the model was not built for one task, so its capabilities are discovered after training rather than specified before it. Dangerous capabilities means the concern is not bias or inaccuracy, which are real but ordinary, but a narrow set of severe harms, chiefly assistance with chemical, biological, radiological and nuclear weapons, offensive cyber capability at scale, and models that can act on the world without a human in the loop.

That is a safety definition rather than a legal one. The legal definitions came later and they are narrower.

The definitions that carry legal force

The EU AI Act is the only regime that turns the idea into hard obligation. It does not use the word frontier. It defines general purpose AI models, then creates a subset presumed to carry systemic risk when the cumulative compute used to train the model exceeds ten to the power twenty five floating point operations. Cross that line and you owe the Commission notification, model evaluation, adversarial testing, incident reporting and cyber security protection for the model weights themselves. Obligations for general purpose models started applying in August 2025.

The United Kingdom took the other road. There is no frontier AI statute. Instead the AI Safety Institute, renamed the AI Security Institute in February 2025, tests models by agreement with the developers, and the safety commitments made at Seoul in 2024 are voluntary undertakings rather than law. The NCSC treats the question as one of secure development and supply chain, which is a familiar frame for anyone who has read its other guidance.

The United States has moved twice in opposite directions. The 2023 executive order set a reporting threshold at ten to the power twenty six operations, an order of magnitude above the European figure, and it was rescinded in January 2025. What survives sits at state level, most visibly California’s transparency requirements on large developers, which oblige published safety frameworks and incident reporting rather than pre-approval.

Pyramid diagram of three tiers of artificial intelligence systems

Pyramid diagram of three tiers of artificial intelligence systems

Almost everything in an ordinary organisation sits in the bottom two tiers. The regulatory weight described here lands on the top one, and on very few developers worldwide.

Regime

What it names

The threshold

What it triggers

EU AI Act

General purpose model with systemic risk

Training compute above ten to the twenty fifth

Notification, evaluation, incident reporting, weight security

UK

Frontier model, by convention only

No statutory figure

Voluntary testing with the AI Security Institute

US federal

Dual use foundation model, order rescinded

Was ten to the twenty sixth

Nothing currently in force

California

Large frontier developer

Compute and revenue tests

Published safety framework, incident reporting

What actually separates frontier from merely large

A compute number is a proxy, not a definition, and everyone involved knows it. Three properties matter more than the arithmetic.

Generality

A fraud model scores transactions. A frontier model will write the fraud policy, argue with you about it, translate it into Polish and then find a way round it if you ask nicely. Capability that was never specified cannot be tested against a specification, which is why evaluation of these systems looks like adversarial red teaming rather than acceptance testing.

Capability that surprises the people who built it

This is the property that makes the category unlike the rest of your estate. Vendors of ordinary software know what their product does. Frontier developers publish capability evaluations after training because they are finding out too, and a supplier who cannot fully enumerate their own product’s behaviour is a supplier assurance problem regardless of the technology involved.

The ability to act rather than answer

A model wired to tools, credentials and a task loop is no longer a text service. It is an unaccountable user account with initiative, and that is a control question your existing framework already knows how to ask.

Why this matters to a risk practitioner

You will almost certainly never train one of these models. You are exposed at three removes and each one belongs somewhere different in your register.

  1. You use one. Your data leaves your boundary, and every prompt is a disclosure to a supplier whose retention terms you probably have not read.
  2. Your suppliers use one. Their support desk, their code, their document review. The exposure arrives through the contract rather than the network.
  3. Your adversaries use one. Capability that once required a skilled operator is now available to anyone who can pay a subscription, which changes the frequency side of every social engineering scenario you have.

The third of those deserves its own treatment and the next post in this series gives it one. The point here is that the first two are procurement and supplier assurance problems in a new costume, and the discipline you already run for cloud services applies with almost no modification.

The honest counterarguments

The threshold is arbitrary and everybody admits it

Ten to the twenty five is a line drawn where the largest models of 2023 sat. Efficiency improves, so a model trained below the line in 2027 may outperform one trained above it in 2024. Regulating by input is regulating the wrong variable, and the Act’s own drafting acknowledges this by allowing the Commission to designate models below the threshold.

The term does commercial work for the people who coined it

A definition that captures only the largest laboratories is also a definition that raises the drawbridge behind them, and the same firms that wrote the early papers benefit from a rule that treats their scale as a regulatory moat. That does not make the safety case wrong. It does mean the word should be read with the same scepticism you would apply to any vendor category.

Most organisations have the opposite problem

The realistic harm in an average British company this year is a member of staff pasting customer records into a free tool, not a laboratory model escaping supervision. Attention spent on the frontier is attention not spent on shadow adoption, and I would rather see a company solve the second before it writes a policy on the first.

What to write down

Answer four questions in one page and you are ahead of most boards.

  1. Which of these models are we using, deliberately or otherwise, and where does the data go?
  2. Which of our suppliers are using one inside a service we depend on?
  3. Which of our decisions could be made or influenced by one without a named human owner?
  4. Who in this organisation is accountable for answering the first three next year, when the answers have changed?

The board member who asked the question deserved a better answer than the room gave. It is this. Frontier AI means a small number of general purpose models capable enough that regulators want to know before they ship, you will never own one, and your exposure runs through the people you buy from and the people who attack you. Write that down, name the owner, and revisit it when the threshold moves, because it will.

Run your GRC programme in your own network.

RaptorGRC Community Edition is free — every module, offline licence activation, nothing phones home.

Register / Download Contact us
An unhandled error has occurred. Reload 🗙

Rejoining the server...

Rejoin failed... trying again in seconds.

Failed to rejoin.
Please retry or reload the page.

The session has been paused by the server.

Failed to resume the session.
Please retry or reload the page.