Framework
NIST AI RMF 1.0: A Practical Audit-Readiness Blueprint
How UK and EU compliance teams turn the NIST AI Risk Management Framework's four functions into evidence that survives a customer assessment or a regulator's questions.
A procurement pack arrives from a large customer, or from a US federal prime contractor, and one line in it stops the room: "Describe your organisation's alignment with the NIST AI Risk Management Framework 1.0." There is no certificate to attach. There is no audit report to point at. Somebody drafts a page of prose about responsible AI, it comes back with follow-up questions, and the deal slows down while the compliance team tries to work out what the buyer actually wants to see.
That is the real shape of this problem for compliance officers, risk leads and DPOs in the UK and EU. The framework is not a rulebook you fail the way you fail a financial audit. It is a set of outcomes, and the assessment that matters — whether it comes from a customer, an internal auditor or eventually a regulator — is whether you can produce evidence that those outcomes are being achieved for named systems. Most organisations already hold some of that evidence, scattered across a model register, a data protection impact assessment, a supplier questionnaire and somebody's inbox. The work is assembling it before it is demanded.
What the framework is, and what it is not
The AI RMF 1.0 is voluntary guidance published by the US National Institute of Standards and Technology. It is written as outcomes to be achieved rather than controls to be implemented, and it is deliberately sector-neutral. A companion Playbook suggests actions against each outcome, and a Generative AI Profile extends the framework to risks specific to generative systems. Underneath it sit seven characteristics of trustworthy AI, including validity and reliability, safety, security and resilience, accountability and transparency, explainability, privacy enhancement, and managed harmful bias.
Three points are worth settling before you plan any work:
- There is no NIST AI RMF certification. No accredited conformity assessment scheme exists for it. Anyone offering to certify you against it is selling their own opinion. Where a customer genuinely needs a certificate, the certifiable route is ISO/IEC 42001, the AI management system standard, which can be audited by an accredited body.
- "Mandatory for US federal agencies" is a simplification. Obligations on agencies flow from executive direction and central budget-office policy that reference the framework; the framework itself confers no compliance status on anyone. For a UK or EU supplier it reaches you through the contract, flowed down by a customer who has their own obligations.
- It is not a substitute for law. Alignment with the RMF discharges no obligation under the EU AI Act or under UK and EU data protection law. It is a way of organising the work and answering questions consistently, not a legal defence.
The four functions, translated into evidence
The framework has four functions: GOVERN, MAP, MEASURE and MANAGE. GOVERN is cross-cutting; the other three describe a lifecycle. The most useful way to read them is as four questions an assessor will ask about one specific system, and then to ask what document in your organisation answers each one today.
| Function | What an assessor is really asking | Evidence that answers it |
|---|---|---|
| GOVERN | Who is accountable if this system causes harm, and do they have the authority to stop it? | Named owner with delegated authority to suspend the system; policy signed off at the right level; minutes showing a real decision was taken, not just a framework adopted |
| MAP | Do you know where this system is used, by whom, and who it can affect? | Inventory entry recording intended purpose, declared out-of-scope uses, affected groups, foreseeable misuse, and third-party model dependencies |
| MEASURE | How do you know it works, and how would you know if it stopped working? | Test results against thresholds set before testing; bias and robustness evaluation with the date it was last run; production monitoring with defined trigger points |
| MANAGE | What happens when a measurement crosses a line? | Recorded risk treatment decisions including accepted risks; incident definition and escalation route; a tested rollback or suspension; supplier notification terms |
In practice the gap is almost always in MEASURE. Teams can describe their governance and can list their systems, but cannot produce a test result with a threshold attached to it, or say when the model was last evaluated against a population that resembles the real one. A governance structure with nothing measured underneath it is the failure mode that assessors find quickest.
One evidence set, three audiences
For a UK or EU organisation the return on this work does not come from NIST alone. The same artefacts answer several demands at once, which is the argument for doing it properly rather than writing a bespoke response to each questionnaire.
The EU AI Act requires providers of high-risk systems to operate a risk management system across the lifecycle and a quality management system, and to conduct post-market monitoring. It places separate duties on deployers, including assigning human oversight to people with the competence, authority and support to exercise it, using input data that is relevant and sufficiently representative where the deployer controls it, and retaining system logs for a defined period. It also imposes transparency duties for certain systems, such as informing people they are interacting with an AI system and marking synthetic content, and expects providers and deployers to ensure a sufficient level of AI literacy among staff operating these systems. Note that a deployer can become a provider in law — for example by putting its own name on a system, substantially modifying it, or changing its intended purpose — so record which role you occupy for each system.
The Act's maximum penalties are up to 7% of global annual turnover for prohibited practices and up to 3% for most other obligations. UK and EU GDPR maxima reach 4%. These are ceilings, not expected outcomes, and they are set by supervisory authorities against the specific facts; quoting them internally as a forecast tends to damage your credibility rather than build urgency.
The UK has no single cross-sector AI statute. Existing regulators apply existing law to AI within their remits, which means for most organisations the immediate legal exposure sits in data protection — lawful basis, DPIAs, and the rules on solely automated decisions with legal or similarly significant effects — alongside sector rules such as model risk expectations in financial services. Public bodies additionally face transparency reporting expectations for algorithmic tools.
The audit-readiness checklist
Build this per system, not per organisation. A named assessor will pick one system and pull the thread.
- An inventory entry with intended purpose, declared out-of-scope uses, deployment status, supplier, and model or version identifier
- A recorded determination of whether you are provider or deployer, and whether the use case sits near a high-risk category or a prohibited practice
- A named accountable owner at a level with authority to suspend the system, plus their deputy
- A risk assessment covering affected people, foreseeable misuse and failure modes — not only information security
- Pre-deployment test results, with the acceptance thresholds recorded before testing rather than after
- A bias and performance evaluation, with the date it was last run and who reviewed it
- Evidence of human oversight in operation: who reviews what, at what sampling rate, and cases where a human actually overrode the system
- Logging sufficient to reconstruct a single decision months later, with a stated retention period
- Production monitoring with defined trigger points and a route for what happens when one fires
- A written definition of an AI incident, an escalation route, and a record of at least one exercise
- Supplier evidence: contractual notice of model changes, the right to test, and what happens on withdrawal of a model version
- Change control showing what changed, who approved it, and what was re-tested
Where this stops
This blueprint covers assurance and evidence. It does not cover several things you may still need. Whether a specific system falls into a high-risk category, or whether a particular use amounts to a prohibited practice, is a legal determination that depends on facts about purpose and deployment — take advice rather than deciding it in a workshop. National implementation, supervisory structures and the harmonised standards that will give concrete conformity criteria are still maturing, so anything you build now should be designed to absorb change.
Sector-specific regimes sit on top and are not addressed here: medical devices, financial services model risk and consumer duty, employment law where screening tools are involved, and any regime governing safety-critical use. Selling into US federal supply chains brings its own security assurance and contracting requirements that are separate from the RMF. If your system processes special category data, makes decisions with legal or similarly significant effects on individuals, or has already produced an outcome someone has complained about, involve your DPO and legal counsel before publishing any assessment. Nothing here is legal advice, and no framework substitutes for a lawyer who knows your facts.
Watch this as a video
- NIST AI RMF
- AI governance
- Audit readiness
- EU AI Act
- ISO 42001
- Evidence