NIST AI Risk Management Framework (AI RMF 1.0)
Evidence request list. 72 controls, 72 carrying auditor artefact guidance. Generated from the compliance knowledge graph on 12 September 2026. Published by The Art of Service.
GOVERN - NIST AI RMF 1.0
Legal and regulatory requirements involving AI are understood, managed, and documented. The organisation identifies which AI-specific and AI-adjacent legal duties bind each system, records how each is discharged, and keeps that record current as the applicable law changes across the jurisdictions the system reaches.
- A register of legal and regulatory requirements applicable to each AI system, naming the instrument and the obligation
- Mapping from each identified obligation to the control, process or artefact that discharges it
- Evidence of legal review at design and before deployment, with reviewer and date
- A record of jurisdictions the system is deployed into and the requirements that follow from each
- Change records showing the register was updated when applicable law changed
- A generic corporate compliance register with no AI-specific entries
- Requirements listed but never mapped to a control, so nothing shows they are met
- Register maintained for the launch jurisdiction only while the system is reachable elsewhere
The characteristics of trustworthy AI are integrated into organizational policies, processes, and procedures. Validity, reliability, safety, security, resilience, accountability, transparency, explainability, privacy and fairness appear as named obligations inside the policies that govern the AI lifecycle, not only in a statement of values.
- Approved AI policy set naming each trustworthiness characteristic as an obligation with an owner
- Procedure documents showing where in the lifecycle each characteristic is assessed
- Approval record with authority and effective date, and the scheduled review interval
- Evidence the policies were communicated to the teams that design, develop and deploy AI
- Trustworthiness stated as principles in a values page with no procedural consequence
- Policy covers security and privacy only, leaving fairness and explainability unaddressed
- No review since first issue, so the policy predates the systems it is meant to govern
Processes and procedures are in place to determine the needed level of risk management activities based on the organization's risk tolerance. Risk management effort is allocated by a documented rule rather than by availability, so a system assessed as higher risk demonstrably receives more scrutiny than a lower-risk one.
- The documented risk tiering criteria and the scale risks are assigned to
- Tier assignment records for the AI systems in the inventory, with the reasoning
- The activity set required at each tier, showing what changes between tiers
- Evidence that a high-tier system actually received the additional activities its tier requires
- Risk tolerance expressed only as a word such as low or moderate, with no scale behind it
- Every system assigned the same tier, which makes the tiering inert
- Tier assigned at intake and never revisited after the system's use expanded
The risk management process and its outcomes are established through transparent policies, procedures, and other controls based on organizational risk priorities. The process and the decisions it produced are both documented in a form a reviewer outside the delivery team can follow, including who decided and on what basis.
- The documented AI risk management process, with the decision points it contains
- Standardised documentation templates used across AI work products
- Completed records for deployed systems showing the process was followed, not just defined
- Named accountable contacts recorded on work products so decisions are traceable to people
- Process defined centrally but the delivery records show a different practice
- Outcomes recorded without the reasoning, so the decision cannot be reviewed
- Documentation held in individual team tools with no organisational retention
Ongoing monitoring and periodic review of the risk management process and its outcomes are planned, organizational roles and responsibilities are clearly defined, including determining the frequency of periodic review. Monitoring and review are scheduled with a stated frequency and a named owner, and the review actually produces recorded findings rather than a confirmation that nothing changed.
- The monitoring plan naming what is monitored, by whom and at what frequency
- Review schedule with the basis for the chosen interval
- Completed review records with findings, decisions and follow-up actions
- Role definitions assigning monitoring and review duties to named functions
- Review frequency stated but no completed review on file for the current period
- Reviews recorded as no change with nothing examined to support that conclusion
- Monitoring owned by the team that built the system, with no independent line
Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities. A maintained inventory identifies the AI systems and models in use, and the resource given to maintaining it is proportionate to the risk the inventoried systems carry.
- The AI system and model inventory with the fields it captures per entry
- The intake process that causes a new system or model to be added
- Evidence of periodic reconciliation between the inventory and systems actually in production
- Named owner and resourcing for inventory maintenance
- Inventory covers models built in house but omits embedded and third-party AI features
- Compiled once for an audit and not maintained since
- No link from an inventory entry to its documentation, so the entry carries no usable detail
Processes and procedures are in place for decommissioning and phasing out of AI systems safely and in a manner that does not increase risks or decrease the organization’s trustworthiness. Retirement is a governed event with retention, user notification and downstream dependency steps, rather than deletion at the point the system stops being useful.
- The documented decommissioning procedure, including retention and legal hold steps
- User and downstream consumer notification records for retired systems
- Dependency analysis showing what consumed the system before it was withdrawn
- Completed decommissioning records for systems actually retired
- Procedure exists but retirements are executed as an infrastructure ticket outside it
- Model artefacts and training data deleted while a regulatory retention duty still applied
- Downstream consumers discovered the retirement from a failure rather than a notice
Roles and responsibilities and lines of communication related to mapping, measuring, and managing AI risks are documented and are clear to individuals and teams throughout the organization. Duties for mapping, measuring and managing AI risk are assigned to identifiable roles with escalation paths, and the people holding them can state what they are accountable for.
- A responsibility assignment covering the map, measure and manage duties
- Organisational chart or charter showing the reporting line for AI risk functions
- Escalation paths with named recipients and thresholds
- Evidence the assignment was communicated and acknowledged by role holders
- Responsibility assigned to a committee rather than to a person who can be asked
- Test and evaluation reporting into the same manager who owns the delivery deadline
- Assignment current at the last reorganisation rather than at the present one
The organization’s personnel and partners receive AI risk management training to enable them to perform their duties and responsibilities consistent with related policies, procedures, and agreements. Training is role-specific and its completion is recorded, and partners bound by agreements are covered as well as employees.
- Training curricula differentiated by AI role, with the content covered
- Completion records by individual and date, including partner personnel where in scope
- The link from a role definition to the training required for it
- Evidence of refresh when policies or applicable law changed
- One general awareness module issued to everyone regardless of AI duty
- Partner and contractor personnel excluded despite performing AI lifecycle tasks
- Completion tracked but no record of what the training actually covered
Executive leadership of the organization takes responsibility for decisions about risks associated with AI system development and deployment. Named executives hold and exercise the decision right over AI risk acceptance, with their decisions minuted, rather than delegating it into the delivery line.
- The charter or terms of reference assigning AI risk decisions to named executives
- Minutes recording executive decisions to accept, mitigate or reject AI risks
- The stated organisational appetite for AI risk, approved at executive level
- Evidence of the authority and budget granted to the accountable officer
- Executive sponsorship claimed with no minuted decision to point to
- Risk acceptance signed by the project owner who benefits from proceeding
- Appetite endorsed once at programme start and never revisited as systems changed
Decision-makings related to mapping, measuring, and managing AI risks throughout the lifecycle is informed by a diverse team (e.g., diversity of demographics, disciplines, experience, expertise, and backgrounds). The people making AI risk decisions span disciplines and backgrounds, and where internal composition is narrow, external input is obtained and recorded.
- Records of the disciplines represented on AI risk decision forums
- Documented consultation with external perspectives where internal expertise was thin
- The organisational commitment on composition of AI risk teams
- Evidence that a dissenting or external view changed a decision
- Diversity asserted from headcount data with no bearing on who decides
- External consultation held after the design was fixed, so it could not affect it
- Same three technical staff constitute every review
Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems. The policy distinguishes the human roles around a system, operator, reviewer, overseer, affected party, and states what each may and must do about its output.
- Policy defining each human role in the human-AI configuration and its authority
- Per-system documentation of the configuration chosen and why
- Oversight procedures stating when a human may override or must intervene
- Competency requirements attached to each oversight role
- Human in the loop asserted without stating what the human is empowered to change
- Oversight assigned to a role with no authority to stop the system
- Configuration documented at design and not updated after automation increased
Organizational policies, and practices are in place to foster a critical thinking and safety-first mindset in the design, development, deployment, and uses of AI systems to minimize negative impacts. Practices exist that make raising a concern normal and consequential, such as separated accountability lines or a recorded route for challenging a deployment decision.
- Policy establishing separated accountability for development, risk and assurance
- The internal route for raising an AI safety concern and the protection attached to it
- Records of concerns raised and how each was resolved
- Evidence a concern changed a design or delayed a deployment
- A stated speak-up culture with no record of anything ever being raised
- Assurance and development reporting to the same accountable owner
- Concerns logged and closed without a recorded resolution
Organizational teams document the risks and potential impacts of the AI technology they design, develop, deploy, evaluate and use, and communicate about the impacts more broadly. Impact documentation is produced by the team that builds the system and is communicated beyond it, so risk knowledge does not stay inside the delivery group.
- Completed impact assessments for AI systems, with author and date
- The template or structure the assessments follow
- Evidence the assessment was communicated outside the delivery team
- The link from an assessment finding to a risk register entry or design change
- Assessment completed as a launch gate and never revisited
- Findings recorded with no route into the organisational risk register
- Only harms to the organisation considered, not harms to individuals or communities
Organizational practices are in place to enable AI testing, identification of incidents, and information sharing. Testing and incident identification are enabled by practice and resourcing, and information about failures is shared with the parties who need it rather than contained.
- In-house testing policy covering the failure modes AI introduces
- The incident identification and recording mechanism for AI systems
- Records of information shared internally or externally about identified issues
- Evidence of resourcing for testing independent of the delivery schedule
- Testing limited to accuracy on a held-out set, missing drift and shortcut learning
- Incidents recorded in a delivery tracker that no risk function reads
- No sharing beyond the team, so the same failure recurs elsewhere
Organizational policies and practices are in place to collect, consider, prioritize, and integrate feedback from those external to the team that developed or deployed the AI system regarding the potential individual and societal impacts related to AI risks. External feedback is solicited under a policy, prioritised on a stated basis and carried into the system, rather than collected and filed.
- Policy establishing how external feedback on AI impacts is collected
- Records of stakeholder engagement activity, including who was engaged
- The prioritisation basis applied to feedback received
- Traceability from a specific piece of feedback to a change made
- Feedback channel published with no process behind it
- Engagement limited to existing customers, missing affected non-users
- Feedback logged with no evidence any of it was acted on
Mechanisms are established to enable AI actors to regularly incorporate adjudicated feedback from relevant AI actors into system design and implementation. There is a mechanism that adjudicates feedback and routes the accepted items into design and implementation on a regular cycle.
- The adjudication mechanism and who holds the decision on what is accepted
- The cadence at which adjudicated feedback enters design or implementation
- Records of feedback adjudicated, with accept or reject reasoning
- Change records showing accepted feedback implemented
- Feedback adjudicated informally, so rejections carry no reasoning
- Accepted items queued indefinitely with no implementation cycle
- Mechanism covers internal actors only and excludes deployers and end users
Policies and procedures are in place that address AI risks associated with third-party entities, including risks of infringement of a third party’s intellectual property or other rights. Third-party AI risk is addressed by policy covering data, models, software and services obtained externally, including the rights position on training data and model outputs.
- Third-party AI policy covering data, pre-trained models, software and services
- Due diligence records for third-party AI components in use
- Contract terms addressing intellectual property, data rights and liability for AI components
- The intellectual property position recorded for training data and model outputs
- Standard vendor due diligence applied with no AI-specific questions
- Open source models adopted with no review of the licence or the training data provenance
- Policy addresses suppliers but not freely obtained models and datasets
Contingency processes are in place to handle failures or incidents in third-party data or AI systems deemed to be high-risk. For third-party components assessed as high risk there is a worked contingency, a fallback, a redundancy or a defined manual path, and it has been tested.
- Identification of third-party AI components assessed as high risk
- The contingency defined for each, naming the fallback or redundancy
- Test or exercise records demonstrating the contingency works
- Trigger criteria and the named role that invokes the contingency
- Contingency documented as revert to manual with no assessment of whether that is possible
- Never exercised, so the fallback capacity is unverified
- Applies to the primary vendor only while a sub-processor carries the same dependency
MANAGE - NIST AI RMF 1.0
A determination is made as to whether the AI system achieves its intended purpose and stated objectives and whether its development or deployment should proceed. A recorded determination weighs measured risk against benefit and decides whether to proceed, so that not proceeding remains an available and evidenced outcome.
- The recorded determination on whether the system achieves its intended purpose
- The measured risk and benefit evidence the determination rested on
- The decision to proceed or not, with the deciding authority named
- Consideration of whether an AI system is the appropriate solution at all
- Determination recorded as approval with no evidence weighed
- No pathway by which the answer could have been not to proceed
- Decision taken by the delivery owner rather than by an authority able to stop it
Treatment of documented AI risks is prioritized based on impact, likelihood, or available resources or methods. Treatment order follows a stated prioritisation basis rather than convenience, and the basis is visible in the treatment record.
- The risk treatment register showing priority assigned per risk
- The prioritisation basis applied, whether impact, likelihood, resource or method availability
- Evidence that higher-priority risks were treated first
- The link from prioritisation to the organisational risk tolerance
- Priority assigned but treatment order does not follow it
- Prioritisation basis unstated, so the ordering cannot be reviewed
- Low-cost treatments completed first regardless of the risk they address
Responses to the AI risks deemed high priority as identified by the Map function, are developed, planned, and documented. Risk response options can include mitigating, transferring, avoiding, or accepting. High-priority risks each carry a planned and documented response naming which of the four options was chosen, with acceptance recorded as a decision rather than as inaction.
- Documented responses for each high-priority risk, naming the option chosen
- Response plans with owner, action and target date
- Acceptance decisions recorded with the accepting authority
- Traceability from the mapped risk to its response
- Acceptance by default because no response was ever planned
- Response option not named, so mitigation and acceptance are indistinguishable
- Plans without owners or dates, so completion cannot be judged
Negative residual risks (defined as the sum of all unmitigated risks) to both downstream acquirers of AI systems and end users are documented. What remains unmitigated is totalled and documented, and communicated to downstream acquirers and end users who inherit it.
- The documented residual risk position after treatment
- Identification of downstream acquirers and end users who bear residual risk
- Evidence residual risk was communicated to those parties
- The authority that accepted the residual position
- Residual risk computed internally and never disclosed downstream
- Individual residual risks listed with no aggregate position
- Residual position stale because treatments changed after it was recorded
Resources required to manage AI risks are taken into account, along with viable non-AI alternative systems, approaches, or methods – to reduce the magnitude or likelihood of potential impacts. The resource cost of managing the risk is accounted for and non-AI alternatives are genuinely analysed as a way of reducing impact, not noted and dismissed.
- The resources required to manage identified AI risks, accounted for and allocated
- Analysis of viable non-AI alternative systems, approaches or methods
- The trade-offs weighed between trustworthiness characteristics and the alternatives
- Interdisciplinary input recorded in the alternatives analysis
- Non-AI alternatives named in a sentence with no analysis
- Risk management resourced from slack capacity rather than allocated
- Trade-offs decided implicitly by the technical team
Mechanisms are in place and applied to sustain the value of deployed AI systems. There are applied mechanisms, retraining, recalibration or refresh, that maintain the system's value against drift, and evidence they are used rather than merely available.
- The mechanisms in place to sustain value, such as retraining or recalibration
- Records showing those mechanisms were applied, with dates
- The trigger conditions that cause them to be applied
- Evidence of the effect on performance after application
- Retraining capability exists but has not been exercised since deployment
- Triggers defined against calendar time rather than observed degradation
- Effect of retraining unmeasured, so sustainment is assumed
Procedures are followed to respond to and recover from a previously unknown risk when it is identified. A followed procedure exists for risks that were not anticipated, covering response and recovery, and it has been used or exercised.
- The documented procedure for responding to and recovering from previously unknown risks
- Records of the procedure being followed for an actual or exercised event
- The route by which a previously unknown risk is escalated
- Post-event review and the changes it produced
- Procedure written for known incident types only
- Never exercised, so recovery capability is untested
- Escalation route defined but not known to the operators who would use it
Mechanisms are in place and applied, responsibilities are assigned and understood to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use. The ability to bypass, disengage or deactivate the system exists technically, is assigned to an understood role, and has been demonstrated to work.
- The technical mechanism to supersede, disengage or deactivate the system
- The role assigned to invoke it and evidence that role understands the duty
- The criteria that trigger disengagement
- Test or exercise evidence that disengagement works and the fallback holds
- Kill switch assumed to exist but never tested
- Deactivation requires a change release, so it cannot be invoked at the speed the risk demands
- Authority to disengage held by someone with a competing delivery incentive
AI risks and benefits from third-party resources are regularly monitored, and risk controls are applied and documented. Third-party data, model, software and hardware dependencies are monitored on an ongoing basis, not assessed once at procurement, and the controls applied are recorded.
- The inventory of third-party resources the AI system depends on
- Monitoring records showing regular review of those dependencies
- The risk controls applied to each and evidence they are in place
- The route by which a third-party change reaches the risk owner
- Third parties assessed at onboarding and not monitored afterwards
- Model or API version changes by the provider go unnoticed
- Controls documented in the contract with no operational evidence
Pre-trained models which are used for development are monitored as part of AI system regular monitoring and maintenance. Pre-trained and transfer-learned models are treated as a monitored component in their own right, since their provenance and behaviour are not controlled by the deploying organisation.
- Identification of pre-trained models used, with version and provenance
- Monitoring records covering those models within regular maintenance
- Assessment of risks carried over from the pre-training data and objective
- The procedure followed when the upstream model is updated or withdrawn
- Pre-trained model treated as a fixed dependency and excluded from monitoring
- Provenance of pre-training data unknown and unrecorded
- No procedure for an upstream model being deprecated
Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management. A single implemented plan covers the post-deployment lifecycle end to end, and each named element, including appeal and override, is actually operating.
- The implemented post-deployment monitoring plan and its scope
- Mechanisms for capturing and evaluating user and AI actor input
- The appeal and override mechanism and records of its use
- Incident response, recovery, change management and decommissioning arrangements for the system
- Plan documents each element but only performance monitoring is running
- Appeal and override present in the plan with no implemented mechanism
- Change management for the model handled outside the plan by the platform team
Measurable activities for continual improvements are integrated into AI system updates and include regular engagement with interested parties, including relevant AI actors. Improvement activities are measurable and attached to the update cycle, and interested parties are engaged as part of that cycle rather than consulted separately.
- The improvement activities integrated into the system update process, with their measures
- Records of engagement with interested parties within the update cycle
- Root cause analyses for degradation, drift, near misses and failures
- Evidence an improvement activity changed a subsequent release
- Improvement recorded as a backlog with no measure of effect
- Engagement conducted annually and disconnected from releases
- Root cause analysis performed for outages only, not for behavioural degradation
Incidents and errors are communicated to relevant AI actors including affected communities. Processes for tracking, responding to, and recovering from incidents and errors are followed and documented. Incidents and errors are communicated outward, including to affected communities, and the tracking record shows how each was identified, repaired and how the repair was distributed.
- The incident and error record showing how each was identified
- Communication records to relevant AI actors and affected communities
- Evidence of whether the error was repaired and how the repair was distributed to impacted users
- The followed process for tracking, response and recovery
- Incidents communicated internally only, never to affected communities
- Repair applied to the primary deployment while other users keep the defective version
- Record shows closure with no account of how the error was identified or fixed
MAP - NIST AI RMF 1.0
Intended purpose, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented. Considerations include: specific set or types of users along with their expectations; potential positive and negative impacts of system uses to individuals, communities, organizations, society, and the planet; assumptions and related limitations about AI system purposes; uses and risks across the development or product AI lifecycle; TEVV and system metrics. Context is captured before build: who the users are and what they expect, the positive and negative impacts of the intended uses, the assumptions and limits on purpose, and the settings the system will actually run in.
- A documented statement of intended purpose and the uses considered beneficial
- Identification of user types and their expectations of the system
- The laws, norms and expectations specific to the deployment setting
- Recorded assumptions and limitations about purpose, and the uses ruled out
- Analysis of foreseeable misuse and repurposing of the deployed system
- Purpose written at product-marketing level, too broad to bound anything
- Foreseeable misuse omitted because it is outside the intended use
- Documented once at concept and never reconciled with how the system is actually used
Inter-disciplinary AI actors, competencies, skills and capacities for establishing context reflect demographic diversity and broad domain and user experience expertise, and their participation is documented. Opportunities for interdisciplinary collaboration are prioritized. The people establishing context bring domain and user-experience expertise beyond the build team, and their participation is recorded so the basis of the context statement is auditable.
- Record of the disciplines and expertise represented in context establishment
- Documented participation of domain and user-experience specialists
- Evidence that interdisciplinary collaboration was prioritised and resourced
- The competency requirements set for context establishment work
- Context established entirely by the engineering team
- Participation asserted with no record of who took part or what they contributed
- Domain expertise consulted once rather than through the mapping work
The organization’s mission and relevant goals for the AI technology are understood and documented. The system's connection to a stated organisational goal is written down, which is what makes a go or no-go decision reviewable later.
- Documented mission and the specific goals the AI technology serves
- The linkage from system objectives to those organisational goals
- The go or no-go decision record referencing that linkage
- Evidence of review when goals or the system changed
- Goal recorded as improve efficiency with nothing measurable behind it
- Alignment written after the build to justify a decision already taken
- No record of the alternatives weighed
The business value or context of business use has been clearly defined or – in the case of assessing existing AI systems – re-evaluated. Business value is stated in terms that can be checked after deployment, and for systems already running it is re-evaluated rather than assumed from the original case.
- The documented business value or use context for the system
- Re-evaluation records for AI systems already in operation
- The measures by which value is judged after deployment
- Analysis of how the social context of use bears on the value claimed
- Value defined at business-case stage and never re-tested against outcomes
- Existing systems inherited with no re-evaluation performed
- Value stated only as cost saving, ignoring how the system is used and by whom
Organizational risk tolerances are determined and documented. Tolerance is written down at a level that lets someone decide whether a measured risk is acceptable, and it is traceable to an authority that can set it.
- The documented risk tolerance statement and its approval authority
- The criteria or scale by which a measured risk is judged against tolerance
- Any sector, regulatory or professional requirements the tolerance is derived from
- Records of decisions taken by applying the tolerance
- Tolerance expressed as a word with no scale, so nothing can be judged against it
- Set by the delivery team rather than by an authority that can accept the consequence
- Documented but never referenced in any deployment decision
System requirements (e.g., “the system shall respect the privacy of its users”) are elicited from and understood by relevant AI actors. Design decisions take socio-technical implications into account to address AI risks. Requirements are elicited from the actors who will be affected and written down, and the design record shows socio-technical implications were weighed rather than only technical ones.
- Written system requirements including non-functional and trustworthiness requirements
- Record of who requirements were elicited from and how
- Design decision records showing socio-technical implications considered
- Traceability from a requirement to the design element that satisfies it
- Requirements captured as model performance targets only
- Elicited from the commissioning stakeholder alone, missing operators and affected parties
- Requirements written but no traceability, so nothing shows they were built
The specific task, and methods used to implement the task, that the AI system will support is defined (e.g., classifiers, generative models, recommenders). The learning or decision task is stated narrowly, together with the method class used to implement it, because a narrow task definition is what makes the risk mapping tractable.
- The documented task the system performs, stated narrowly
- The method or model class used to implement the task
- The boundary of what the system does not do
- Evidence the definition was updated when the task changed
- Task defined as the product name rather than as a learning or decision task
- Method recorded as machine learning with no class identified
- Definition unchanged after a general-purpose model replaced a narrow one
Information about the AI system’s knowledge limits and how system output may be utilized and overseen by humans is documented. Documentation provides sufficient information to assist relevant AI actors when making informed decisions and taking subsequent actions. The system's knowledge limits are written down alongside how output should be used and overseen, in enough detail for a downstream actor to make an informed decision.
- Documented knowledge limits, including the conditions the system was not built for
- Guidance on permitted use and interpretation of system output
- The human oversight expected over the output, and by whom
- Evidence this documentation reaches downstream deployers and operators
- Limits known to the build team and absent from anything a deployer receives
- Output guidance written for the ideal case, silent on out-of-distribution input
- Documentation exists but is not delivered with the system
Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection (e.g., availability, representativeness, suitability), system trustworthiness, and construct validation. The evaluation design is documented as a scientific claim: what was measured, on what data, and whether the measure validly stands for the property claimed.
- Documented experimental design for evaluation of the system
- Data collection and selection decisions, with availability, representativeness and suitability addressed
- Construct validation showing the metric stands for the property claimed
- The TEVV considerations identified and how each is handled
- Benchmark accuracy reported with no argument that it measures the property claimed
- Data selection undocumented, so representativeness cannot be assessed
- Evaluation designed by the same people optimising against it
Potential benefits of intended AI system functionality and performance are examined and documented. Claimed benefits are examined and written down rather than assumed, so they can later be weighed against measured costs and residual risk.
- Documented analysis of the benefits the system is expected to deliver
- The performance basis on which each claimed benefit rests
- Identification of who receives the benefit
- Evidence the benefit analysis is revisited against realised outcomes
- Benefits asserted in a business case with no performance basis
- Benefit to the organisation documented, benefit or detriment to users not considered
- Never revisited, so claimed benefits are never confirmed
Potential costs, including non-monetary costs, which result from expected or realized AI errors or system functionality and trustworthiness - as connected to organizational risk tolerance - are examined and documented. Costs of error are examined including non-monetary ones such as harm to individuals and communities, and connected explicitly to the tolerance the organisation has set.
- Documented analysis of costs arising from system error or failure
- Non-monetary costs identified, including harm to individuals and communities
- The connection drawn between those costs and the stated risk tolerance
- Input from parties outside the delivery team on what costs matter
- Only direct financial and remediation cost considered
- Costs listed with no reference to tolerance, so no threshold is crossed or not crossed
- Analysis performed once, before the deployment scope widened
Targeted application scope is specified and documented based on the system’s capability, established context, and AI system categorization. Scope is bounded in writing and justified by capability and context, which is what keeps the risk mapping and the evaluation resources tractable.
- The documented targeted application scope and its boundaries
- The capability and context basis for the scope chosen
- Controls that keep use inside the stated scope
- Records of any scope extension and the re-assessment that accompanied it
- Scope broad enough that no use is outside it
- Boundary documented with nothing enforcing it
- Scope extended in practice without re-assessment
Processes for operator and practitioner proficiency with AI system performance and trustworthiness – and relevant technical standards and certifications – are defined, assessed and documented. Operator proficiency is defined as a requirement, assessed rather than assumed, and recorded, because the human-AI configuration only works if the human can do the job assigned.
- Defined proficiency requirements for operators and practitioners of the system
- Assessment records showing proficiency was tested, not assumed
- Relevant technical standards or certifications identified for the role
- Refresh arrangements when the system or its performance changes
- Proficiency assumed from job title
- Training delivered but proficiency never assessed
- No refresh after a model update changed system behaviour
Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from GOVERN function. Oversight is designed for the specific configuration, assessed for whether it is exercisable in practice, and consistent with the governing policy rather than invented per project.
- The documented human oversight process for the system
- Assessment of whether oversight is exercisable at the pace and volume of operation
- The link to the governing organisational policy on oversight
- The authority the overseer holds, including whether they can stop the system
- Oversight designed at a volume the reviewer cannot sustain
- Overseer able to observe but not to intervene
- Process differs from the governing policy with no recorded deviation
Approaches for mapping AI technology and legal risks of its components – including the use of third-party data or software – are in place, followed, and documented, as are risks of infringement of a third-party’s intellectual property or other rights. There is a followed approach for mapping the technology and legal risk carried by each component, including data and software obtained from third parties and the rights position attached to them.
- The documented approach for mapping component technology and legal risk
- Component inventory identifying third-party data, models and software
- Intellectual property and rights analysis for each third-party component
- Evidence the approach was followed for the components actually in use
- Approach documented but not applied to components adopted since
- Pre-trained models used with no analysis of the provenance of their training data
- Rights reviewed for commercial components only, not for freely obtained ones
Internal risk controls for components of the AI system including third-party AI technologies are identified and documented. For each component carrying risk, the internal control applied to it is identified and written down, so the risk mapping produces controls rather than a list.
- The internal controls identified for each AI system component
- Controls specific to third-party and open-source AI technologies
- The pre-adoption evaluation practice for third-party material
- Evidence the controls named are actually in place
- Risks identified for components with no control named against them
- Freely available third-party material adopted outside the evaluation practice
- Controls documented centrally but absent in the deployed pipeline
Likelihood and magnitude of each identified impact (both potentially beneficial and harmful) based on expected use, past uses of AI systems in similar contexts, public incident reports, feedback from those external to the team that developed or deployed the AI system, or other data are identified and documented. Each identified impact carries a likelihood and a magnitude derived from stated evidence, including comparable past deployments and public incident reports, not from unsupported estimation.
- Impact register with likelihood and magnitude recorded per impact
- The evidence base cited for each estimate, including comparable systems and incident reports
- Beneficial as well as harmful impacts characterised
- The use of these estimates in a go or no-go decision
- Likelihood assigned by consensus in a workshop with no evidence cited
- Only harmful impacts characterised, so trade-offs cannot be weighed
- Estimates produced and never used in any decision
Practices and personnel for supporting regular engagement with relevant AI actors and integrating feedback about positive, negative, and unanticipated impacts are in place and documented. Engagement with affected actors is a resourced standing practice with named personnel, and unanticipated impacts have a route back into the impact record.
- The documented engagement practice and its cadence
- Personnel assigned to conduct and integrate engagement
- Records of impacts reported through engagement, including unanticipated ones
- Evidence reported impacts were integrated into the impact record
- Engagement run as a one-off consultation at launch
- No named personnel, so engagement depends on individual initiative
- Unanticipated impacts reported with no route into the risk record
MEASURE - NIST AI RMF 1.0
Approaches and metrics for measurement of AI risks enumerated during the Map function are selected for implementation starting with the most significant AI risks. The risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented. Measurement approaches are selected against the risks that mapping produced, most significant first, and anything left unmeasured is named as unmeasured rather than silently omitted.
- The selected measurement approaches and metrics, traced to mapped risks
- The prioritisation showing the most significant risks addressed first
- An explicit record of risks and characteristics that will not or cannot be measured, with reasons
- Evidence the selection was implemented rather than only planned
- Metrics chosen from what the tooling already emits rather than from the mapped risks
- Unmeasured characteristics simply absent, so their absence reads as a pass
- Selection documented with no evidence of implementation
Appropriateness of AI metrics and effectiveness of existing controls is regularly assessed and updated including reports of errors and impacts on affected communities. The metrics themselves are re-examined on a cycle for whether they remain appropriate, informed by reported errors and by impacts on affected communities.
- Records of periodic assessment of metric appropriateness and control effectiveness
- Error reports and community impact reports considered in that assessment
- Changes made to metrics or controls as a result
- The trigger conditions, such as drift or changed operating setting, that force a re-assessment
- Metrics fixed at launch and carried unchanged through model updates
- Assessment considers internal error rates only, not reported impacts
- Re-assessment scheduled but no completed record for the current period
Internal experts who did not serve as front-line developers for the system and/or independent assessors are involved in regular assessments and updates. Domain experts, users, AI actors external to the team that developed or deployed the AI system, and affected communities are consulted in support of assessments as necessary per organizational risk tolerance. Assessment includes people who did not build the system, with the depth of external consultation set by the organisation's risk tolerance.
- Assessment records naming assessors and their independence from development
- Records of consultation with domain experts, users and affected communities
- The link from risk tolerance to the level of independence required
- Findings raised by independent assessors and their disposition
- Independent review performed by another team under the same delivery owner
- Consultation limited to internal users
- Independent findings recorded but closed without action
Test sets, metrics, and details about the tools used during test, evaluation, validation, and verification (TEVV) are documented. The TEVV record is complete enough for the evaluation to be repeated: which test sets, which metrics, which tools and which versions.
- Documentation of the test sets used, including provenance and composition
- The metrics computed and their definitions
- Tools and versions used during evaluation
- Sufficient detail for the evaluation to be repeated by someone else
- Results reported with the test set identified only by filename
- Metric named but not defined, so results are not comparable across runs
- Tooling versions unrecorded, so a result cannot be reproduced
Privacy risk of the AI system – as identified in the MAP function – is examined and documented. Privacy examination covers what the AI system makes possible, inference and re-identification from training data and outputs, not only the lawfulness of the input data.
- Privacy risk examination covering training data, inference and outputs
- Assessment of re-identification and memorisation risk
- Privacy-enhancing measures applied and their assessed effect
- Documentation traced to the privacy risks mapping identified
- Privacy assessed as lawful basis for input data only
- Memorisation and training data extraction not considered
- Assessment completed for the original data set and not repeated after retraining
Fairness and bias – as identified in the MAP function – is evaluated and results are documented. Fairness evaluation states which fairness definition was applied and why, evaluates against it, and records the results including where the definition itself is contested.
- The fairness definition or definitions applied and the reason for choosing them
- Evaluation results disaggregated across the relevant groups
- Identification of the groups assessed and the basis for that selection
- Documentation of trade-offs where fairness definitions conflict
- A single fairness metric applied with no argument that it fits the context
- Groups assessed limited to those for which attribute data happened to exist
- Disparity measured and documented with no decision recorded about it
Environmental impact and sustainability of AI model training and management activities – as identified in the MAP function – are assessed and documented. The environmental cost of training and operating the system is assessed on stated measures, energy, water and greenhouse gas emissions, rather than treated as out of scope.
- Assessment of energy consumption for training and inference
- Water consumption and greenhouse gas emissions attributable to the system where applicable
- The measurement basis and any published metrics adopted
- Documentation of the assessment and its bearing on design decisions
- Environmental impact declared immaterial with no measurement
- Training cost assessed while ongoing inference cost is ignored
- Assessment held by the infrastructure team with no link to the AI system record
Effectiveness of the employed TEVV metrics and processes in the MEASURE function are evaluated and documented. The evaluation apparatus is itself evaluated, for whether the metrics still discriminate, whether they are being optimised against, and whether they carry unexamined assumptions.
- Evaluation of whether the TEVV metrics remain effective and discriminating
- Consideration of gaming, saturation and drift in the metrics themselves
- Review of assumptions embedded in the measurement approach
- Changes made to TEVV processes as a result
- Metrics never questioned once adopted
- Metric saturation read as system improvement
- Effectiveness review performed by the team whose work the metrics judge
Evaluations involving human subjects meet applicable requirements (including human subject protection) and are representative of the relevant population. Where evaluation involves human subjects or data captured from them, the applicable protection requirements are met and the subject population is representative of the deployment population.
- Identification of evaluations involving human subjects or human subject data
- Evidence the applicable human subject protection requirements were met, including any approval obtained
- Analysis of representativeness against the relevant population
- Consent and data handling records for subject data used in evaluation
- Human subject involvement not recognised because the data was already held
- Convenience sample used with no representativeness analysis
- Protection requirements treated as applying only to funded research
AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s). Measures are documented. Performance is demonstrated under conditions resembling deployment rather than only on held-out test data, and the measures are recorded.
- The performance or assurance criteria set for the system
- Measurement results obtained under conditions similar to the deployment setting
- The argument that the test conditions resemble deployment
- Documentation of the measures used, qualitative as well as quantitative
- Performance demonstrated only in silico on a static test set
- Deployment conditions asserted to be similar with no analysis
- Qualitative assurance criteria stated but never measured
The functionality and behavior of the AI system and its components – as identified in the MAP function – are monitored when in production. Production monitoring covers the functionality and behaviour that mapping identified as consequential, so drift away from the design assumptions is detected while running.
- The production monitoring configuration and what it observes
- Traceability from monitored signals to the behaviours mapping identified
- Drift detection results and the thresholds applied
- Alerting and the named recipients of monitoring output
- Monitoring covers infrastructure availability rather than model behaviour
- Drift thresholds set but nothing acts when they are crossed
- Component behaviour unmonitored where the component is third-party
The AI system to be deployed is demonstrated to be valid and reliable. Limitations of the generalizability beyond the conditions under which the technology was developed are documented. Validity and reliability are demonstrated before deployment, and the boundary beyond which the demonstration does not carry is written down.
- Validation results demonstrating validity and reliability prior to deployment
- Documented limitations on generalisability beyond the development conditions
- The conditions under which the system was developed and tested
- Evidence that validation failure would prevent deployment
- Validity claimed from training performance rather than independent validation
- Generalisability limits known informally and not documented
- Validation performed after deployment as a formality
AI system is evaluated regularly for safety risks – as identified in the MAP function. The AI system to be deployed is demonstrated to be safe, its residual negative risk does not exceed the risk tolerance, and can fail safely, particularly if made to operate beyond its knowledge limits. Safety metrics implicate system reliability and robustness, real-time monitoring, and response times for AI system failures. Safety is evaluated on a recurring basis against the mapped safety risks, residual risk is compared to the stated tolerance, and behaviour beyond the knowledge limits is tested for safe failure.
- Safety evaluation results traced to the safety risks mapping identified
- The residual safety risk recorded and compared against the stated tolerance
- Test results for behaviour beyond the system's knowledge limits
- Safety metrics covering reliability, robustness, real-time monitoring and failure response time
- Safety evaluated once before launch and not repeated
- Residual risk documented but never compared to a tolerance
- Out-of-distribution behaviour untested, so fail-safe behaviour is unverified
AI system security and resilience – as identified in the MAP function – are evaluated and documented. Security and resilience evaluation addresses the AI-specific attack surface, adversarial input, poisoning, extraction and supply chain, and records whether the system degrades gracefully.
- Evaluation results for AI-specific security risks, including adversarial input and data poisoning
- Resilience testing showing behaviour under unexpected adverse events or environment change
- Model and data supply chain integrity checks
- Documentation of the evaluation traced to the security risks mapping identified
- Conventional application security testing substituted for AI-specific evaluation
- Resilience assessed as infrastructure failover only
- Model extraction and inversion risks not considered
Risks associated with transparency and accountability – as identified in the MAP function – are examined and documented. The examination is of transparency and accountability as risks: what information asymmetry remains between the organisation and those affected, and who can be held to account for an outcome.
- Examination of what information about the system is and is not disclosed, and to whom
- Identification of the accountable party for system outcomes
- Assessment of residual information asymmetry with operators and affected communities
- Documentation traced to the transparency risks mapping identified
- Transparency treated as equivalent to publishing model documentation
- Accountability assigned to the system rather than to a person or function
- Disclosure assessed for regulators only, not for affected individuals
The AI model is explained, validated, and documented, and AI system output is interpreted within its context – as identified in the MAP function – and to inform responsible use and governance. Explanations are validated for fidelity to actual model behaviour and are pitched at the audience that must act on them, and the interpretation of output is bound to the mapped context.
- The explanation method used and validation of its fidelity to model behaviour
- Model documentation covering how the model reaches its outputs
- Guidance on interpreting output within the deployment context
- Identification of the audiences for explanation and what each needs
- Explanation method adopted with no check that it reflects the model
- Feature attributions produced for developers and never translated for the people affected
- Limitations of the explanation method not stated
Approaches, personnel, and documentation are in place to regularly identify and track existing, unanticipated, and emergent AI risks based on factors such as intended and actual performance in deployed contexts. Risk identification continues after deployment with assigned personnel and a tracking record, so risks that emerge in real use are captured rather than only those anticipated at design.
- The approach for identifying emergent and unanticipated risks in deployment
- Personnel assigned to that identification and tracking
- The tracking record showing risks identified after deployment
- Comparison of actual against intended performance in the deployed context
- Risk identification treated as a design-phase activity that ends at launch
- Emergent risks noted in incident tickets but never entered as risks
- No one assigned, so identification depends on something going visibly wrong
Risk tracking approaches are considered for settings where AI risks are difficult to assess using currently available measurement techniques or where metrics are not yet available. Where no adequate metric exists the risk is still tracked by some stated means, so difficulty of measurement does not become silent omission.
- Identification of risks for which adequate measurement techniques do not exist
- The tracking approach adopted for each such risk
- Any novel or qualitative measurement approaches trialled
- Review of whether measurement has since become possible
- Hard-to-measure risks dropped from the register rather than tracked qualitatively
- Approach considered once and not revisited as techniques matured
- Absence of a metric reported as absence of risk
Feedback processes for end users and impacted communities to report problems and appeal system outcomes are established and integrated into AI system evaluation metrics. End users and impacted communities have a working route to report problems and appeal an outcome, and what arrives through it feeds the evaluation metrics rather than a separate queue.
- The reporting and appeal route available to end users and impacted communities
- Records of problems reported and appeals lodged, with outcomes
- The integration of that feedback into evaluation metrics
- Evidence the route is discoverable by the people expected to use it
- Appeal route exists in policy but is not reachable from the point of the decision
- Reports handled as customer service tickets with no route into evaluation
- No appeal available where the decision is automated and consequential
Measurement approaches for identifying AI risks are connected to deployment context(s) and informed through consultation with domain experts and other end users. Approaches are documented. The measurement design is informed by people who understand the deployment context, because the risks that matter there are often not visible to those running the evaluation.
- Documentation of the measurement approaches and their connection to deployment contexts
- Records of consultation with domain experts and end users on measurement design
- Evidence that consultation changed the measurement approach
- Identification of the deployment contexts the measurement is meant to cover
- Measurement designed entirely by the evaluation team
- Consultation held after metrics were fixed
- One measurement approach applied across materially different deployment contexts
Measurement results regarding AI system trustworthiness in deployment context(s) and across AI lifecycle are informed by input from domain experts and other relevant AI actors to validate whether the system is performing consistently as intended. Results are documented. Measured results are validated against the judgement of people who know the context, so a result inside its operational limits on paper is confirmed as adequate in practice.
- Measurement results with domain expert and AI actor input recorded against them
- The pre-defined operational limits the results are judged against
- Documentation of whether the system is performing consistently as intended
- Disposition of any expert view that conflicted with the measured result
- Results published with no expert validation of what they mean in context
- Operational limits set after the results were known
- Conflicting expert judgement recorded and then disregarded without reasoning
Measurable performance improvements or declines based on consultations with relevant AI actors including affected communities, and field data about context-relevant risks and trustworthiness characteristics, are identified and documented. Change in performance over time is measured against a baseline using field data and consultation, so decline is detected as decline rather than absorbed as normal variation.
- Baseline measures for the trustworthiness characteristics being tracked
- Field data showing performance over time against that baseline
- Consultation records with AI actors and affected communities on observed change
- Documented identification of improvement or decline and the action taken
- No baseline, so change cannot be identified
- Field data collected but never compared across periods
- Decline attributed to data quality without investigation
Assembled from the framework’s own control set, so this list is regenerated rather than written and stays current as the graph does.