Skip to content

Evidence request lists

NIST AI Risk Management Framework (AI RMF 1.0)

Evidence request list. 72 controls, 72 carrying auditor artefact guidance. Generated from the compliance knowledge graph on 12 September 2026. Published by The Art of Service.

GOVERN - NIST AI RMF 1.0

AIRMF-GV-1.1
Legal and regulatory requirements involving AI are understood, managed, and documented

Legal and regulatory requirements involving AI are understood, managed, and documented. The organisation identifies which AI-specific and AI-adjacent legal duties bind each system, records how each is discharged, and keeps that record current as the applicable law changes across the jurisdictions the system reaches.

Artefacts an auditor will ask for
  • A register of legal and regulatory requirements applicable to each AI system, naming the instrument and the obligation
  • Mapping from each identified obligation to the control, process or artefact that discharges it
  • Evidence of legal review at design and before deployment, with reviewer and date
  • A record of jurisdictions the system is deployed into and the requirements that follow from each
  • Change records showing the register was updated when applicable law changed
Where this commonly fails
  • A generic corporate compliance register with no AI-specific entries
  • Requirements listed but never mapped to a control, so nothing shows they are met
  • Register maintained for the launch jurisdiction only while the system is reachable elsewhere
AIRMF-GV-1.2
The characteristics of trustworthy AI are integrated into organizational policies, processes, and procedures

The characteristics of trustworthy AI are integrated into organizational policies, processes, and procedures. Validity, reliability, safety, security, resilience, accountability, transparency, explainability, privacy and fairness appear as named obligations inside the policies that govern the AI lifecycle, not only in a statement of values.

Artefacts an auditor will ask for
  • Approved AI policy set naming each trustworthiness characteristic as an obligation with an owner
  • Procedure documents showing where in the lifecycle each characteristic is assessed
  • Approval record with authority and effective date, and the scheduled review interval
  • Evidence the policies were communicated to the teams that design, develop and deploy AI
Where this commonly fails
  • Trustworthiness stated as principles in a values page with no procedural consequence
  • Policy covers security and privacy only, leaving fairness and explainability unaddressed
  • No review since first issue, so the policy predates the systems it is meant to govern
AIRMF-GV-1.3
Processes and procedures are in place to determine the needed level of risk management activities based on the organization's risk tolerance

Processes and procedures are in place to determine the needed level of risk management activities based on the organization's risk tolerance. Risk management effort is allocated by a documented rule rather than by availability, so a system assessed as higher risk demonstrably receives more scrutiny than a lower-risk one.

Artefacts an auditor will ask for
  • The documented risk tiering criteria and the scale risks are assigned to
  • Tier assignment records for the AI systems in the inventory, with the reasoning
  • The activity set required at each tier, showing what changes between tiers
  • Evidence that a high-tier system actually received the additional activities its tier requires
Where this commonly fails
  • Risk tolerance expressed only as a word such as low or moderate, with no scale behind it
  • Every system assigned the same tier, which makes the tiering inert
  • Tier assigned at intake and never revisited after the system's use expanded
AIRMF-GV-1.4
The risk management process and its outcomes are established through transparent policies, procedures, and other controls based on organizational risk priorities

The risk management process and its outcomes are established through transparent policies, procedures, and other controls based on organizational risk priorities. The process and the decisions it produced are both documented in a form a reviewer outside the delivery team can follow, including who decided and on what basis.

Artefacts an auditor will ask for
  • The documented AI risk management process, with the decision points it contains
  • Standardised documentation templates used across AI work products
  • Completed records for deployed systems showing the process was followed, not just defined
  • Named accountable contacts recorded on work products so decisions are traceable to people
Where this commonly fails
  • Process defined centrally but the delivery records show a different practice
  • Outcomes recorded without the reasoning, so the decision cannot be reviewed
  • Documentation held in individual team tools with no organisational retention
AIRMF-GV-1.5
Ongoing monitoring and periodic review of the risk management process and its outcomes are planned, organizational roles and responsibilities are clearly defined, including determining the frequency of periodic review

Ongoing monitoring and periodic review of the risk management process and its outcomes are planned, organizational roles and responsibilities are clearly defined, including determining the frequency of periodic review. Monitoring and review are scheduled with a stated frequency and a named owner, and the review actually produces recorded findings rather than a confirmation that nothing changed.

Artefacts an auditor will ask for
  • The monitoring plan naming what is monitored, by whom and at what frequency
  • Review schedule with the basis for the chosen interval
  • Completed review records with findings, decisions and follow-up actions
  • Role definitions assigning monitoring and review duties to named functions
Where this commonly fails
  • Review frequency stated but no completed review on file for the current period
  • Reviews recorded as no change with nothing examined to support that conclusion
  • Monitoring owned by the team that built the system, with no independent line
AIRMF-GV-1.6
Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities

Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities. A maintained inventory identifies the AI systems and models in use, and the resource given to maintaining it is proportionate to the risk the inventoried systems carry.

Artefacts an auditor will ask for
  • The AI system and model inventory with the fields it captures per entry
  • The intake process that causes a new system or model to be added
  • Evidence of periodic reconciliation between the inventory and systems actually in production
  • Named owner and resourcing for inventory maintenance
Where this commonly fails
  • Inventory covers models built in house but omits embedded and third-party AI features
  • Compiled once for an audit and not maintained since
  • No link from an inventory entry to its documentation, so the entry carries no usable detail
AIRMF-GV-1.7
Processes and procedures are in place for decommissioning and phasing out of AI systems safely and in a manner that does not increase risks or decrease the organization's trustworthiness

Processes and procedures are in place for decommissioning and phasing out of AI systems safely and in a manner that does not increase risks or decrease the organization’s trustworthiness. Retirement is a governed event with retention, user notification and downstream dependency steps, rather than deletion at the point the system stops being useful.

Artefacts an auditor will ask for
  • The documented decommissioning procedure, including retention and legal hold steps
  • User and downstream consumer notification records for retired systems
  • Dependency analysis showing what consumed the system before it was withdrawn
  • Completed decommissioning records for systems actually retired
Where this commonly fails
  • Procedure exists but retirements are executed as an infrastructure ticket outside it
  • Model artefacts and training data deleted while a regulatory retention duty still applied
  • Downstream consumers discovered the retirement from a failure rather than a notice
AIRMF-GV-2.1
Roles and responsibilities and lines of communication related to mapping, measuring, and managing AI risks are documented and are clear to individuals and teams throughout the organization

Roles and responsibilities and lines of communication related to mapping, measuring, and managing AI risks are documented and are clear to individuals and teams throughout the organization. Duties for mapping, measuring and managing AI risk are assigned to identifiable roles with escalation paths, and the people holding them can state what they are accountable for.

Artefacts an auditor will ask for
  • A responsibility assignment covering the map, measure and manage duties
  • Organisational chart or charter showing the reporting line for AI risk functions
  • Escalation paths with named recipients and thresholds
  • Evidence the assignment was communicated and acknowledged by role holders
Where this commonly fails
  • Responsibility assigned to a committee rather than to a person who can be asked
  • Test and evaluation reporting into the same manager who owns the delivery deadline
  • Assignment current at the last reorganisation rather than at the present one
AIRMF-GV-2.2
The organization's personnel and partners receive AI risk management training to enable them to perform their duties and responsibilities consistent with related policies, procedures, and agreements

The organization’s personnel and partners receive AI risk management training to enable them to perform their duties and responsibilities consistent with related policies, procedures, and agreements. Training is role-specific and its completion is recorded, and partners bound by agreements are covered as well as employees.

Artefacts an auditor will ask for
  • Training curricula differentiated by AI role, with the content covered
  • Completion records by individual and date, including partner personnel where in scope
  • The link from a role definition to the training required for it
  • Evidence of refresh when policies or applicable law changed
Where this commonly fails
  • One general awareness module issued to everyone regardless of AI duty
  • Partner and contractor personnel excluded despite performing AI lifecycle tasks
  • Completion tracked but no record of what the training actually covered
AIRMF-GV-2.3
Executive leadership of the organization takes responsibility for decisions about risks associated with AI system development and deployment

Executive leadership of the organization takes responsibility for decisions about risks associated with AI system development and deployment. Named executives hold and exercise the decision right over AI risk acceptance, with their decisions minuted, rather than delegating it into the delivery line.

Artefacts an auditor will ask for
  • The charter or terms of reference assigning AI risk decisions to named executives
  • Minutes recording executive decisions to accept, mitigate or reject AI risks
  • The stated organisational appetite for AI risk, approved at executive level
  • Evidence of the authority and budget granted to the accountable officer
Where this commonly fails
  • Executive sponsorship claimed with no minuted decision to point to
  • Risk acceptance signed by the project owner who benefits from proceeding
  • Appetite endorsed once at programme start and never revisited as systems changed
AIRMF-GV-3.1
Decision-making related to mapping, measuring, and managing AI risks throughout the lifecycle is informed by a diverse team

Decision-makings related to mapping, measuring, and managing AI risks throughout the lifecycle is informed by a diverse team (e.g., diversity of demographics, disciplines, experience, expertise, and backgrounds). The people making AI risk decisions span disciplines and backgrounds, and where internal composition is narrow, external input is obtained and recorded.

Artefacts an auditor will ask for
  • Records of the disciplines represented on AI risk decision forums
  • Documented consultation with external perspectives where internal expertise was thin
  • The organisational commitment on composition of AI risk teams
  • Evidence that a dissenting or external view changed a decision
Where this commonly fails
  • Diversity asserted from headcount data with no bearing on who decides
  • External consultation held after the design was fixed, so it could not affect it
  • Same three technical staff constitute every review
AIRMF-GV-3.2
Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems

Policies and procedures are in place to define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems. The policy distinguishes the human roles around a system, operator, reviewer, overseer, affected party, and states what each may and must do about its output.

Artefacts an auditor will ask for
  • Policy defining each human role in the human-AI configuration and its authority
  • Per-system documentation of the configuration chosen and why
  • Oversight procedures stating when a human may override or must intervene
  • Competency requirements attached to each oversight role
Where this commonly fails
  • Human in the loop asserted without stating what the human is empowered to change
  • Oversight assigned to a role with no authority to stop the system
  • Configuration documented at design and not updated after automation increased
AIRMF-GV-4.1
Organizational policies and practices are in place to foster a critical thinking and safety-first mindset in the design, development, deployment, and uses of AI systems to minimize negative impacts

Organizational policies, and practices are in place to foster a critical thinking and safety-first mindset in the design, development, deployment, and uses of AI systems to minimize negative impacts. Practices exist that make raising a concern normal and consequential, such as separated accountability lines or a recorded route for challenging a deployment decision.

Artefacts an auditor will ask for
  • Policy establishing separated accountability for development, risk and assurance
  • The internal route for raising an AI safety concern and the protection attached to it
  • Records of concerns raised and how each was resolved
  • Evidence a concern changed a design or delayed a deployment
Where this commonly fails
  • A stated speak-up culture with no record of anything ever being raised
  • Assurance and development reporting to the same accountable owner
  • Concerns logged and closed without a recorded resolution
AIRMF-GV-4.2
Organizational teams document the risks and potential impacts of the AI technology they design, develop, deploy, evaluate and use, and communicate about the impacts more broadly

Organizational teams document the risks and potential impacts of the AI technology they design, develop, deploy, evaluate and use, and communicate about the impacts more broadly. Impact documentation is produced by the team that builds the system and is communicated beyond it, so risk knowledge does not stay inside the delivery group.

Artefacts an auditor will ask for
  • Completed impact assessments for AI systems, with author and date
  • The template or structure the assessments follow
  • Evidence the assessment was communicated outside the delivery team
  • The link from an assessment finding to a risk register entry or design change
Where this commonly fails
  • Assessment completed as a launch gate and never revisited
  • Findings recorded with no route into the organisational risk register
  • Only harms to the organisation considered, not harms to individuals or communities
AIRMF-GV-4.3
Organizational practices are in place to enable AI testing, identification of incidents, and information sharing

Organizational practices are in place to enable AI testing, identification of incidents, and information sharing. Testing and incident identification are enabled by practice and resourcing, and information about failures is shared with the parties who need it rather than contained.

Artefacts an auditor will ask for
  • In-house testing policy covering the failure modes AI introduces
  • The incident identification and recording mechanism for AI systems
  • Records of information shared internally or externally about identified issues
  • Evidence of resourcing for testing independent of the delivery schedule
Where this commonly fails
  • Testing limited to accuracy on a held-out set, missing drift and shortcut learning
  • Incidents recorded in a delivery tracker that no risk function reads
  • No sharing beyond the team, so the same failure recurs elsewhere
AIRMF-GV-5.1
Organizational policies and practices are in place to collect, consider, prioritize, and integrate feedback from those external to the team that developed or deployed the AI system regarding the potential individual and societal impacts related to AI risks

Organizational policies and practices are in place to collect, consider, prioritize, and integrate feedback from those external to the team that developed or deployed the AI system regarding the potential individual and societal impacts related to AI risks. External feedback is solicited under a policy, prioritised on a stated basis and carried into the system, rather than collected and filed.

Artefacts an auditor will ask for
  • Policy establishing how external feedback on AI impacts is collected
  • Records of stakeholder engagement activity, including who was engaged
  • The prioritisation basis applied to feedback received
  • Traceability from a specific piece of feedback to a change made
Where this commonly fails
  • Feedback channel published with no process behind it
  • Engagement limited to existing customers, missing affected non-users
  • Feedback logged with no evidence any of it was acted on
AIRMF-GV-5.2
Mechanisms are established to enable AI actors to regularly incorporate adjudicated feedback from relevant AI actors into system design and implementation

Mechanisms are established to enable AI actors to regularly incorporate adjudicated feedback from relevant AI actors into system design and implementation. There is a mechanism that adjudicates feedback and routes the accepted items into design and implementation on a regular cycle.

Artefacts an auditor will ask for
  • The adjudication mechanism and who holds the decision on what is accepted
  • The cadence at which adjudicated feedback enters design or implementation
  • Records of feedback adjudicated, with accept or reject reasoning
  • Change records showing accepted feedback implemented
Where this commonly fails
  • Feedback adjudicated informally, so rejections carry no reasoning
  • Accepted items queued indefinitely with no implementation cycle
  • Mechanism covers internal actors only and excludes deployers and end users
AIRMF-GV-6.1
Policies and procedures are in place that address AI risks associated with third-party entities, including risks of infringement of a third party's intellectual property or other rights

Policies and procedures are in place that address AI risks associated with third-party entities, including risks of infringement of a third party’s intellectual property or other rights. Third-party AI risk is addressed by policy covering data, models, software and services obtained externally, including the rights position on training data and model outputs.

Artefacts an auditor will ask for
  • Third-party AI policy covering data, pre-trained models, software and services
  • Due diligence records for third-party AI components in use
  • Contract terms addressing intellectual property, data rights and liability for AI components
  • The intellectual property position recorded for training data and model outputs
Where this commonly fails
  • Standard vendor due diligence applied with no AI-specific questions
  • Open source models adopted with no review of the licence or the training data provenance
  • Policy addresses suppliers but not freely obtained models and datasets
AIRMF-GV-6.2
Contingency processes are in place to handle failures or incidents in third-party data or AI systems deemed to be high-risk

Contingency processes are in place to handle failures or incidents in third-party data or AI systems deemed to be high-risk. For third-party components assessed as high risk there is a worked contingency, a fallback, a redundancy or a defined manual path, and it has been tested.

Artefacts an auditor will ask for
  • Identification of third-party AI components assessed as high risk
  • The contingency defined for each, naming the fallback or redundancy
  • Test or exercise records demonstrating the contingency works
  • Trigger criteria and the named role that invokes the contingency
Where this commonly fails
  • Contingency documented as revert to manual with no assessment of whether that is possible
  • Never exercised, so the fallback capacity is unverified
  • Applies to the primary vendor only while a sub-processor carries the same dependency

MANAGE - NIST AI RMF 1.0

AIRMF-MN-1.1
A determination is made as to whether the AI system achieves its intended purpose and stated objectives and whether its development or deployment should proceed

A determination is made as to whether the AI system achieves its intended purpose and stated objectives and whether its development or deployment should proceed. A recorded determination weighs measured risk against benefit and decides whether to proceed, so that not proceeding remains an available and evidenced outcome.

Artefacts an auditor will ask for
  • The recorded determination on whether the system achieves its intended purpose
  • The measured risk and benefit evidence the determination rested on
  • The decision to proceed or not, with the deciding authority named
  • Consideration of whether an AI system is the appropriate solution at all
Where this commonly fails
  • Determination recorded as approval with no evidence weighed
  • No pathway by which the answer could have been not to proceed
  • Decision taken by the delivery owner rather than by an authority able to stop it
AIRMF-MN-1.2
Treatment of documented AI risks is prioritized based on impact, likelihood, or available resources or methods

Treatment of documented AI risks is prioritized based on impact, likelihood, or available resources or methods. Treatment order follows a stated prioritisation basis rather than convenience, and the basis is visible in the treatment record.

Artefacts an auditor will ask for
  • The risk treatment register showing priority assigned per risk
  • The prioritisation basis applied, whether impact, likelihood, resource or method availability
  • Evidence that higher-priority risks were treated first
  • The link from prioritisation to the organisational risk tolerance
Where this commonly fails
  • Priority assigned but treatment order does not follow it
  • Prioritisation basis unstated, so the ordering cannot be reviewed
  • Low-cost treatments completed first regardless of the risk they address
AIRMF-MN-1.3
Responses to the AI risks deemed high priority as identified by the MAP function are developed, planned, and documented, and risk response options can include mitigating, transferring, avoiding, or accepting

Responses to the AI risks deemed high priority as identified by the Map function, are developed, planned, and documented. Risk response options can include mitigating, transferring, avoiding, or accepting. High-priority risks each carry a planned and documented response naming which of the four options was chosen, with acceptance recorded as a decision rather than as inaction.

Artefacts an auditor will ask for
  • Documented responses for each high-priority risk, naming the option chosen
  • Response plans with owner, action and target date
  • Acceptance decisions recorded with the accepting authority
  • Traceability from the mapped risk to its response
Where this commonly fails
  • Acceptance by default because no response was ever planned
  • Response option not named, so mitigation and acceptance are indistinguishable
  • Plans without owners or dates, so completion cannot be judged
AIRMF-MN-1.4
Negative residual risks, defined as the sum of all unmitigated risks, to both downstream acquirers of AI systems and end users are documented

Negative residual risks (defined as the sum of all unmitigated risks) to both downstream acquirers of AI systems and end users are documented. What remains unmitigated is totalled and documented, and communicated to downstream acquirers and end users who inherit it.

Artefacts an auditor will ask for
  • The documented residual risk position after treatment
  • Identification of downstream acquirers and end users who bear residual risk
  • Evidence residual risk was communicated to those parties
  • The authority that accepted the residual position
Where this commonly fails
  • Residual risk computed internally and never disclosed downstream
  • Individual residual risks listed with no aggregate position
  • Residual position stale because treatments changed after it was recorded
AIRMF-MN-2.1
Resources required to manage AI risks are taken into account, along with viable non-AI alternative systems, approaches, or methods, to reduce the magnitude or likelihood of potential impacts

Resources required to manage AI risks are taken into account, along with viable non-AI alternative systems, approaches, or methods – to reduce the magnitude or likelihood of potential impacts. The resource cost of managing the risk is accounted for and non-AI alternatives are genuinely analysed as a way of reducing impact, not noted and dismissed.

Artefacts an auditor will ask for
  • The resources required to manage identified AI risks, accounted for and allocated
  • Analysis of viable non-AI alternative systems, approaches or methods
  • The trade-offs weighed between trustworthiness characteristics and the alternatives
  • Interdisciplinary input recorded in the alternatives analysis
Where this commonly fails
  • Non-AI alternatives named in a sentence with no analysis
  • Risk management resourced from slack capacity rather than allocated
  • Trade-offs decided implicitly by the technical team
AIRMF-MN-2.2
Mechanisms are in place and applied to sustain the value of deployed AI systems

Mechanisms are in place and applied to sustain the value of deployed AI systems. There are applied mechanisms, retraining, recalibration or refresh, that maintain the system's value against drift, and evidence they are used rather than merely available.

Artefacts an auditor will ask for
  • The mechanisms in place to sustain value, such as retraining or recalibration
  • Records showing those mechanisms were applied, with dates
  • The trigger conditions that cause them to be applied
  • Evidence of the effect on performance after application
Where this commonly fails
  • Retraining capability exists but has not been exercised since deployment
  • Triggers defined against calendar time rather than observed degradation
  • Effect of retraining unmeasured, so sustainment is assumed
AIRMF-MN-2.3
Procedures are followed to respond to and recover from a previously unknown risk when it is identified

Procedures are followed to respond to and recover from a previously unknown risk when it is identified. A followed procedure exists for risks that were not anticipated, covering response and recovery, and it has been used or exercised.

Artefacts an auditor will ask for
  • The documented procedure for responding to and recovering from previously unknown risks
  • Records of the procedure being followed for an actual or exercised event
  • The route by which a previously unknown risk is escalated
  • Post-event review and the changes it produced
Where this commonly fails
  • Procedure written for known incident types only
  • Never exercised, so recovery capability is untested
  • Escalation route defined but not known to the operators who would use it
AIRMF-MN-2.4
Mechanisms are in place and applied, and responsibilities are assigned and understood, to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use

Mechanisms are in place and applied, responsibilities are assigned and understood to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use. The ability to bypass, disengage or deactivate the system exists technically, is assigned to an understood role, and has been demonstrated to work.

Artefacts an auditor will ask for
  • The technical mechanism to supersede, disengage or deactivate the system
  • The role assigned to invoke it and evidence that role understands the duty
  • The criteria that trigger disengagement
  • Test or exercise evidence that disengagement works and the fallback holds
Where this commonly fails
  • Kill switch assumed to exist but never tested
  • Deactivation requires a change release, so it cannot be invoked at the speed the risk demands
  • Authority to disengage held by someone with a competing delivery incentive
AIRMF-MN-3.1
AI risks and benefits from third-party resources are regularly monitored, and risk controls are applied and documented

AI risks and benefits from third-party resources are regularly monitored, and risk controls are applied and documented. Third-party data, model, software and hardware dependencies are monitored on an ongoing basis, not assessed once at procurement, and the controls applied are recorded.

Artefacts an auditor will ask for
  • The inventory of third-party resources the AI system depends on
  • Monitoring records showing regular review of those dependencies
  • The risk controls applied to each and evidence they are in place
  • The route by which a third-party change reaches the risk owner
Where this commonly fails
  • Third parties assessed at onboarding and not monitored afterwards
  • Model or API version changes by the provider go unnoticed
  • Controls documented in the contract with no operational evidence
AIRMF-MN-3.2
Pre-trained models which are used for development are monitored as part of AI system regular monitoring and maintenance

Pre-trained models which are used for development are monitored as part of AI system regular monitoring and maintenance. Pre-trained and transfer-learned models are treated as a monitored component in their own right, since their provenance and behaviour are not controlled by the deploying organisation.

Artefacts an auditor will ask for
  • Identification of pre-trained models used, with version and provenance
  • Monitoring records covering those models within regular maintenance
  • Assessment of risks carried over from the pre-training data and objective
  • The procedure followed when the upstream model is updated or withdrawn
Where this commonly fails
  • Pre-trained model treated as a fixed dependency and excluded from monitoring
  • Provenance of pre-training data unknown and unrecorded
  • No procedure for an upstream model being deprecated
AIRMF-MN-4.1
Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management

Post-deployment AI system monitoring plans are implemented, including mechanisms for capturing and evaluating input from users and other relevant AI actors, appeal and override, decommissioning, incident response, recovery, and change management. A single implemented plan covers the post-deployment lifecycle end to end, and each named element, including appeal and override, is actually operating.

Artefacts an auditor will ask for
  • The implemented post-deployment monitoring plan and its scope
  • Mechanisms for capturing and evaluating user and AI actor input
  • The appeal and override mechanism and records of its use
  • Incident response, recovery, change management and decommissioning arrangements for the system
Where this commonly fails
  • Plan documents each element but only performance monitoring is running
  • Appeal and override present in the plan with no implemented mechanism
  • Change management for the model handled outside the plan by the platform team
AIRMF-MN-4.2
Measurable activities for continual improvements are integrated into AI system updates and include regular engagement with interested parties, including relevant AI actors

Measurable activities for continual improvements are integrated into AI system updates and include regular engagement with interested parties, including relevant AI actors. Improvement activities are measurable and attached to the update cycle, and interested parties are engaged as part of that cycle rather than consulted separately.

Artefacts an auditor will ask for
  • The improvement activities integrated into the system update process, with their measures
  • Records of engagement with interested parties within the update cycle
  • Root cause analyses for degradation, drift, near misses and failures
  • Evidence an improvement activity changed a subsequent release
Where this commonly fails
  • Improvement recorded as a backlog with no measure of effect
  • Engagement conducted annually and disconnected from releases
  • Root cause analysis performed for outages only, not for behavioural degradation
AIRMF-MN-4.3
Incidents and errors are communicated to relevant AI actors including affected communities, and processes for tracking, responding to, and recovering from incidents and errors are followed and documented

Incidents and errors are communicated to relevant AI actors including affected communities. Processes for tracking, responding to, and recovering from incidents and errors are followed and documented. Incidents and errors are communicated outward, including to affected communities, and the tracking record shows how each was identified, repaired and how the repair was distributed.

Artefacts an auditor will ask for
  • The incident and error record showing how each was identified
  • Communication records to relevant AI actors and affected communities
  • Evidence of whether the error was repaired and how the repair was distributed to impacted users
  • The followed process for tracking, response and recovery
Where this commonly fails
  • Incidents communicated internally only, never to affected communities
  • Repair applied to the primary deployment while other users keep the defective version
  • Record shows closure with no account of how the error was identified or fixed

MAP - NIST AI RMF 1.0

AIRMF-MP-1.1
Intended purpose, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented

Intended purpose, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented. Considerations include: specific set or types of users along with their expectations; potential positive and negative impacts of system uses to individuals, communities, organizations, society, and the planet; assumptions and related limitations about AI system purposes; uses and risks across the development or product AI lifecycle; TEVV and system metrics. Context is captured before build: who the users are and what they expect, the positive and negative impacts of the intended uses, the assumptions and limits on purpose, and the settings the system will actually run in.

Artefacts an auditor will ask for
  • A documented statement of intended purpose and the uses considered beneficial
  • Identification of user types and their expectations of the system
  • The laws, norms and expectations specific to the deployment setting
  • Recorded assumptions and limitations about purpose, and the uses ruled out
  • Analysis of foreseeable misuse and repurposing of the deployed system
Where this commonly fails
  • Purpose written at product-marketing level, too broad to bound anything
  • Foreseeable misuse omitted because it is outside the intended use
  • Documented once at concept and never reconciled with how the system is actually used
AIRMF-MP-1.2
Inter-disciplinary AI actors, competencies, skills and capacities for establishing context reflect demographic diversity and broad domain and user experience expertise, and their participation is documented

Inter-disciplinary AI actors, competencies, skills and capacities for establishing context reflect demographic diversity and broad domain and user experience expertise, and their participation is documented. Opportunities for interdisciplinary collaboration are prioritized. The people establishing context bring domain and user-experience expertise beyond the build team, and their participation is recorded so the basis of the context statement is auditable.

Artefacts an auditor will ask for
  • Record of the disciplines and expertise represented in context establishment
  • Documented participation of domain and user-experience specialists
  • Evidence that interdisciplinary collaboration was prioritised and resourced
  • The competency requirements set for context establishment work
Where this commonly fails
  • Context established entirely by the engineering team
  • Participation asserted with no record of who took part or what they contributed
  • Domain expertise consulted once rather than through the mapping work
AIRMF-MP-1.3
The organization's mission and relevant goals for the AI technology are understood and documented

The organization’s mission and relevant goals for the AI technology are understood and documented. The system's connection to a stated organisational goal is written down, which is what makes a go or no-go decision reviewable later.

Artefacts an auditor will ask for
  • Documented mission and the specific goals the AI technology serves
  • The linkage from system objectives to those organisational goals
  • The go or no-go decision record referencing that linkage
  • Evidence of review when goals or the system changed
Where this commonly fails
  • Goal recorded as improve efficiency with nothing measurable behind it
  • Alignment written after the build to justify a decision already taken
  • No record of the alternatives weighed
AIRMF-MP-1.4
The business value or context of business use has been clearly defined or, in the case of assessing existing AI systems, re-evaluated

The business value or context of business use has been clearly defined or – in the case of assessing existing AI systems – re-evaluated. Business value is stated in terms that can be checked after deployment, and for systems already running it is re-evaluated rather than assumed from the original case.

Artefacts an auditor will ask for
  • The documented business value or use context for the system
  • Re-evaluation records for AI systems already in operation
  • The measures by which value is judged after deployment
  • Analysis of how the social context of use bears on the value claimed
Where this commonly fails
  • Value defined at business-case stage and never re-tested against outcomes
  • Existing systems inherited with no re-evaluation performed
  • Value stated only as cost saving, ignoring how the system is used and by whom
AIRMF-MP-1.5
Organizational risk tolerances are determined and documented

Organizational risk tolerances are determined and documented. Tolerance is written down at a level that lets someone decide whether a measured risk is acceptable, and it is traceable to an authority that can set it.

Artefacts an auditor will ask for
  • The documented risk tolerance statement and its approval authority
  • The criteria or scale by which a measured risk is judged against tolerance
  • Any sector, regulatory or professional requirements the tolerance is derived from
  • Records of decisions taken by applying the tolerance
Where this commonly fails
  • Tolerance expressed as a word with no scale, so nothing can be judged against it
  • Set by the delivery team rather than by an authority that can accept the consequence
  • Documented but never referenced in any deployment decision
AIRMF-MP-1.6
System requirements are elicited from and understood by relevant AI actors, and design decisions take socio-technical implications into account to address AI risks

System requirements (e.g., “the system shall respect the privacy of its users”) are elicited from and understood by relevant AI actors. Design decisions take socio-technical implications into account to address AI risks. Requirements are elicited from the actors who will be affected and written down, and the design record shows socio-technical implications were weighed rather than only technical ones.

Artefacts an auditor will ask for
  • Written system requirements including non-functional and trustworthiness requirements
  • Record of who requirements were elicited from and how
  • Design decision records showing socio-technical implications considered
  • Traceability from a requirement to the design element that satisfies it
Where this commonly fails
  • Requirements captured as model performance targets only
  • Elicited from the commissioning stakeholder alone, missing operators and affected parties
  • Requirements written but no traceability, so nothing shows they were built
AIRMF-MP-2.1
The specific task, and methods used to implement the task, that the AI system will support is defined

The specific task, and methods used to implement the task, that the AI system will support is defined (e.g., classifiers, generative models, recommenders). The learning or decision task is stated narrowly, together with the method class used to implement it, because a narrow task definition is what makes the risk mapping tractable.

Artefacts an auditor will ask for
  • The documented task the system performs, stated narrowly
  • The method or model class used to implement the task
  • The boundary of what the system does not do
  • Evidence the definition was updated when the task changed
Where this commonly fails
  • Task defined as the product name rather than as a learning or decision task
  • Method recorded as machine learning with no class identified
  • Definition unchanged after a general-purpose model replaced a narrow one
AIRMF-MP-2.2
Information about the AI system's knowledge limits and how system output may be utilized and overseen by humans is documented

Information about the AI system’s knowledge limits and how system output may be utilized and overseen by humans is documented. Documentation provides sufficient information to assist relevant AI actors when making informed decisions and taking subsequent actions. The system's knowledge limits are written down alongside how output should be used and overseen, in enough detail for a downstream actor to make an informed decision.

Artefacts an auditor will ask for
  • Documented knowledge limits, including the conditions the system was not built for
  • Guidance on permitted use and interpretation of system output
  • The human oversight expected over the output, and by whom
  • Evidence this documentation reaches downstream deployers and operators
Where this commonly fails
  • Limits known to the build team and absent from anything a deployer receives
  • Output guidance written for the ideal case, silent on out-of-distribution input
  • Documentation exists but is not delivered with the system
AIRMF-MP-2.3
Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection, system trustworthiness, and construct validation

Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection (e.g., availability, representativeness, suitability), system trustworthiness, and construct validation. The evaluation design is documented as a scientific claim: what was measured, on what data, and whether the measure validly stands for the property claimed.

Artefacts an auditor will ask for
  • Documented experimental design for evaluation of the system
  • Data collection and selection decisions, with availability, representativeness and suitability addressed
  • Construct validation showing the metric stands for the property claimed
  • The TEVV considerations identified and how each is handled
Where this commonly fails
  • Benchmark accuracy reported with no argument that it measures the property claimed
  • Data selection undocumented, so representativeness cannot be assessed
  • Evaluation designed by the same people optimising against it
AIRMF-MP-3.1
Potential benefits of intended AI system functionality and performance are examined and documented

Potential benefits of intended AI system functionality and performance are examined and documented. Claimed benefits are examined and written down rather than assumed, so they can later be weighed against measured costs and residual risk.

Artefacts an auditor will ask for
  • Documented analysis of the benefits the system is expected to deliver
  • The performance basis on which each claimed benefit rests
  • Identification of who receives the benefit
  • Evidence the benefit analysis is revisited against realised outcomes
Where this commonly fails
  • Benefits asserted in a business case with no performance basis
  • Benefit to the organisation documented, benefit or detriment to users not considered
  • Never revisited, so claimed benefits are never confirmed
AIRMF-MP-3.2
Potential costs, including non-monetary costs, which result from expected or realized AI errors or system functionality and trustworthiness are examined and documented, as connected to organizational risk tolerance

Potential costs, including non-monetary costs, which result from expected or realized AI errors or system functionality and trustworthiness - as connected to organizational risk tolerance - are examined and documented. Costs of error are examined including non-monetary ones such as harm to individuals and communities, and connected explicitly to the tolerance the organisation has set.

Artefacts an auditor will ask for
  • Documented analysis of costs arising from system error or failure
  • Non-monetary costs identified, including harm to individuals and communities
  • The connection drawn between those costs and the stated risk tolerance
  • Input from parties outside the delivery team on what costs matter
Where this commonly fails
  • Only direct financial and remediation cost considered
  • Costs listed with no reference to tolerance, so no threshold is crossed or not crossed
  • Analysis performed once, before the deployment scope widened
AIRMF-MP-3.3
Targeted application scope is specified and documented based on the system's capability, established context, and AI system categorization

Targeted application scope is specified and documented based on the system’s capability, established context, and AI system categorization. Scope is bounded in writing and justified by capability and context, which is what keeps the risk mapping and the evaluation resources tractable.

Artefacts an auditor will ask for
  • The documented targeted application scope and its boundaries
  • The capability and context basis for the scope chosen
  • Controls that keep use inside the stated scope
  • Records of any scope extension and the re-assessment that accompanied it
Where this commonly fails
  • Scope broad enough that no use is outside it
  • Boundary documented with nothing enforcing it
  • Scope extended in practice without re-assessment
AIRMF-MP-3.4
Processes for operator and practitioner proficiency with AI system performance and trustworthiness, and relevant technical standards and certifications, are defined, assessed and documented

Processes for operator and practitioner proficiency with AI system performance and trustworthiness – and relevant technical standards and certifications – are defined, assessed and documented. Operator proficiency is defined as a requirement, assessed rather than assumed, and recorded, because the human-AI configuration only works if the human can do the job assigned.

Artefacts an auditor will ask for
  • Defined proficiency requirements for operators and practitioners of the system
  • Assessment records showing proficiency was tested, not assumed
  • Relevant technical standards or certifications identified for the role
  • Refresh arrangements when the system or its performance changes
Where this commonly fails
  • Proficiency assumed from job title
  • Training delivered but proficiency never assessed
  • No refresh after a model update changed system behaviour
AIRMF-MP-3.5
Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from the GOVERN function

Processes for human oversight are defined, assessed, and documented in accordance with organizational policies from GOVERN function. Oversight is designed for the specific configuration, assessed for whether it is exercisable in practice, and consistent with the governing policy rather than invented per project.

Artefacts an auditor will ask for
  • The documented human oversight process for the system
  • Assessment of whether oversight is exercisable at the pace and volume of operation
  • The link to the governing organisational policy on oversight
  • The authority the overseer holds, including whether they can stop the system
Where this commonly fails
  • Oversight designed at a volume the reviewer cannot sustain
  • Overseer able to observe but not to intervene
  • Process differs from the governing policy with no recorded deviation
AIRMF-MP-4.1
Approaches for mapping AI technology and legal risks of its components, including the use of third-party data or software, are in place, followed, and documented, as are risks of infringement of a third party's intellectual property or other rights

Approaches for mapping AI technology and legal risks of its components – including the use of third-party data or software – are in place, followed, and documented, as are risks of infringement of a third-party’s intellectual property or other rights. There is a followed approach for mapping the technology and legal risk carried by each component, including data and software obtained from third parties and the rights position attached to them.

Artefacts an auditor will ask for
  • The documented approach for mapping component technology and legal risk
  • Component inventory identifying third-party data, models and software
  • Intellectual property and rights analysis for each third-party component
  • Evidence the approach was followed for the components actually in use
Where this commonly fails
  • Approach documented but not applied to components adopted since
  • Pre-trained models used with no analysis of the provenance of their training data
  • Rights reviewed for commercial components only, not for freely obtained ones
AIRMF-MP-4.2
Internal risk controls for components of the AI system including third-party AI technologies are identified and documented

Internal risk controls for components of the AI system including third-party AI technologies are identified and documented. For each component carrying risk, the internal control applied to it is identified and written down, so the risk mapping produces controls rather than a list.

Artefacts an auditor will ask for
  • The internal controls identified for each AI system component
  • Controls specific to third-party and open-source AI technologies
  • The pre-adoption evaluation practice for third-party material
  • Evidence the controls named are actually in place
Where this commonly fails
  • Risks identified for components with no control named against them
  • Freely available third-party material adopted outside the evaluation practice
  • Controls documented centrally but absent in the deployed pipeline
AIRMF-MP-5.1
Likelihood and magnitude of each identified impact are identified and documented, based on expected use, past uses of AI systems in similar contexts, public incident reports, feedback from those external to the team, or other data

Likelihood and magnitude of each identified impact (both potentially beneficial and harmful) based on expected use, past uses of AI systems in similar contexts, public incident reports, feedback from those external to the team that developed or deployed the AI system, or other data are identified and documented. Each identified impact carries a likelihood and a magnitude derived from stated evidence, including comparable past deployments and public incident reports, not from unsupported estimation.

Artefacts an auditor will ask for
  • Impact register with likelihood and magnitude recorded per impact
  • The evidence base cited for each estimate, including comparable systems and incident reports
  • Beneficial as well as harmful impacts characterised
  • The use of these estimates in a go or no-go decision
Where this commonly fails
  • Likelihood assigned by consensus in a workshop with no evidence cited
  • Only harmful impacts characterised, so trade-offs cannot be weighed
  • Estimates produced and never used in any decision
AIRMF-MP-5.2
Practices and personnel for supporting regular engagement with relevant AI actors and integrating feedback about positive, negative, and unanticipated impacts are in place and documented

Practices and personnel for supporting regular engagement with relevant AI actors and integrating feedback about positive, negative, and unanticipated impacts are in place and documented. Engagement with affected actors is a resourced standing practice with named personnel, and unanticipated impacts have a route back into the impact record.

Artefacts an auditor will ask for
  • The documented engagement practice and its cadence
  • Personnel assigned to conduct and integrate engagement
  • Records of impacts reported through engagement, including unanticipated ones
  • Evidence reported impacts were integrated into the impact record
Where this commonly fails
  • Engagement run as a one-off consultation at launch
  • No named personnel, so engagement depends on individual initiative
  • Unanticipated impacts reported with no route into the risk record

MEASURE - NIST AI RMF 1.0

AIRMF-MS-1.1
Approaches and metrics for measurement of AI risks enumerated during the MAP function are selected for implementation starting with the most significant AI risks, and the risks or trustworthiness characteristics that will not or cannot be measured are properly documented

Approaches and metrics for measurement of AI risks enumerated during the Map function are selected for implementation starting with the most significant AI risks. The risks or trustworthiness characteristics that will not – or cannot – be measured are properly documented. Measurement approaches are selected against the risks that mapping produced, most significant first, and anything left unmeasured is named as unmeasured rather than silently omitted.

Artefacts an auditor will ask for
  • The selected measurement approaches and metrics, traced to mapped risks
  • The prioritisation showing the most significant risks addressed first
  • An explicit record of risks and characteristics that will not or cannot be measured, with reasons
  • Evidence the selection was implemented rather than only planned
Where this commonly fails
  • Metrics chosen from what the tooling already emits rather than from the mapped risks
  • Unmeasured characteristics simply absent, so their absence reads as a pass
  • Selection documented with no evidence of implementation
AIRMF-MS-1.2
Appropriateness of AI metrics and effectiveness of existing controls is regularly assessed and updated, including reports of errors and impacts on affected communities

Appropriateness of AI metrics and effectiveness of existing controls is regularly assessed and updated including reports of errors and impacts on affected communities. The metrics themselves are re-examined on a cycle for whether they remain appropriate, informed by reported errors and by impacts on affected communities.

Artefacts an auditor will ask for
  • Records of periodic assessment of metric appropriateness and control effectiveness
  • Error reports and community impact reports considered in that assessment
  • Changes made to metrics or controls as a result
  • The trigger conditions, such as drift or changed operating setting, that force a re-assessment
Where this commonly fails
  • Metrics fixed at launch and carried unchanged through model updates
  • Assessment considers internal error rates only, not reported impacts
  • Re-assessment scheduled but no completed record for the current period
AIRMF-MS-1.3
Internal experts who did not serve as front-line developers for the system and independent assessors are involved in regular assessments and updates, and domain experts, users, AI actors external to the team, and affected communities are consulted in support of assessments as necessary per organizational risk tolerance

Internal experts who did not serve as front-line developers for the system and/or independent assessors are involved in regular assessments and updates. Domain experts, users, AI actors external to the team that developed or deployed the AI system, and affected communities are consulted in support of assessments as necessary per organizational risk tolerance. Assessment includes people who did not build the system, with the depth of external consultation set by the organisation's risk tolerance.

Artefacts an auditor will ask for
  • Assessment records naming assessors and their independence from development
  • Records of consultation with domain experts, users and affected communities
  • The link from risk tolerance to the level of independence required
  • Findings raised by independent assessors and their disposition
Where this commonly fails
  • Independent review performed by another team under the same delivery owner
  • Consultation limited to internal users
  • Independent findings recorded but closed without action
AIRMF-MS-2.1
Test sets, metrics, and details about the tools used during test, evaluation, validation, and verification are documented

Test sets, metrics, and details about the tools used during test, evaluation, validation, and verification (TEVV) are documented. The TEVV record is complete enough for the evaluation to be repeated: which test sets, which metrics, which tools and which versions.

Artefacts an auditor will ask for
  • Documentation of the test sets used, including provenance and composition
  • The metrics computed and their definitions
  • Tools and versions used during evaluation
  • Sufficient detail for the evaluation to be repeated by someone else
Where this commonly fails
  • Results reported with the test set identified only by filename
  • Metric named but not defined, so results are not comparable across runs
  • Tooling versions unrecorded, so a result cannot be reproduced
AIRMF-MS-2.10
Privacy risk of the AI system as identified in the MAP function is examined and documented

Privacy risk of the AI system – as identified in the MAP function – is examined and documented. Privacy examination covers what the AI system makes possible, inference and re-identification from training data and outputs, not only the lawfulness of the input data.

Artefacts an auditor will ask for
  • Privacy risk examination covering training data, inference and outputs
  • Assessment of re-identification and memorisation risk
  • Privacy-enhancing measures applied and their assessed effect
  • Documentation traced to the privacy risks mapping identified
Where this commonly fails
  • Privacy assessed as lawful basis for input data only
  • Memorisation and training data extraction not considered
  • Assessment completed for the original data set and not repeated after retraining
AIRMF-MS-2.11
Fairness and bias as identified in the MAP function is evaluated and results are documented

Fairness and bias – as identified in the MAP function – is evaluated and results are documented. Fairness evaluation states which fairness definition was applied and why, evaluates against it, and records the results including where the definition itself is contested.

Artefacts an auditor will ask for
  • The fairness definition or definitions applied and the reason for choosing them
  • Evaluation results disaggregated across the relevant groups
  • Identification of the groups assessed and the basis for that selection
  • Documentation of trade-offs where fairness definitions conflict
Where this commonly fails
  • A single fairness metric applied with no argument that it fits the context
  • Groups assessed limited to those for which attribute data happened to exist
  • Disparity measured and documented with no decision recorded about it
AIRMF-MS-2.12
Environmental impact and sustainability of AI model training and management activities as identified in the MAP function are assessed and documented

Environmental impact and sustainability of AI model training and management activities – as identified in the MAP function – are assessed and documented. The environmental cost of training and operating the system is assessed on stated measures, energy, water and greenhouse gas emissions, rather than treated as out of scope.

Artefacts an auditor will ask for
  • Assessment of energy consumption for training and inference
  • Water consumption and greenhouse gas emissions attributable to the system where applicable
  • The measurement basis and any published metrics adopted
  • Documentation of the assessment and its bearing on design decisions
Where this commonly fails
  • Environmental impact declared immaterial with no measurement
  • Training cost assessed while ongoing inference cost is ignored
  • Assessment held by the infrastructure team with no link to the AI system record
AIRMF-MS-2.13
Effectiveness of the employed TEVV metrics and processes in the MEASURE function are evaluated and documented

Effectiveness of the employed TEVV metrics and processes in the MEASURE function are evaluated and documented. The evaluation apparatus is itself evaluated, for whether the metrics still discriminate, whether they are being optimised against, and whether they carry unexamined assumptions.

Artefacts an auditor will ask for
  • Evaluation of whether the TEVV metrics remain effective and discriminating
  • Consideration of gaming, saturation and drift in the metrics themselves
  • Review of assumptions embedded in the measurement approach
  • Changes made to TEVV processes as a result
Where this commonly fails
  • Metrics never questioned once adopted
  • Metric saturation read as system improvement
  • Effectiveness review performed by the team whose work the metrics judge
AIRMF-MS-2.2
Evaluations involving human subjects meet applicable requirements including human subject protection and are representative of the relevant population

Evaluations involving human subjects meet applicable requirements (including human subject protection) and are representative of the relevant population. Where evaluation involves human subjects or data captured from them, the applicable protection requirements are met and the subject population is representative of the deployment population.

Artefacts an auditor will ask for
  • Identification of evaluations involving human subjects or human subject data
  • Evidence the applicable human subject protection requirements were met, including any approval obtained
  • Analysis of representativeness against the relevant population
  • Consent and data handling records for subject data used in evaluation
Where this commonly fails
  • Human subject involvement not recognised because the data was already held
  • Convenience sample used with no representativeness analysis
  • Protection requirements treated as applying only to funded research
AIRMF-MS-2.3
AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment settings, and measures are documented

AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s). Measures are documented. Performance is demonstrated under conditions resembling deployment rather than only on held-out test data, and the measures are recorded.

Artefacts an auditor will ask for
  • The performance or assurance criteria set for the system
  • Measurement results obtained under conditions similar to the deployment setting
  • The argument that the test conditions resemble deployment
  • Documentation of the measures used, qualitative as well as quantitative
Where this commonly fails
  • Performance demonstrated only in silico on a static test set
  • Deployment conditions asserted to be similar with no analysis
  • Qualitative assurance criteria stated but never measured
AIRMF-MS-2.4
The functionality and behavior of the AI system and its components, as identified in the MAP function, are monitored when in production

The functionality and behavior of the AI system and its components – as identified in the MAP function – are monitored when in production. Production monitoring covers the functionality and behaviour that mapping identified as consequential, so drift away from the design assumptions is detected while running.

Artefacts an auditor will ask for
  • The production monitoring configuration and what it observes
  • Traceability from monitored signals to the behaviours mapping identified
  • Drift detection results and the thresholds applied
  • Alerting and the named recipients of monitoring output
Where this commonly fails
  • Monitoring covers infrastructure availability rather than model behaviour
  • Drift thresholds set but nothing acts when they are crossed
  • Component behaviour unmonitored where the component is third-party
AIRMF-MS-2.5
The AI system to be deployed is demonstrated to be valid and reliable, and limitations of the generalizability beyond the conditions under which the technology was developed are documented

The AI system to be deployed is demonstrated to be valid and reliable. Limitations of the generalizability beyond the conditions under which the technology was developed are documented. Validity and reliability are demonstrated before deployment, and the boundary beyond which the demonstration does not carry is written down.

Artefacts an auditor will ask for
  • Validation results demonstrating validity and reliability prior to deployment
  • Documented limitations on generalisability beyond the development conditions
  • The conditions under which the system was developed and tested
  • Evidence that validation failure would prevent deployment
Where this commonly fails
  • Validity claimed from training performance rather than independent validation
  • Generalisability limits known informally and not documented
  • Validation performed after deployment as a formality
AIRMF-MS-2.6
AI system is evaluated regularly for safety risks as identified in the MAP function, is demonstrated to be safe, its residual negative risk does not exceed the risk tolerance, and it can fail safely, particularly if made to operate beyond its knowledge limits

AI system is evaluated regularly for safety risks – as identified in the MAP function. The AI system to be deployed is demonstrated to be safe, its residual negative risk does not exceed the risk tolerance, and can fail safely, particularly if made to operate beyond its knowledge limits. Safety metrics implicate system reliability and robustness, real-time monitoring, and response times for AI system failures. Safety is evaluated on a recurring basis against the mapped safety risks, residual risk is compared to the stated tolerance, and behaviour beyond the knowledge limits is tested for safe failure.

Artefacts an auditor will ask for
  • Safety evaluation results traced to the safety risks mapping identified
  • The residual safety risk recorded and compared against the stated tolerance
  • Test results for behaviour beyond the system's knowledge limits
  • Safety metrics covering reliability, robustness, real-time monitoring and failure response time
Where this commonly fails
  • Safety evaluated once before launch and not repeated
  • Residual risk documented but never compared to a tolerance
  • Out-of-distribution behaviour untested, so fail-safe behaviour is unverified
AIRMF-MS-2.7
AI system security and resilience as identified in the MAP function are evaluated and documented

AI system security and resilience – as identified in the MAP function – are evaluated and documented. Security and resilience evaluation addresses the AI-specific attack surface, adversarial input, poisoning, extraction and supply chain, and records whether the system degrades gracefully.

Artefacts an auditor will ask for
  • Evaluation results for AI-specific security risks, including adversarial input and data poisoning
  • Resilience testing showing behaviour under unexpected adverse events or environment change
  • Model and data supply chain integrity checks
  • Documentation of the evaluation traced to the security risks mapping identified
Where this commonly fails
  • Conventional application security testing substituted for AI-specific evaluation
  • Resilience assessed as infrastructure failover only
  • Model extraction and inversion risks not considered
AIRMF-MS-2.8
Risks associated with transparency and accountability as identified in the MAP function are examined and documented

Risks associated with transparency and accountability – as identified in the MAP function – are examined and documented. The examination is of transparency and accountability as risks: what information asymmetry remains between the organisation and those affected, and who can be held to account for an outcome.

Artefacts an auditor will ask for
  • Examination of what information about the system is and is not disclosed, and to whom
  • Identification of the accountable party for system outcomes
  • Assessment of residual information asymmetry with operators and affected communities
  • Documentation traced to the transparency risks mapping identified
Where this commonly fails
  • Transparency treated as equivalent to publishing model documentation
  • Accountability assigned to the system rather than to a person or function
  • Disclosure assessed for regulators only, not for affected individuals
AIRMF-MS-2.9
The AI model is explained, validated, and documented, and AI system output is interpreted within its context as identified in the MAP function and to inform responsible use and governance

The AI model is explained, validated, and documented, and AI system output is interpreted within its context – as identified in the MAP function – and to inform responsible use and governance. Explanations are validated for fidelity to actual model behaviour and are pitched at the audience that must act on them, and the interpretation of output is bound to the mapped context.

Artefacts an auditor will ask for
  • The explanation method used and validation of its fidelity to model behaviour
  • Model documentation covering how the model reaches its outputs
  • Guidance on interpreting output within the deployment context
  • Identification of the audiences for explanation and what each needs
Where this commonly fails
  • Explanation method adopted with no check that it reflects the model
  • Feature attributions produced for developers and never translated for the people affected
  • Limitations of the explanation method not stated
AIRMF-MS-3.1
Approaches, personnel, and documentation are in place to regularly identify and track existing, unanticipated, and emergent AI risks based on factors such as intended and actual performance in deployed contexts

Approaches, personnel, and documentation are in place to regularly identify and track existing, unanticipated, and emergent AI risks based on factors such as intended and actual performance in deployed contexts. Risk identification continues after deployment with assigned personnel and a tracking record, so risks that emerge in real use are captured rather than only those anticipated at design.

Artefacts an auditor will ask for
  • The approach for identifying emergent and unanticipated risks in deployment
  • Personnel assigned to that identification and tracking
  • The tracking record showing risks identified after deployment
  • Comparison of actual against intended performance in the deployed context
Where this commonly fails
  • Risk identification treated as a design-phase activity that ends at launch
  • Emergent risks noted in incident tickets but never entered as risks
  • No one assigned, so identification depends on something going visibly wrong
AIRMF-MS-3.2
Risk tracking approaches are considered for settings where AI risks are difficult to assess using currently available measurement techniques or where metrics are not yet available

Risk tracking approaches are considered for settings where AI risks are difficult to assess using currently available measurement techniques or where metrics are not yet available. Where no adequate metric exists the risk is still tracked by some stated means, so difficulty of measurement does not become silent omission.

Artefacts an auditor will ask for
  • Identification of risks for which adequate measurement techniques do not exist
  • The tracking approach adopted for each such risk
  • Any novel or qualitative measurement approaches trialled
  • Review of whether measurement has since become possible
Where this commonly fails
  • Hard-to-measure risks dropped from the register rather than tracked qualitatively
  • Approach considered once and not revisited as techniques matured
  • Absence of a metric reported as absence of risk
AIRMF-MS-3.3
Feedback processes for end users and impacted communities to report problems and appeal system outcomes are established and integrated into AI system evaluation metrics

Feedback processes for end users and impacted communities to report problems and appeal system outcomes are established and integrated into AI system evaluation metrics. End users and impacted communities have a working route to report problems and appeal an outcome, and what arrives through it feeds the evaluation metrics rather than a separate queue.

Artefacts an auditor will ask for
  • The reporting and appeal route available to end users and impacted communities
  • Records of problems reported and appeals lodged, with outcomes
  • The integration of that feedback into evaluation metrics
  • Evidence the route is discoverable by the people expected to use it
Where this commonly fails
  • Appeal route exists in policy but is not reachable from the point of the decision
  • Reports handled as customer service tickets with no route into evaluation
  • No appeal available where the decision is automated and consequential
AIRMF-MS-4.1
Measurement approaches for identifying AI risks are connected to deployment contexts and informed through consultation with domain experts and other end users, and approaches are documented

Measurement approaches for identifying AI risks are connected to deployment context(s) and informed through consultation with domain experts and other end users. Approaches are documented. The measurement design is informed by people who understand the deployment context, because the risks that matter there are often not visible to those running the evaluation.

Artefacts an auditor will ask for
  • Documentation of the measurement approaches and their connection to deployment contexts
  • Records of consultation with domain experts and end users on measurement design
  • Evidence that consultation changed the measurement approach
  • Identification of the deployment contexts the measurement is meant to cover
Where this commonly fails
  • Measurement designed entirely by the evaluation team
  • Consultation held after metrics were fixed
  • One measurement approach applied across materially different deployment contexts
AIRMF-MS-4.2
Measurement results regarding AI system trustworthiness in deployment contexts and across the AI lifecycle are informed by input from domain experts and other relevant AI actors to validate whether the system is performing consistently as intended, and results are documented

Measurement results regarding AI system trustworthiness in deployment context(s) and across AI lifecycle are informed by input from domain experts and other relevant AI actors to validate whether the system is performing consistently as intended. Results are documented. Measured results are validated against the judgement of people who know the context, so a result inside its operational limits on paper is confirmed as adequate in practice.

Artefacts an auditor will ask for
  • Measurement results with domain expert and AI actor input recorded against them
  • The pre-defined operational limits the results are judged against
  • Documentation of whether the system is performing consistently as intended
  • Disposition of any expert view that conflicted with the measured result
Where this commonly fails
  • Results published with no expert validation of what they mean in context
  • Operational limits set after the results were known
  • Conflicting expert judgement recorded and then disregarded without reasoning
AIRMF-MS-4.3
Measurable performance improvements or declines based on consultations with relevant AI actors including affected communities, and field data about context-relevant risks and trustworthiness characteristics, are identified and documented

Measurable performance improvements or declines based on consultations with relevant AI actors including affected communities, and field data about context-relevant risks and trustworthiness characteristics, are identified and documented. Change in performance over time is measured against a baseline using field data and consultation, so decline is detected as decline rather than absorbed as normal variation.

Artefacts an auditor will ask for
  • Baseline measures for the trustworthiness characteristics being tracked
  • Field data showing performance over time against that baseline
  • Consultation records with AI actors and affected communities on observed change
  • Documented identification of improvement or decline and the action taken
Where this commonly fails
  • No baseline, so change cannot be identified
  • Field data collected but never compared across periods
  • Decline attributed to data quality without investigation
Assembled from the framework's own control set. Every line traces to a control in the graph, so this pack is regenerated rather than written, and stays current as the graph does.

Assembled from the framework’s own control set, so this list is regenerated rather than written and stays current as the graph does.