Data for AI Systems - ISO 42001 Annex A Controls

ISO 42001 Annex A.7

Data is the foundation of any AI system. Annex A.7 makes data management an explicit part of the management system, not a downstream technical concern.

ISO 42001 Annex A.7 - Data for AI Systems Explained

Annex A.7 recognises that data quality and data management are the foundation of AI system performance and trustworthiness. The five controls work together to make sure data used by AI systems is appropriate, traceable and prepared in a way that supports the AI system's intended use.

Control A.7.2 - Data for development and enhancement of AI system

The organisation must define, document and implement data management processes related to the development of AI systems. The implementation guidance under Annex B.7.2 highlights the topics typically covered - privacy and security implications, security and safety threats from data dependence, transparency and explainability including data provenance, representativeness of training data compared to operational use, and accuracy and integrity of the data.

Control A.7.3 - Acquisition of data

The organisation must determine and document details about the acquisition and selection of data used in AI systems. The details typically include the categories of data needed, the quantity, the data sources (internal, purchased, shared, open data, synthetic), the characteristics of the data source, data subject demographics and characteristics, prior handling of the data, data rights including personal data and copyright, associated metadata and the provenance of the data.

Control A.7.4 - Quality of data for AI systems

The organisation must define and document requirements for data quality and make sure data used to develop and operate the AI system meets those requirements. Data quality is the degree to which the characteristics of data satisfy stated and implied needs when used under specified conditions. The organisation must consider the impact of bias on system performance and fairness, and adjust the model and data as needed to improve performance and fairness for the use case.

Control A.7.5 - Data provenance

The organisation must define and document a process for recording the provenance of data used in its AI systems over the life cycle of the data and the AI system. Data provenance covers the creation, update, transcription, abstraction, validation and transferring of control of data. Data sharing and data transformations are also part of provenance.

Control A.7.6 - Data preparation

The organisation must define and document its criteria for selecting data preparation methods. Common preparation methods include statistical exploration, cleaning, imputation of missing entries, normalisation, scaling, labelling of target variables, and encoding of categorical variables. The criteria should reflect the AI task and the expected operational environment.

Application to deployers

For AI deployers using third-party AI services, the data controls in A.7 apply differently. The developer of the AI system handles the training data. The deployer handles the production data put into the AI system in operation - the inputs the AI system receives from the deployer's environment, the prompts entered by users, the documents or images uploaded for analysis. Deployers document the categories of production data they use, the data quality requirements they apply, the provenance of the production data, and any preparation done before passing data to the AI system.

Data controls are where ISO 42001 intersects most heavily with data protection and information security obligations. AI training data often includes personal data, which brings the UK GDPR into play. Production data can include personal data too, particularly where AI is used in customer-facing services. The data controls in A.7 need to be coherent with the privacy controls under ISO 27701 and the information security controls under ISO 27001.

Bias in training data is the AI-specific concern that distinguishes A.7 from generic data management controls. The standard expects the organisation to consider bias actively, document the assessment, and take appropriate action where the impact assessment or risk assessment identifies bias as a concern.

When auditing Annex A.7, I look for documented data management processes and evidence of their implementation. For developers, I expect to see training data records, quality assessment, bias considerations and provenance records. For deployers, I expect to see production data documentation, data quality controls in operation, and any preparation steps applied before passing data to the AI system.

The integration with the impact assessment is important. If the impact assessment identifies a risk of bias, I expect A.7.4 to address it specifically, with the bias assessment and any mitigations recorded.

For the inspection AI, the supplier handled the training data. We documented their description and limitations. For our production data, we record the products being inspected, the inspection conditions and the calibration state of the cameras. That is our piece of the data picture, and the supplier has theirs.

Practical Compliance Guidance

The dedicated F-Q112 AI Data Matrix is the operational document for capturing data resources and data management arrangements for each AI system. It supports compliance with all five Annex A.7 controls.

The following alphaZ documents support compliance with ISO 42001 Annex A.7.

alphaZ document How to use it
ISO 42001 AI Management System Toolkit The full toolkit containing the AI management system documentation including the AI Data Matrix and supporting templates.
F-Q112 AI Data Matrix Records the data resources used by each AI system including categories, sources, quality requirements, provenance, preparation methods and bias considerations.
F-IMS40 AI Process Register Records the AI systems within scope and links each to the relevant data documentation.
F-IMS70 Annex A Controls Records the Statement of Applicability including the A.7 controls with the implementation status and supporting evidence.

Note - all the above files can be downloaded with an alphaZ subscription.

Frequently Asked Questions

Training data is typically the developer's responsibility. The deployer relies on the documentation provided by the supplier and references it in the AI Data Matrix. A.7 applies to the deployer for the production data used in operation, the data quality requirements applied at the deployer's end, and the integration of training data documentation from the supplier into the deployer's records.
The control requires the organisation to consider the impact of bias on system performance and fairness. The implementation guidance recognises that adjustments to the model and data may be needed to improve fairness for the use case. The bias assessment should be documented, with mitigations applied where the impact assessment or risk assessment identifies bias as a concern.
Data provenance is the record of where data came from and what has happened to it - creation, updates, transcription, abstraction, validation and transfers of control. It matters for AI systems because it supports trust in the AI system's outputs, makes it possible to investigate problems back to their source, and provides evidence for regulators and customers about how the AI system has been built and operated.

UK Legislation

The following UK legislation is directly relevant to data used in AI systems. Organisations outside the UK should identify the equivalent legislation applicable in their jurisdiction.

Further Resources

payment logos