A supplier can provide an impressive model card, security certificate and sales presentation, yet still leave your organisation unable to explain how an AI system was classified, tested, monitored or controlled. That is the gap AI supplier due diligence must close. It is not a procurement formality. It is the process that establishes whether a supplier’s claims can support your own legal, risk and assurance obligations.
For organisations deploying AI in the UK and Europe, the question is no longer simply whether a vendor is secure or commercially viable. You need to know whether its system can be governed in practice, whether evidence will remain available throughout the lifecycle, and whether contractual rights match the accountability your organisation retains.
Why AI supplier due diligence is a governance control
Buying an AI system does not transfer accountability for its use. Under the EU AI Act, responsibilities depend on the role your organisation performs. A provider has extensive obligations for a high-risk AI system, including risk management, technical documentation, data governance, quality management and post-market monitoring. A deployer also has defined obligations, including using the system in accordance with instructions, monitoring its operation, retaining logs where applicable, and ensuring human oversight.
The division is not always neat. A deployer that substantially modifies a system, changes its intended purpose, or places its own AI product on the market may take on provider obligations. A generic AI model embedded in an internal process can also create risks that neither the procurement team nor the business owner initially anticipated.
This is why a standard third-party security questionnaire is insufficient. It may establish that a supplier has encryption, access controls and incident response arrangements. It will not necessarily establish whether the training data is suitable for the intended use, whether performance has been validated for your population, or whether the supplier can supply the evidence needed for a high-risk assessment.
Effective due diligence connects supplier evidence to the AI system record, legal classification, risk assessment and control plan. It gives compliance teams a defensible basis for approval, conditional approval or rejection.
Start with the use case, not the supplier’s marketing
The same model can create markedly different obligations depending on how it is deployed. A language model used to draft internal meeting notes is not assessed in the same way as one used to rank job applicants, support credit decisions or triage patients. Due diligence must therefore begin with a clear description of the intended purpose, affected individuals, decision rights, jurisdictions, data categories and integration points.
Ask the business owner what the system actually does, not what it is called. Identify whether it produces recommendations, scores, classifications, generated content or automated decisions. Establish who can override outputs, what happens when an output is wrong, and whether the output has a material effect on a person.
This initial scoping determines the depth of review. A low-impact internal productivity tool may justify a proportionate assessment focused on confidentiality, data handling and acceptable use. A system that may fall within an EU AI Act high-risk use case requires deeper technical and legal evidence. Treating both reviews identically either wastes resource or misses material risk.
What to test in AI supplier due diligence
A useful assessment examines evidence across five connected areas:
- Legal status and intended purpose: Confirm the supplier’s role, the system’s stated purpose, geographical availability, applicable AI Act classification, and any prohibited-practice concerns. Where a supplier claims that a system is not high risk, request the reasoning rather than accepting a label.
- Data governance and privacy: Establish what data is used for training, fine-tuning, retrieval and inference; where it is stored; whether customer inputs are retained; and whether personal data may be used to improve the service. Check sub-processors, international transfers, deletion arrangements and data subject rights support.
- Technical performance and limitations: Request evaluation methodology, test results, known failure modes, accuracy measures relevant to the use case, bias testing where appropriate, version history and documented limitations. Aggregate benchmark scores are rarely enough.
- Security, resilience and operational control: Review identity and access management, environment separation, vulnerability management, logging, incident handling, business continuity and supplier dependency. For generative systems, test controls for prompt injection, data leakage and unauthorised tool use.
- Governance and accountability: Assess named owners, quality management arrangements, change controls, human oversight guidance, complaints handling, audit support and post-market monitoring. Evidence should show an operating process, not merely a policy statement.
The strength of the evidence matters as much as its existence. A policy that says a supplier tests for bias is weak evidence without scope, methodology, results, thresholds, remediation records and approval ownership. A declaration of EU AI Act alignment is not a substitute for traceable documentation.
Request evidence that can be retained and tested
Due diligence often fails because teams collect documents in email threads, complete a spreadsheet, then cannot locate the basis for approval six months later. The assessment should create an evidence file tied to the particular AI use case, supplier and version being approved.
For systems that may be high risk, review the elements that support the provider’s technical documentation, including the system description, intended purpose, risk management measures, data governance approach, testing information, human oversight measures, cybersecurity provisions and performance information. Annex IV of the EU AI Act provides a practical reference point for the type of information technical documentation should contain.
Do not assume the supplier will hand over its full technical documentation. Some information will be commercially sensitive, particularly for foundation models and proprietary systems. The practical question is whether the supplier can provide enough evidence for your organisation to assess and govern the deployment. That may involve controlled auditor access, a detailed assurance pack, independent assessment reports, contractual disclosure commitments or a supervised demonstration of controls.
Where evidence is unavailable, record the gap explicitly. A conditional approval can be appropriate where the use case is limited, compensating controls are credible and remediation has a defined deadline. It is not appropriate where the missing evidence prevents a meaningful assessment of a material risk.
Build a workflow that survives audit scrutiny
The most defensible process is repeatable and owned. First, register the proposed AI system and its supplier in a central inventory. Assign a business owner, technical owner, risk owner and approval authority. Capture the use case, data flows, affected groups and system dependencies before questionnaires are sent.
Next, perform a preliminary legal classification and risk triage. This determines which evidence is mandatory, which specialist reviewers are required and whether deployment must wait for formal approval. Legal, privacy, information security, model risk and procurement should not each run isolated reviews of the same system. They need a shared record with clear findings and actions.
Then assess supplier responses against defined controls. For each control, record the evidence received, reviewer judgement, residual risk, required action, target date and accountable owner. Avoid a simple pass-or-fail score where the decision depends on context. A supplier may have mature security controls but inadequate transparency for a sensitive employment use case. That should result in a specific governance decision, not a misleading average score.
Approval should state the permitted use, any data restrictions, required human oversight, monitoring expectations and re-assessment trigger. The decision record should be available to the board, internal audit, customers and regulators in a form each can understand.
Put the evidence requirements into the contract
Pre-contract diligence is only as durable as the contractual rights behind it. Supplier terms should address notice of material model, data-processing or sub-processor changes; access to relevant assurance evidence; incident notification; cooperation with regulatory enquiries; audit or audit-report rights; retention and deletion obligations; and termination support.
The right level of contractual control depends on bargaining power and risk. A major cloud or foundation-model provider may not accept bespoke audit rights. In that case, decide whether its standard assurance package, certifications and reporting commitments are sufficient for the intended use. If they are not, the correct answer may be to limit the use case or select another supplier.
Due diligence continues after go-live
AI systems change more frequently than traditional software. A model update can alter output quality, data handling, safety behaviour and legal classification assumptions without changing the product name. Monitoring cannot be left to annual supplier renewal.
Set review triggers for material model releases, new integrations, expansion to new users or countries, significant incidents, changes to training or retention practices, adverse audit findings and changes in the intended purpose. Capture operational feedback from users, override rates, error reports and complaints. These are governance signals, not merely service-management metrics.
ISO/IEC 42001 supports this discipline by treating third-party relationships, risk treatment, documented information, performance evaluation and continual improvement as management-system activities. It helps turn supplier assurance from a one-off questionnaire into a controlled lifecycle process.
A practical platform such as Endaxi AIG can keep the inventory, classification rationale, evidence, control actions, approvals and review dates in one system of record. That is materially stronger than relying on disconnected procurement folders and spreadsheets when an auditor asks why a particular AI supplier was approved.
The useful test is simple: if a regulator, customer or internal auditor asked tomorrow why this supplier and this AI use case were acceptable, could you produce the decision, evidence, controls and current ownership without reconstructing the story from inboxes? If not, the due diligence is not finished.

