Enterprise AI Lab · Research note

Red-teaming high-risk AI
under Article 9.
A working method.

Article 9 of the EU AI Act requires testing that finds the risks worth managing. This note sets out how the Lab runs adversarial testing against high-risk systems — and is honest about where the method’s limits sit.

Article 9
The EU AI Act duty this serves
5
Stages in the working method
Open
Published for challenge, limits included
Free
No gate, no registration, no email

Why we are publishing this

This is the Enterprise AI Lab’s first published output, and it is deliberately modest in what it claims. It is not a standard and it is not a compliance checklist. It is a description of the method we currently use to run adversarial testing against high-risk AI systems — written down so that clients, peers and, where relevant, regulators can see the reasoning and challenge it.

Two reasons to publish. First, Article 9 of the EU AI Act requires testing but does not prescribe how to test, and the harmonised standards that will eventually fill that gap are still being drafted. Organisations deploying high-risk systems have to make defensible methodological choices now, not once the standards land. Second, writing a method down improves it. Where our practice changes, this note will change with it, and revisions will be dated.

What Article 9 actually asks for

Article 9 requires every high-risk AI system to sit inside a risk management system: established, implemented, documented and maintained as a continuous, iterative process across the system’s whole lifecycle. Within that process, providers must identify the known and reasonably foreseeable risks the system poses to health, safety and fundamental rights; estimate the risks that arise under intended use and under reasonably foreseeable misuse; and test the system — against metrics defined in advance — to identify the most appropriate risk-management measures. Deployers are not spared: they must operate systems in line with the provider’s instructions and monitor them, and in regulated financial services the practical burden of demonstrating oversight lands on the firm using the system.

Two phrases in that text do most of the work. The first is reasonably foreseeable misuse. Deliberate adversarial attacks on AI systems are well documented, publicly discussed and inexpensive to mount; it is difficult to argue they are not foreseeable. A risk process that never attempts them has a visible gap. The second is the requirement that testing identifies risk-management measures — testing exists to change the system and its controls, not to produce a report.

Article 15 removes any remaining doubt about the attack surface: it requires technical measures against AI-specific vulnerabilities and names data poisoning, adversarial examples and model evasion explicitly. Read together, Articles 9, 14 and 15 imply adversarial testing as part of ordinary risk management for high-risk systems. For UK organisations the Act reaches further than many assume — our EU AI Act guide covers scope — and FCA and PRA operational-resilience expectations point in the same direction.

The working method, in five stages

Stage one is scoping. We agree what the system is for — its intended purpose as documented, not as assumed — and where the boundaries of the test sit: which components are in scope (the model, the retrieval layer, the tooling it can call, the integrations around it), what environment we test in, and what counts as a finding. Article 9 defines risk relative to intended purpose, so a test that is vague about purpose cannot say anything useful about risk.

Stage two is threat modelling. Before any attack is attempted, we ask who would realistically attack this system and why: the user population, the incentive to defraud or manipulate, the data the system touches, and the decisions downstream of its output — credit, underwriting, employment and access to services being the classifications that matter most in financial services. The output is a prioritised set of harm scenarios, each traceable to the risk categories Article 9 names.

Stage three is attack execution. We work through attack classes in a defined order rather than improvising:

  • Prompt injection — direct, and indirect via documents, emails or retrieved content the system ingests;
  • Data poisoning — manipulation of the training, fine-tuning or retrieval corpora the system depends on;
  • Evasion — inputs crafted so the system misclassifies, misprices or bypasses its own output constraints;
  • Extraction — probing for training data, system prompts or confidential context.

Stage four tests human-oversight bypass. Article 14 assumes a person can understand the system, intervene and interrupt it. We test whether that holds under adversarial conditions: whether outputs can be made to look confident and well-formed while wrong, whether volume and pacing defeat meaningful review, and whether the stop mechanisms work when someone actually uses them.

Stage five is evidence capture. Every attempt is logged — payload, configuration, model version, outcome, reproduction steps — and every material finding is mapped to a risk-management measure or an explicit acceptance of residual risk. For Sentinel clients this evidence lands directly in Citadel, so the test becomes part of the governance record rather than a PDF sitting beside it.

What “adequate” looks like

There is no published threshold for adequate adversarial testing under the Act, and we will not pretend one exists. Our working position is that adequacy is a property of the process, not of the pass rate. In practice we look for five things. Every test traces to a foreseeable risk, and every material risk either has a test against it or a documented reason it does not. Metrics are defined before testing begins, as Article 9(8) requires, rather than fitted to the results. Findings are reproducible by someone who was not in the room. Testing demonstrably changes something — a control adopted, a design revised, or a residual risk accepted at the right level of seniority and recorded as such. And depth is proportionate: a system informing credit decisions warrants more adversarial attention than one drafting internal correspondence.

An organisation that can show that chain — risk, test, finding, measure, sign-off — cannot prove its systems are invulnerable. Nobody can. It can demonstrate the risk management process the Act actually asks for, which is the more defensible claim.

The limits of this method

An honest method states what it cannot do. Red-teaming samples a failure space; it does not enumerate it. An absence of findings is evidence of effort, not of safety, and we write our reports accordingly. Results are also version-bound: a model update, a fine-tune or a change to a retrieval corpus can invalidate previous findings, which is why we treat adversarial testing as a recurring control inside the risk management cycle rather than a one-off certification exercise.

The attack classes above reflect the current landscape and will age. Attackers adapt faster than methodologies, and a commissioned red team works with bounded time, knowledge and permission — real adversaries accept none of those constraints. Finally, the regulatory ground is still moving. We test this method against NIST’s AI Risk Management Framework and the management-system discipline of ISO 42001, but the harmonised standards being drafted under the Act may draw the line for “adequate” somewhere else. When they do, this note will be revised rather than defended.

The Lab takes commissioned engagements to run this method against live systems — scoped per engagement, under NDA, with findings you can act on. Details are on the Lab overview.

Talk to the people who take AI seriously.

Thirty minutes with a founder. No sales deck, no obligation.

Price agreed before we start · Founder-led