Factory Floor · 04

Attack it on purpose,
before someone does it
for free.

Two related builds. The evaluation infrastructure that tells you whether a model does what someone claims, in the conditions it will actually meet. And the adversarial programme that goes looking for the ways it fails — on your schedule rather than an attacker’s.

You hold: the evaluation suite, the adversarial findings and evidence assembled for the obligation, not a summary of it.

Article 9
The EU AI Act obligation this method is built against
4
Published papers and research notes from the Lab
Live
Evaluation suites run against a system’s real endpoint
0
Certifications claimed on your behalf

Two problems, usually confused

Evaluation asks whether it works. Red teaming asks how it breaks.

Evaluation infrastructure is a measurement problem. Does this model do what the vendor, or your own team, says it does — against real cases, at the latency you need, at a cost you can carry, and when the input is messy rather than curated? Most organisations discover they cannot answer this about a system they have already deployed, because the only evidence is a demo somebody else controlled.

Red teaming is an adversarial problem, and it is a different discipline. It asks what an intelligent, motivated party could make this system do. Prompt injection through content the system ingests. Nested instructions layered into a memory system so the dangerous prompt is one nobody wrote — the Inception Hack is our published work on exactly that. Extraction of data the model should not surface. Behaviour outside the envelope the guardrails assume.

Both are buildable as infrastructure rather than as one-off exercises, and that is the point of this class. An evaluation you ran once is a snapshot; an evaluation suite wired into CI is a control.

The method is free. The work is not.

The Lab publishes its working methods — the staged Article 9 approach, from threat modelling to evidence capture, is on the research page. Read it, use it, criticise it. What you buy is us running it on your system.

How it runs

Failures become findings, with an owner and a date.

A red-team report that ends in a PDF has changed nothing. Findings have to land somewhere they get worked.

  • Threat model the system, not the technology

    What would someone want from this system, and what does it have access to. Generic AI threat lists produce generic findings nobody prioritises.

  • Build the evaluation suite

    Golden sets from real cases, adversarial cases, latency and cost envelopes. Run against the live endpoint so the result reflects the deployed system rather than a copy of it.

  • Run the adversarial programme

    Injection through ingested content, nested instructions in layered memory, extraction attempts, guardrail evasion. Documented as it runs, with the attempts that failed recorded too.

  • Capture evidence for the obligation

    Under Article 9 the question is what adequate testing means in practice. The output is designed to be the evidence, not a summary of it.

  • Wire it in, so it keeps running

    Suites in continuous integration; regression alerts as new adversarial cases are added. Eval maintenance is what a Lab Retainer buys.

The two lists that matter

What you hold, and what we will not do.

Both are in the engagement letter before you sign it. The second list is the one worth reading twice — it is where most disappointment in this market actually comes from.

What you hold at the end

  • The evaluation suite, runnable against your live endpoint
  • The adversarial findings — including the attempts that failed
  • Evidence assembled for the obligation, not a summary of it
  • Each failure as a risk finding with a named owner and a due date
  • The suite wired into CI, with regression alerting

What we won’t do

  • Red-team reports that end in a PDF and no owner
  • Evaluations against curated inputs the system will never meet
  • Claiming a certification or compliance outcome on your behalf
  • Testing a copy of the system rather than the deployed one
  • One-off exercises sold as a control

Where this work comes from

The Lab runs ahead of the rulebook on purpose.

Red-teaming methodology for high-risk systems under Article 9, governance models for agentic AI that operate outside human-in-the-loop assumptions, and evaluation of whether a model does what someone claims. Published where the work is ready to share, and commissionable when you need it pointed at your system.

Explore the Lab

It feeds assurance

Test results are the evidence layer under governance and assurance — and in Citadel, evaluation failures become tracked risk findings against the system that produced them.

What would an attacker get out of it?

Thirty minutes with a founder. Bring the system you are least comfortable defending, and we will tell you where we would start looking.

Fixed fee · Quoted in writing before we start · NDA available