Factory Floor
We build four kinds
of thing.
All of them run.
Not a capability list — a shortlist. These are the classes of system we have built, can staff on Monday, and will show you working. Anything outside them we will tell you to buy, or tell you we are the wrong firm.
The rule of the floor
Evals before prompt code. Rollback before the first commit.
Every build on this floor starts with two artefacts that are not code: the evaluation harness that decides whether the thing works, and the rollback plan for when it does not. Both are agreed before a line is written, because a system whose success criteria were invented afterwards cannot be defended to anyone — not a board, not an auditor, and not the team who has to run it at 3am.
What follows is the same for all four classes. Architecture with failure-mode analysis from day one. Evals in continuous integration on every commit. Monitoring, alerting and drift detection tied to business impact rather than to model metrics. Handover documentation a team that was not in the room can operate from.
What we won’t build
Production systems without evals. Chatbots without guardrails or monitoring. Black-box models you cannot audit. Builds without a rollback plan. Autonomous agents without human decision gates. Handoffs to teams who were not in the room.
The four classes
What comes off this floor.
Each has its own page: what it is, when it is the right answer, what you hold at the end, and where it goes wrong. Read the one that sounds like your problem.
01
Production LLM Systems
Retrieval, reasoning and generation pipelines taken to production quality — golden sets, adversarial cases, latency budgets and cost ceilings in the harness before the prompt code exists. The class most clients arrive asking for.
02
Governed Multi-Agent Workforce
The thing we run ourselves, stood up inside your organisation for a defined function — code review, incident triage, compliance checking, research synthesis. Every agent action logged; every claim carrying an evidence state.
03
Decision Support Tooling
Systems that put a number in front of a decision-maker and can defend it — registers, exposure engines, portfolio views, board packs. Citadel is the one we built for ourselves; the pattern generalises.
04
Model Evals & Red Teaming
The infrastructure that tells you whether a model does what someone claims, in the conditions it will actually meet — and the adversarial programme that attacks it on purpose before somebody else does it for free.
Who builds it
Two founders. Seventy-one agent seats.
One human decision.
An AI board governs — four officers and nine engineering seats, working enterprise briefs. Each founder then runs their own lab dev team of nine, orchestrated through Claude and Codex, and their own four delivery streams inside the Hermes AI Factory, where every stream carries the same five stages. It all runs on Google Cloud Platform and Vertex AI, with one audit trail across the whole estate. Nothing is deployed, published or purchased by an agent: a named human answers for what leaves the building.
- AI BoardFour officers — CFO, CTO, COO, CMO — over nine engineering seats: Eng, AI/ML, Data, DevOps, API, QA, Security, Docs and UI/UX. This is the tier that works enterprise briefs.
- The foundersDeclan and Austin, each orchestrating their side of the estate through Claude and Codex. Every gate that matters closes here.
- Lab dev teamsNine seats each, one per discipline, separate from the board and pointed at lab development rather than client delivery.
- Hermes AI FactoryFour delivery streams per founder, eight in total, each staffed across the same five stages: orchestrate, plan, build, review, document.
- One platformGoogle Cloud Platform and Vertex AI under all of it, and a single audit trail across it — every agent action logged, every claim carrying an evidence state.
How a build gets commissioned
Nothing on this floor starts without a decision behind it.
If the build/no-build question is still open, a Feasibility Sprint answers it in two to three weeks and can honestly recommend not building. Once it is answered, a Production Build is the contract shape — 6–12 weeks, fixed scope, fixed fee. A workforce deployment runs longer at 8–16 weeks because your operating model has to change with it.
Past handover, a Lab Retainer buys standing capacity for evolution and eval maintenance, and Team Enablement leaves the discipline inside your own engineers.
If it can be bought, buy it
We build what does not exist yet. Where the market already sells the thing properly, we will say so and help you choose — a Build, Buy or Partner call costs a fraction of a build you did not need.
Tell us what you are trying to build.
Thirty minutes with a founder. We will tell you how we would decompose it — before you commit to anything. If it is outside the four classes, we will say that too.
Price agreed before we start · Founder-led · NDA available