How we work · 02
One concrete problem.
Tested against a harness
we agreed first.
We map the problem to a named decision the model output must feed, build a working proof of concept against an agreed evals harness, and stress-test the retrieval, reasoning or generation pipeline with adversarial cases. If the answer is don’t build this, you hold the evidence that says so.
You hold: the running proof of concept, the evals harness and a written build/no-build recommendation.
The question this shape exists to answer
Should this be built at all? Answered in weeks, for a fixed fee.
The expensive mistake in AI is not a failed build. It is a successful build of something that was never going to be worth running — discovered eight months and a capital allocation later. A Feasibility Sprint puts a decision gate in front of that, and prices the gate low enough that using it is obviously rational.
The mechanism is the harness. Before anything is built we agree how we will know whether it works: golden cases with the right answers settled by whoever owns the process, adversarial cases designed to break it, a latency budget, and a cost ceiling. Then the proof of concept gets built and measured against that, not against a demo script.
You keep the harness either way. That matters more than it sounds: if you later run a procurement, the harness is how you test a vendor’s claim against your own cases rather than watching a demonstration they controlled.
Adversarial from the start
Retrieval that returns nothing. Inputs designed to mislead the reasoning. Injection through content the system ingests. A proof of concept tested only on cases someone chose to show you is not evidence of anything.
How it runs
Harness first. Then build. Then try to break it.
Two to three weeks, fixed. The output is a decision, supported by something running.
- Duration2–3 weeks, fixed
- FeeFixed, quoted in writing before work begins
- MethodEvals-first
- GateA written recommendation either way
- ProducesA running proof of concept and its harness
Map the problem to a decision
What decision does the model output feed, who makes it, and what would “good enough” mean to them. Without that, accuracy targets are arbitrary.
Agree the harness
Golden cases and their correct answers, adversarial cases, latency budget, cost ceiling. Agreeing golden answers often surfaces real internal disagreement — better now than in production.
Build the proof of concept
Enough system to answer the question honestly. Not a prototype dressed up for a demo, and not a production build in disguise.
Stress it
Adversarial cases against retrieval, reasoning and generation. Documented, including the attempts that failed to break it.
Write the recommendation
Build or do not build, with costs, risks and a path to production if the answer is build. A negative recommendation is a completed engagement.
The two lists that matter
What you hold, and what we will not do.
Both are in the engagement letter before you sign it. The second list is the one worth reading twice — it is where most disappointment in this market actually comes from.
What you hold at the end
- The running proof of concept
- The evals harness — yours to keep, whichever way the answer goes
- A written build/no-build recommendation, with costs and risks
- A path to production, where the recommendation is to build
- The adversarial findings, including what did not break it
What we won’t do
- Production systems without evals
- Chatbots without guardrails or monitoring
- Proofs of concept with no path to production
- Projects where the problem is not understood yet
- Engagements that would need ten or more people to deliver
If the answer is build
The harness carries straight into the production build.
A Production Build takes six to twelve weeks from green light, and the harness you already own becomes the CI gate on every commit. That continuity is why the sprint is worth running even when everyone is fairly sure of the answer.
How production builds run →Also the way into the Lab
Where the question is genuinely unanswered — not just unanswered here — the Feasibility Sprint is the usual route into a Lab engagement, under NDA.
Should you build it at all?
Thirty minutes with a founder. Tell us the problem and we will tell you what the harness would have to contain to answer it honestly.
Fixed fee · Quoted in writing before we start · NDA available