How we work · 03
Built far enough
to answer honestly.
And no further.
The artefact a Feasibility Sprint produces: one use case built to the point where the build/no-build question can be answered, and deliberately no further. Working code against an agreed evals harness, attacked with adversarial cases, with a measured cost per transaction and a written call either way.
You hold: the running proof of concept, the evals harness and a measured cost per transaction, not an estimate.
Where most proofs of concept go wrong
Built to persuade, or built to answer. It cannot be both.
A proof of concept built to persuade gets demonstrated on cases that suit it, and its costs get discussed in round numbers. It succeeds, everyone is pleased, and the production build then discovers the awkward parts at ten times the price. That is the normal life cycle of AI pilots in large organisations, and it is not caused by incompetence — it is caused by what the artefact was for.
This one is built to answer. Which means it is deliberately not polished, is measured against cases agreed before it existed, is attacked with adversarial inputs, and carries a real cost-per-transaction figure rather than an estimate. It is also deliberately stopped once the question is answered: continuing to build past that point is spending your money to make a decision that has already been made.
The honest negative is the reason to buy this. A written, evidenced “do not build this” is worth considerably more than the fee, and it is contracted as a completed engagement rather than as a failure that needs explaining internally.
What “far enough” means
Far enough that the harness gives a real reading, the failure modes are visible, and the cost per transaction is measured rather than modelled. Not far enough to be production — no monitoring estate, no rollback rehearsal, no handover pack. Those belong to a Production Build.
How it is judged
Five tests a proof of concept has to pass before it counts as evidence.
The method is the Feasibility Sprint’s, and it is described there. What belongs here is the standard the artefact is held to when it lands on your desk.
- EngagementFeasibility Sprint — 2–3 weeks
- MethodEvals-first
- GateA written recommendation either way
- Not includedMonitoring estate, handover pack, production hardening
- Next stepProduction Build, if the answer is build
Measured against cases agreed before it existed
Golden cases with answers settled by the process owner, and adversarial cases designed to break it. A proof of concept scored on cases chosen after the fact has been marked by its own author.
Attacked, not demonstrated
Empty retrieval, misleading inputs, injection through the content it ingests. The write-up records what held as well as what failed; a demo shows you only the first.
Costed per transaction, from real calls
Model calls, retrieval and tokens as actually incurred, not a spreadsheet estimate. If the number would be embarrassing at production volume, this is where you find out.
Stopped at the answer
Rough edges left rough on purpose. Building past the point where the question is answered spends your money on a decision already made — and produces a prototype someone will be tempted to ship.
Written up either way
Build or do not build, with costs, risks, failure modes and a path to production where the answer is build. The negative is a completed engagement, not a failure to be explained.
The two lists that matter
What you hold, and what we will not do.
Both are in the engagement letter before you sign it. The second list is the one worth reading twice — it is where most disappointment in this market actually comes from.
What you hold at the end
- The running proof of concept, as built
- The evals harness, and the golden and adversarial case sets
- A measured cost per transaction, not an estimate
- The failure modes found, and the attacks that did not work
- A written build/no-build recommendation with a path to production
What we won’t do
- Proofs of concept with no path to production
- Demos tuned to cases we selected
- Continuing to build after the question is answered
- Cost figures modelled rather than measured
- Calling a positive result the only successful outcome
What happens to it afterwards
A proof of concept is not a production system, and we will not pretend otherwise.
If the recommendation is build, the production build starts from the harness rather than from the prototype code — architecture with failure-mode analysis, evals in CI, monitoring tied to business impact, rollback written before the first commit. Promoting a proof of concept straight into production is the single most common way this class of system fails.
See the Production Build →The pilot nobody closed
Proofs of concept that quietly became load-bearing are a category we find in almost every estate review — no owner, no monitoring, frequently in front of customers. Stopping properly is part of the discipline.
Test it before you fund it.
Thirty minutes with a founder. Tell us the use case the decision turns on, and we will tell you what two to three weeks would actually establish.
Fixed fee · Quoted in writing before we start · NDA available