← All insights

Architecture

Why AI pilots stall at the fourth one

The first pilot is easy because nothing depends on it. The fourth is hard because everything does — and by then the architecture decision has already been made by accident.

6 min read

A first pilot succeeds on enthusiasm. It is scoped small, staffed with volunteers, and measured generously. Nothing else in the business depends on it, so nothing else constrains it.

By the fourth, the situation has inverted. Each pilot has chosen its own model provider, its own retrieval approach, its own evaluation method, and its own place to put credentials. None of that was a decision anybody made deliberately — it accumulated. The cost of the fifth use case is now higher than the first, which is the exact opposite of what a platform is supposed to do.

The diagnostic question is not how the models are performing. It is what the marginal cost of the next use case looks like, and whether anyone can answer that without a week of investigation. If the second use case cost more than the first, the program does not have an architecture — it has a collection.

The fix is rarely to rebuild everything. It is to pick the strongest use case, rebuild that one properly as a reference implementation on shared infrastructure, and make every subsequent build reuse it. That is slower for one quarter and considerably faster thereafter.

Next step

Bring us the constraint, not the brief.

The first conversation is diagnostic: what is actually blocking the outcome, and whether we are the right people to unblock it. If we are not, we will say so.