Why we build AI-first, and what that rules out.
A test for AI-first that a buyer can actually apply, what it rules out, and the place where our own product does yet pass.

Of all the AI-first pitches you sat through this year, how many of those claims could you test? Probably none. It’s no fault of yours; the bar to pass isn’t very high but it can be hard to see if you don’t know what to look for.
Therefore, I would like to give you something you can use to test the next one, including on us, if we pitch to you. If you think that this might cost me more than it costs you, you’re right. And that’s why it’s worth to continue to read.
The test
First, take a capability. Remove the AI from it. What’s left?
If what remains is the same outcome you wanted, just arriving more slowly, then AI was an accelerator. It’s useful, often worth paying for, and not just architecture. By architecture I mean a decision the rest of the system is built around, one you cannot reverse later without rebuilding.
If what remains is a different outcome, or no outcome at all, then AI is load-bearing, meaning the capability collapses without it. That is the version of AI-first that is worthy of the name.
Here’s an example that fails, and I want to pick one we would be tempted by.
A document summariser sitting on top of a claims queue. It saves real time, adjusters like it, and I would happily build it. Switching it off changes how long the work takes and nothing else. By definition, this marks it as a feature, not architecture.
Where my own product currently does not pass
I would like to give you a breach-risk scorer as the pass: something that runs ahead of a service-level deadline, predicts which claims will miss it, and warns the person accountable while there is still time to act.
There are two problems with that, and I would rather put them in this section than bury them.
The first is that it is not built. It is specified, it is the fifth of five AI surfaces our claims architecture reserves, and it is last in the delivery order. Reserving is not shipping. As of now, none of the five is live.
The second is more interesting, because my own test argues against my own example.
A rule-based check can trigger at any fraction of the allotted time, not only at the deadline. Eighty percent elapsed and the step still open is a perfectly good early warning. It is deterministic, cheap, and auditable. My own next section says a rule that can be read and checked beats a model that is neither. So the threshold rule is the baseline the scorer has to beat, on calibration and on expected time-to-close. It is not something the scorer obviously improves on.
I am keeping the example because that is what applying a test to yourself looks like when the test is real.
What it rules out
The constraint costs three things.
It rules out shipping a capability whose AI can be switched off without anyone noticing, because that tells me the design started somewhere else and the AI arrived afterwards.
It rules out AI where the honest answer is a rule. If a region gets 4 hours to accept a claim, that is a number somebody wrote in a table. It is not something a machine needs to predict, and in front of a regulator asking why a claim was handled as it was, a written rule is a better answer than a model.
And it rules out claims we cannot check. An AI capability in a regulated operation has to arrive with a way to evaluate it, so its performance is measured rather than asserted. That is the expensive one, and it is why these surfaces ship in stages rather than being announced together.
What reserving them actually changes
If none of the five is live, you would fairly ask what the AI-first claim rests on today. The answer is narrower than the phrase implies, and personally, I think it is the part that is genuinely hard to add later.
Every one of the five has to write into the audit log on the same terms as a human: what it decided, how certain it was, which model version decided it, and what evidence it drew on. Designing that in from the start changes the shape of the log, the identity model and the operator portal. Retrofitting it means changing all three at once, in a system already carrying live claims.
So the defensible claim is not that our AI agents are working today. It is that the system they will land in was built to hold them accountable.
That is a decision, not a feature. It’s architecture.
The same test on our own delivery
Our engineers direct agents through the development lifecycle rather than treating them as a typing shortcut. For a team this size, it’s how we are able to ship rapidly.
And if I were to apply my own test to that?
We take the AI agents away and the same team ships the same software but more slowly. That is acceleration, not architecture. By my definition our delivery process is not AI-first, it is AI-accelerated, and I would rather say so than let the two blur together because the second sounds better.
What to do with the test
Use it on the next vendor who sits in front of you, and use it on us.