← who we help
Client under NDA
Building safety, regulated · Explainable AI

Teaching software to know when not to answer

Five days of specialist work now takes under twenty minutes, with a regulator-ready record behind every decision the client could not produce before. Standing start to production in under two weeks, by a single senior builder, at roughly a quarter of the cost of a traditional build.

Client
Under NDA
Sector
Building safety, regulated
Service line
Build →
Engagement
Fixed-price product build
Timeline
Under two weeks to production
Team
One senior builder

By the numbers

Under 20 minper schedule, down from five days
95%match with the client's own experts in testing
Under 2 wksfrom a standing start to production
~¼ the costof a traditional engineering build

The problem

Some industries cannot tolerate a confident wrong answer. In safety-critical, heavily regulated work, choosing the wrong option can put people at risk and breach the law; the cost is the least of it.

Our client works in building safety: the trade of making sure a building is put together so it protects the people inside it. Their specialists spend their days matching detailed technical requirements to the single correct option from a large catalogue of approved, independently tested and certified solutions. Then they price each one against a fiddly model of banded rates and stacking adjustments.

A day looks like this. A job lands as a dense, sprawling spreadsheet, one row for every point in the building that has to be made safe. A single job can run to several hundred rows. For each row a specialist reads off the surrounding construction, the service passing through it, the exact dimensions and the safety rating it has to meet. Then they find the one certified solution independently tested for that precise combination, and price it. Every row is a slightly different mix of a dozen interacting variables, and a building-safety rule sits behind each one.

It was painstaking work. It relied on a handful of very experienced people, it was slow, and because it lived mostly in someone’s head and a spreadsheet, it was almost impossible to show your working afterwards. That last point matters when the regulator expects a documented trail behind every decision.

The client wanted to automate this without losing the judgement that made their people trustworthy. Speed was never the hard part. The hard part was that a wrong match has real consequences, so the system could not be allowed to bluff. It had to give the right answer or admit it did not have one, and it had to leave a record either way.

What we did

The obvious approach today is to hand the whole problem to an AI model and trust what comes back. We did not do that, and that decision is what made the project work.

We started by getting the expertise out of people’s heads and into rules. Most of these decisions are not ambiguous at all to an expert; they follow firm logic. We captured that logic in a deterministic engine that narrows hundreds of options down to one by applying the same checks a senior specialist would, in the same order. Same input, same output, every time.

AI only enters when the rules genuinely cannot separate two or more valid options, and even then it is kept on a short leash. It gets the shortlist and the rules, not free rein, and it has explicit permission to say it is not sure. When it lacks confidence, the item goes to a human instead of being forced into an answer.

We built a system whose default, when in doubt, is to stop and ask rather than guess. In most software that would be a failure state. Here it is the most important feature, because one quiet wrong answer does far more damage than a hundred honest “I don’t knows”.

On a live job, it works like this. A schedule arrives as a dense spreadsheet of several hundred rows. For every row it is sure about, the system reads the inputs, then picks the single certified solution that fits the construction, the service and the rating. It prices that on the rate card and writes back a matched, priced, fully reasoned answer. Anything it cannot separate with confidence, or that is missing something it needs, it sets aside for a specialist to finish. The specialist opens the job and sees only those rows, each with the system’s working laid out, not a blank sheet of several hundred.

None of this works without the data underneath it, and that was a large part of the job. The catalogue alone runs to more than 600 independently tested products. Much of its detail sat locked in unstructured PDFs and scattered reference material that no system could query. We pulled all of it, with the inconsistent free-text on every incoming schedule, into one clean, structured database the engine could reason over. On top of that, we recorded every decision with its inputs, its reasoning and a confidence level. The audit trail builds itself as a by-product of normal work, and the system shows the working behind every answer. That is exactly what could not be done before, when the knowledge lived in someone’s head and a spreadsheet.

The result

Work that used to take a specialist five days now runs in under twenty minutes, and the experienced people who used to do it spend their time on the genuinely tricky cases rather than routine lookups.

We saw what that did for the team. Freed from grinding through these schedules by hand, they now turn around far more quotes in a week than they used to. The small data-entry slips that creep in when a person keys several hundred rows into a complex spreadsheet have largely gone. The machine does the careful keying; the people bring the judgement.

The bigger win is trust. The same request always produces the same result, and every answer carries a complete record of how it was reached, which is what the regulator wants to see. In extensive testing the system reached the same conclusions as the client’s own experts 95% of the time. The remaining one in twenty were extreme edge cases, which is exactly what the system is built to hand to a person rather than guess on. None of this is tied to the client’s particular catalogue, so the same approach can extend to new sources and adjacent problems without being rebuilt.

The second story is how quickly that value arrived, and how little it took. We delivered the whole thing with a single senior builder, in production in under two weeks, for roughly a quarter of what a traditional engineering project would have cost. By the point a conventional build would still be agreeing its specification, this one was already doing the job. Fast and cheap usually means corners cut, which in work like this would be disqualifying. Here the speed came from doing the thinking properly up front, in the rules, rather than skipping it.

Why it matters

Most “AI automation” works by always producing an answer. In high-stakes work, that is the dangerous part. The unusual thing here was building restraint into the system. Hard-won expert rules do most of the work; AI is allowed in only where true ambiguity remains, and abstaining is always an acceptable answer. Software that knows the limits of its own knowledge turns out to be far more useful, and far safer, than software that always has something to say.

Restraint is one half of that. Transparency is the other. Most AI hands back an answer but cannot tell you why, because the reasoning sits inside a model no one can interrogate. This system works the other way around. Every answer carries the inputs, the reasoning and the confidence behind it, so anyone can see how it got there and challenge it if they need to. That is what explainable AI is meant to deliver, and in regulated work it is not optional: a regulator will not trust a system that cannot show its working.

Technical card

AI
Claude (Sonnet 4) via the Anthropic API, used only where the deterministic rules cannot separate valid options, with explicit permission to abstain
Cloud
Google Cloud: Cloud Run with Cloud SQL
Backend
Python 3.12, FastAPI
Data
PostgreSQL, SQLAlchemy, Alembic migrations
Documents
PDF and spreadsheet ingestion, generated PDF output
Explainable AI
Every decision recorded with its inputs, reasoning and confidence level; the same input always produces the same output, so any answer can be reproduced and challenged
Quality
Full automated test suite, static typing

Some details have been generalised to protect client confidentiality.

A problem like this?

If this is close to where you are, let's talk.

or email us directly at [email protected]

Ask LevelFive

What would you like to know?

Ask anything about how we work, what we build, who we help, or how an engagement runs. Answers come from this site.

Enter to send, Esc to close. Answers are generated and can be imperfect. For anything specific, talk to us.