Examiners and adverse-action lawsuits don't want 'explainable'. They want defensible. The difference is the audit trail: versioned system prompts, prompt-data libraries, model cards, decision logs, and the documentation gap most lenders only notice when regulators arrive. A field-tested checklist before you scale a pilot.

Every AI underwriting pilot we have seen reaches the same meeting. The model scores well, the credit team likes the summaries, and someone from risk asks the question: if an applicant is declined on this output, can we show why? The answer that comes back is usually 'the model is explainable'. That is not wrong. It is just not the question.

Examiners, and the lawyers who bring adverse-action claims, do not want an explanation. They want a record: which version of the system produced the decision, what it was told, what it had been checked against, and who signed it off. Explainability is a property of a model. Defensibility is a property of an organisation, and it is built almost entirely out of paperwork the pilot never produced.

We build AI systems on spec for lenders and underwriting platforms, and the audit trail is the part of the brief that arrives late, if at all. Where a point below turns on how a rule applies to your book, take legal advice. We are engineers.

What is the difference between explainable and defensible?

Explainable means you can produce a plausible account of why the model said what it said: the features that weighed heaviest, the rationale it wrote. Useful, and the vendor demo will show it. But an explanation generated after the fact by the same system that made the decision is testimony from an interested witness.

Defensible means a third party with no goodwill towards you can reconstruct the decision from records you kept at the time: the configuration the system ran under, the inputs it was given, the checks it had passed, and the human accountable for it. An examiner wants to see a controlled process. A claimant's lawyer wants to find the moment it was not one. Both read the same trail, and neither is impressed by a rationale paragraph. Explainability is a feature you buy or prompt for. Defensibility is a discipline you run, and the model is the least interesting component in it.

An explanation generated after the fact by the same system that made the decision is testimony from an interested witness.

What does a defensible audit trail contain?

Five artefacts, only useful together. A lender holding three of the five has a partial record, and a partial record is what cross-examination is for.

  • Versioned system prompts. Every prompt that has ever run in production, with a version identifier, the date it went live and the name of whoever approved it. The prompt is the policy; if you could not produce the credit policy that applied on the day, you would not call the decision defensible.
  • A prompt-data library. A curated bank of reference cases that every prompt change is tested against. For an AI-native underwriting platform we worked with, that library held 84 reference items across 21 underwriting dimensions, and no prompt change reached production without running against all of them. It turns 'we tweaked the prompt' into 'we changed the prompt and here is the evidence it still behaves'.
  • Model cards. One per model version: what it is, why it was chosen, its known weaknesses on your data, when you adopted it and when you retired it. When a provider deprecates a model, the card is the record that the migration was a decision and not a drift.
  • Decision logs. For every decision the system contributed to: prompt version, model version, the inputs, the output, and any human who reviewed or overrode it. Stored so they cannot be quietly edited, kept for as long as the decision could be challenged. This is the artefact examiners ask for first and the one most pilots cannot produce.
  • Evidence of evaluation. The reference-library results over time, tied to prompt and model versions. Not a dashboard screenshot; the underlying records. If output quality moved, you should be able to say when, by how much, and what was done.

Where is the documentation gap?

Always the same place: the pilot was built to prove the model works, so the team instrumented the model and not the process. Prompts lived in a repository, which is not the same as knowing which commit was live on a given day. The reference cases lived in a spreadsheet anyone could edit. Model upgrades happened because the provider sent an email. Decisions were logged as outcomes, without the inputs and configuration that produced them.

None of that is negligence; it is what an experiment looks like. The gap opens when the experiment becomes the system of record and the documentation never catches up. Lenders notice at one of three moments: a regulator's request for information, an adverse-action challenge on a specific file, or due diligence before a funding round or an acquisition. All three arrive with a deadline, and a trail reconstructed after the fact is expensive when it is possible at all.

The fix is cheap at the start. The engagement above, a system-prompting strategy plus the prompt-data library, took two weeks from kickoff to a production-ready system. That is not a large number against the first request for information you cannot answer.

What should you have before scaling a pilot?

The checklist we would want signed off, by name, before a pilot moves from supervised use to something the business leans on every day.

  1. A named owner for the system prompt, with approval authority, and a versioned history that can answer 'what was live on this date' in minutes.
  2. A prompt-data library of reference cases spanning every underwriting dimension the system touches, held under change control, run in full before any prompt or model change ships.
  3. A model card for every model version in use, including the retired ones, with the reason each was adopted and the date each was replaced.
  4. A decision log that captures configuration and inputs, not just outcomes: prompt version, model version, the context the model saw, the output, and any human review.
  5. A written account of where humans sit in the loop: which decisions the system makes alone, which it recommends, and who is accountable for each. If the answer is 'it depends', the pilot is not ready.
  6. A retention and access policy for all of the above, agreed with whoever owns your regulatory relationships and, for adverse-action exposure, with counsel. Take legal advice on the periods; we can only tell you the records must exist.
  7. A drift process: what happens when evaluation results move, who is told, and what triggers a rollback. A trail that shows you noticed and acted is good; one that shows you noticed and did nothing is worse than none.

If more than two of these are missing, scaling the pilot is not a technical decision. It is a decision to build the audit trail later, under pressure, from fragments.

What to do next

Start with the decision log, because it is the one artefact you cannot backfill: a decision not recorded with its configuration at the time cannot be reconstructed later. Then put the system prompt under real version control with a named approver. Then build the reference library, one dimension at a time, and refuse to ship a prompt change that has not run against it. Model cards and the drift process follow once those three exist.

That is the order we work in. A 20 or 45-minute discovery call to find out which of the five artefacts exist and which are assumed. Where it makes sense, a four-day Spec from €5,000 that turns the checklist into a written brief for your own system, with an owner against each piece. From there, a fixed-scope, fixed-price build; the average ships in eight weeks. We build on spec and do not sell a product, so the trail you end up with is yours, in a shape an examiner can read.

Or skip ahead and talk through it directly