Twelve questions a board should ask before approving an AI project

Approve AI projects with written answers, measured evidence, clear ownership, data controls and a plan for failure.

Boards are being asked to approve AI projects because a demonstration looked good. That is weak evidence. A demonstration is the least informative evidence available, because it is a curated sample chosen by the person seeking approval. It tells you nothing about the ninety-ninth case.

These are the twelve questions we would want asked about any AI project, including our own. They are deliberately non-technical. Each question has a good answer and a bad one. You can recognise the bad answers without knowing anything about machine learning.

1. Ask for the exact output in writing

The first answer should make clear what the system will produce. If the answer describes a general capability, nobody has specified anything and there is nothing to hold anyone to.

  • Good: a specific deliverable with a defined shape. Eleven fields extracted from a supplier invoice with a confidence score on each.
  • Bad: a capability. It will help the team handle invoices more efficiently.

2. Ask how the team will measure whether it worked

A good answer measures correctness against data the system has never seen. That measurement should come from someone who knows the domain, and it should have a number attached.

  • Good: a measurement against data the system has never seen, produced by someone who knows the domain, with a number attached.
  • Bad: user feedback, adoption rates, or a pilot that felt positive.

User feedback, adoption rates and positive pilots measure enthusiasm. Correctness remains unchecked, and enthusiasm reliably precedes the discovery that correctness was never checked.

3. Ask what happens when the system is wrong

Every probabilistic system, one that works in likelihoods rather than certainties, is wrong sometimes. A project that has not designed for that has simply decided the error will be somebody else’s problem later.

  • Good: a described path. Below a confidence threshold it declines and routes to a person. Above it, errors are caught by this specific downstream check.
  • Bad: it is highly accurate.

High accuracy does not answer the question. The board needs to know what happens after a wrong output is produced.

4. Ask who is accountable for any output that reaches a customer

When an AI output reaches a customer, a person needs to be accountable for it. If no person’s name attaches to an output, nobody can be asked to justify it, and eventually someone will be.

  • Good: a named role, with a record showing they accepted it.
  • Bad: the system, the vendor, or a diffuse reference to the team.

5. Ask what data leaves your environment and who receives it

This question is about who receives the data. Security language about how the data travels does not answer it.

  • Good: a named provider, under a signed agreement, processing in a named jurisdiction, with a documented retention position.
  • Bad: it is secure, or it uses enterprise-grade encryption.

6. Check whether the system can expose content a user is not entitled to see

This is the most common serious defect we find. The business risk is simple: a user may see information they should not see.

  • Good: permissions are applied at retrieval, using the same authorisation source as the underlying system, before anything is generated.
  • Bad: we filter the results, or only authorised users have access.

Filtering after generation is different from applying permissions before retrieval. Access to the tool is also different from entitlement to the content it can reach.

7. Ask for the cost per unit of work at real volume

The cost needs to match the process being replaced. A monthly platform figure or a per-seat price cannot be compared against that process, and that comparison is the only one that decides whether the project is worth doing.

  • Good: a figure per invoice, claim or ticket, projected at production volume including the awkward inputs, the retries and the human review queue.
  • Bad: a monthly platform figure, or a per-seat price.

8. Ask what happens if the model changes underneath you

In a business process, a system whose output can change without notice is an uncontrolled dependency. The answer should show how a behaviour change will be detected.

  • Good: versions are pinned, and the evaluation runs on a schedule so a behaviour change is detected within a day.
  • Bad: the provider handles that, or we always use the latest.

9. Ask how the team turns it off

Turning the system off requires more than rolling back a deployment. A rollback answers a different question, and it does not address the outputs already produced and consumed.

  • Good: a flag, a defined fallback, and a drill that has actually been run.
  • Bad: we would roll back the deployment.

10. Ask what happens if it has been wrong for three weeks

This is the single most revealing question on the list. Almost nobody has considered it before launch, and it takes twenty minutes to settle in advance versus days during an incident.

  • Good: we can identify every affected output and where it went, and there is an agreed position on correction and notification with a named decision-maker.
  • Bad: a blank look.

11. Ask what the system is replacing and what that costs today

A project needs a measured baseline. Without a baseline there is no way to evaluate the result, and projects without baselines are declared successful by default.

  • Good: a measured baseline. Three people, this many hours, this error rate, this much rework.
  • Bad: an assumption that the current process is expensive without anyone having measured it.

12. Ask what would make the project stop

A stopping condition should be agreed before starting. A project with no stopping condition does not stop; it acquires momentum and becomes something nobody can question without appearing to be against progress.

  • Good: a defined threshold agreed before starting. If accuracy on the held-out set is below this, or the escalation rate exceeds that, we stop and reconsider.
  • Bad: no answer.

A demonstration hides the cases that matter

The standard evidence is weak, and boards are shown demonstrations constantly without being told what they are looking at.

A demonstration is a sample of size one to five, chosen by the party seeking approval, from a distribution they have seen and you have not. Every incentive points toward selecting cases the system handles well, and this happens without any dishonesty. The person demonstrating believes those cases are representative, because they are the cases they have been working with.

The cases that matter are the ones nobody chose to show: the malformed input, the ambiguous phrasing, the document in the wrong format, the request that should have been refused. Those are where cost and risk live, and they are absent from every demonstration by construction.

The alternative evidence is cheap to ask for. Ask for a measured pass rate on a held-out set, with the failures shown rather than the successes. A held-out set is data kept back from the build so the test has not already been practised. Any team that has done the work can produce it in an afternoon. Any team that cannot has not done the work.

Two answers should stop approval

Most weak answers deserve a follow-up. Two deserve a decision.

"We will work that out during the pilot." Applied to accountability, permissions or the failure path, this means the project intends to discover its own governance after it is running. Pilots become production through momentum rather than decision, and the questions get harder to ask the longer the thing has been working.

"The vendor handles that." Applied to data handling, model changes or correctness, this is a delegation of responsibility that does not survive contact with a regulator or a customer. The vendor’s obligations are whatever the contract says, and almost nobody proposing an AI project has read it closely enough to know.

Both answers can appear in a project that may be sound. Both leave it short of approval.

Use the questions before the meeting

Send the list ahead rather than asking cold. That makes the project team do the thinking, and most of the value is generated before the meeting.

Ask for written answers. A confident verbal response to question six sounds fine and is much harder to give in writing, because writing forces precision about who checks what and when.

Revisit the answers at the point the pilot becomes production, which is the transition where governance is most often skipped. The answers that were adequate for fifty documents a month are frequently inadequate for forty thousand, and nobody re-asks because the project is already considered approved.

The answers separate a specified project from enthusiasm

Read together, the twelve separate two kinds of project.

One has been specified: somebody has written down what it does, how it is measured, what happens when it fails, who is responsible and what would make them stop. It may still be a bad idea, but it can be evaluated as a business proposition.

The other is an enthusiasm. It has a demonstration, a vendor and a sense that something must be done. It will be approved, will produce activity, and in eighteen months nobody will be able to say what it achieved because nothing was ever measured.

The questions do not require technical knowledge to ask or to assess. That is deliberate. Boards have been assessing supplier risk, process change and operational controls for decades, and this is the same exercise with an unfamiliar noun.

Different people should own different answers

A practical point saves a great deal of circular discussion: these questions have different owners. Sending all twelve to the technical lead produces confident answers to four of them and guesses at the rest.

Questions one, eleven and twelve belong to whoever owns the business process. What it produces, what it replaces and what would make us stop are commercial judgements, and a technical team answering them is inventing a business case.

Questions two, three, eight and nine belong to engineering. Measurement, failure paths, version control and the off switch are build decisions with build answers.

Questions four, five, six and ten belong to risk, compliance or legal, informed by engineering. Accountability, data movement, entitlement and the three-week scenario are governance questions that engineers should not settle alone, and frequently do because nobody else asked.

Question seven belongs to finance, with numbers from engineering. Cost per unit against the current baseline is an ordinary investment appraisal wearing unfamiliar vocabulary.

Distributing them that way tends to surface disagreements early, which is the point. The projects that go wrong are usually ones where four functions each assumed another had thought about question ten.

The three-week error question tests the whole design

If there is only time for one, ask the tenth: if this has been wrong for three weeks, what do we do?

Answering it requires having thought about detection, lineage, blast radius, correction and notification. Lineage means the record of every affected output and where it went. A team that answers it well has almost certainly answered the others. A team that cannot has usually built a demonstration.

These questions apply to our own proposals

These questions apply to our own proposals and we would rather be asked them than not. Our answers are that every engagement carries a written scope and a definition of done, an evaluation set built before implementation, a refusal path rather than a guessing one, named human accountability on anything that reaches a customer, and a rollback drill executed before go-live.

Where we cannot answer well, that is worth knowing too. On a genuinely novel problem we may not be able to give a defensible accuracy figure at proposal stage, in which case the honest structure is a bounded discovery increment that produces the evaluation set, priced separately, before anyone commits to building.

A supplier who answers all twelve confidently on day one about a problem nobody has solved before is telling you something about their relationship with certainty.

More on our commitments in the trust centre, and on how engagements are structured on engagement models.

Written by Brilliant Systems

Our engineers write these between projects. If something here is relevant to a decision you are making, we are happy to talk it through without it becoming a pitch.

Certified, partnered and awarded