Capital · coordination · constructionCareers

Technical · 13 min · reviewed August 2026

Evaluating a model before a pilot

A test suite built from labelled history, run before anyone sees a demo — and an honest account of what a passing score does not prove.

Why before

A pilot that starts without an evaluation suite has no way to distinguish a good system from a persuasive one, and the demo will be persuasive. Building the suite first also settles an argument that is otherwise deferred indefinitely: what counts as a correct answer, decided by the business rather than by whoever is holding the model.

Building the set from labelled history

Take decisions the institution has already made, with their outcomes, across a window long enough to include a bad quarter. Hold out a slice nobody may look at. Where the labels are contested — and in exception handling they usually are — the process of resolving them is itself the most valuable output of the engagement, because it produces a written decision policy where there was a convention.

The three numbers

  • The error rate the business will acceptAgreed in advance, in writing, by the person accountable for the process. Not chosen after seeing the results.
  • The cost asymmetryWhat a false positive costs against a false negative. These are rarely equal and the threshold should reflect it.
  • The coverageThe share of cases the system is allowed to handle at all. Narrow and reliable beats broad and supervised.

What a passing score does not prove

It does not prove the system will behave the same on next quarter’s distribution, that the corpus will still be current, or that the humans around it will keep reviewing what they are meant to review. Those are operating properties, and they are measured after deployment or not at all. The suite is the entry condition for a pilot, not evidence that the pilot will succeed.

Terms used here

Author

Name pending · practice lead. Reviewed by the editorial owner.

Cite

MLG Blockchain, “Evaluating a model before a pilot,” 2026. TechArticle, machine-readable. https://mlgblockchain.com/insights/ai-transformation/evaluating-a-model-before-a-pilot

Prints cleanly, with URL and date in the running head.