Angulith

June 3, 2026

Evaluation is the product

A fluent demo is not a system. If you cannot rerun the test when the model or the corpus moves, you shipped a conversation, not a capability.

The pattern is familiar. A vendor or an internal lab shows a chat window over the company corpus. Leadership is impressed. Someone asks about accuracy. The room agrees to “keep an eye on it.” Three months later the corpus has grown, the model has been quietly upgraded, and nobody can say whether the thing is better or worse than the week it was praised.

Evaluation is not a research nicety. It is the product. Before we put a model in a workflow, we want a baseline a skeptic would accept: a fixed set of tasks, a definition of wrong that matches the cost of being wrong, and a way to rerun the set in the same pipeline that deploys the feature. If that sounds heavy, the feature is not ready.

The other half of evaluation is the fallback. Models will be wrong, rate-limited, or deprecated. A clerk, a dispatcher, or a nurse needs a path that does not require the model to be up. If the only path is “try again,” you have built an outage into the operation.

We have declined engagements whose success metric was executive delight in a conference room. Delight is cheap. A measured task in the system of record is not. If you cannot name the user, the task, and the cost of a bad answer, start there. The model can wait.