Skip to content
LUNTA

Own the evaluation, rent the model

LUNTA · · 2 min read

Every quarter a better model ships, and somewhere an enterprise rewrites its roadmap around it. The organisations that come through this churn intact share one property: they own the thing that decides whether any model is good enough for their operation. That thing is not the model.

The model is a component, not the system

In a production AI system, the model is the part you exchange most often and control least. The parts you keep are the ones nobody demos: the retrieval layer over your data, the integration into your systems of record, the governance controls — and above all, the evaluation suite that measures whether the whole assembly does the job to the standard you set. Enterprises routinely buy the replaceable part and improvise the durable one.

The consequence surfaces at every model release. An organisation without its own evaluation asks the only question it can answer: ‘is the new model better on the vendor’s benchmarks?’ An organisation with one asks the question that matters: ‘does it clear our thresholds, on our data, on our failure modes, at our cost ceiling?’ The first organisation upgrades on faith. The second upgrades on evidence — or declines, on evidence.

What an evaluation asset looks like

It is not a benchmark score and it is not a vibe check in a spreadsheet. A working evaluation asset has four parts: an adjudicated test set drawn from production traffic, large enough to detect the differences you care about; thresholds someone accountable has signed, set before the results were known; a failure taxonomy, because ‘accuracy 94%’ hides whether the 6% is embarrassing or catastrophic; and a harness that runs on every change — model version, prompt, retrieval index — the way a test suite runs on every commit. Built once, it turns every future model decision from a debate into a measurement.

The procurement consequence

Owning the evaluation changes your commercial position. Vendor switching stops being a leap of faith and becomes a bounded measurement exercise, which means you can credibly threaten to switch, which means you negotiate from strength. Model price drops — which the market has delivered year after year — become savings you can actually bank, because you can verify the cheaper tier still clears your thresholds. And when a vendor deprecates the model you depend on, you have a selection procedure instead of an emergency.

This is why every system we deliver ships with the evaluation suite that judges it, and why the thresholds are agreed before the pilot starts. Rent the model — it will change under you regardless. Own the judgement. A programme that rents both has no position at all.

Read next

A demo budget is not a run rate

Pilots are judged on capability and killed by arithmetic. The cost ceiling belongs in the threshold schedule, not in the rollout post-mortem.