---
title: "Evaluation-Driven Development: Write the Test Before Choosing the Model"
description: Build AI products around evaluation cases, business acceptance thresholds, and regression gates instead of subjective demos.
---

[Cayru Blog](https://www.cayru.cr/en-us/blog)

# [Evaluation-Driven Development: Write the Test Before Choosing the Model](https://www.cayru.cr/en-us/blog/evaluation-driven-development-write-the-test-before-choosing-the-model)

 Written by [Tony Ruiz](https://www.cayru.cr/en-us/blog/author/tony-ruiz) | Sep 30, 2026, 2:53:47 AM

Traditional software teams can often define a correct output exactly. AI products are different: quality may be probabilistic, context-dependent, and spread across several tool calls.

That does not make testing optional. It makes the evaluation design part of the product.

Anthropic describes evaluations as tests that give an AI system an input and apply grading logic to measure success, emphasizing that their value compounds across an agent's lifecycle.\[^5\] OpenAI has also published an end-to-end example that uses evaluations as the core process for turning a human workflow into a production system.\[^6\]

The executive implication is simple: approve an AI project only when the team can explain how it will know the product is getting better.

## Begin with real cases

Create the first evaluation set before model selection. Use examples from the actual workflow:

- common requests;
- high-value cases;
- ambiguous or incomplete information;
- policy conflicts;
- unusual but legitimate edge cases;
- known historical failures; and
- inputs that should be refused or escalated.

Twenty carefully selected cases can be more useful at the start than a thousand generic examples. The goal is not statistical completeness. It is to make the team's assumptions visible and testable.

## Grade the outcome in layers

Do not reduce quality to one subjective score. Evaluate several layers:

1. **Task result:** Was the answer, classification, or artifact useful and correct?
2. **Process:** Did the system select the right tools and use the right evidence?
3. **Policy:** Did it stay within data, safety, and authority rules?
4. **Business state:** Did the intended record or transaction end correctly?
5. **Experience:** Was latency, explanation, and escalation acceptable to the user?

Use deterministic graders wherever possible: schemas, calculations, expected records, policy rules, and integration tests. Use expert or model-based judgment for qualities that genuinely require it, and calibrate automated judges against human labels.

## Define release thresholds

An evaluation suite becomes operational when the team specifies what must pass.

For example:

- no critical policy violation;
- at least 95% correct routing on approved intents;
- 100% escalation on a defined high-risk set;
- no regression beyond an agreed tolerance;
- P95 latency under the user threshold; and
- cost below the workflow's unit-economic ceiling.

Different risk tiers can have different thresholds. What matters is that the decision is made before a persuasive demo creates pressure to launch.

## Turn production failures into assets

When a user corrects the system or an incident occurs, preserve the case, expected behavior, context, configuration, and outcome. Add it to the regression suite after removing sensitive data appropriately.

The result is a learning loop: real failures make future releases harder to break in the same way.

Run the suite when the model, prompt, context source, tool, policy, or orchestration changes. AI behavior can move even when application code does not.

Cayru builds evaluation harnesses that connect technical behavior with business acceptance, giving product and engineering leaders evidence for model selection, release, and optimization.

[Define the first 25 evaluation cases](https://www.cayru.cr/contact-us) for one high-value AI workflow.

[View full post](https://www.cayru.cr/en-us/blog/evaluation-driven-development-write-the-test-before-choosing-the-model)

```json
{
  "@context" : "http://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Tony Ruiz"
  },
  "dateModified" : "2026-09-30T02:57:33.831Z",
  "datePublished" : "2026-09-30T02:53:47Z",
  "headline" : "Evaluation-Driven Development: Write the Test Before Choosing the Model",
  "image" : {
    "@type" : "ImageObject",
    "height" : 941,
    "url" : "https://51106415.fs1.hubspotusercontent-na1.net/hubfs/51106415/07-evaluation-driven-development.png",
    "width" : 1672
  },
  "mainEntityOfPage" : "https://www.cayru.cr/en-us/blog/evaluation-driven-development-write-the-test-before-choosing-the-model",
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject",
      "height" : 60,
      "url" : "/hs/hsstatic/content_shared_assets/static-1.4092/img/default-amp-logo.png",
      "width" : 60
    },
    "name" : "Cayru Blog"
  }
}
```