Search

Search Cadence Lab

1 published insight

Open full search

Free evaluation tool

Decide what “good” means before you test the AI.

Turn a task, desired result, and important failure modes into a practical rubric, test set, review rules, and stop conditions your team can challenge.

Build an evaluation plan

How it works

A demo shows possibility. An evaluation shows evidence.

Describe the real task, the result people need, and which errors matter. The builder turns that context into observable review criteria and cases that challenge more than the happy path.

It does not run a model, judge outputs, approve a system, or replace security, privacy, legal, accessibility, or domain review.

Build the plan

Start with the work and its consequences.

Describe categories and roles, not real customer, employee, confidential, regulated, or proprietary information.

01Define the task

Use a task-based name, such as “Support reply quality.”

Name a role or audience, not a person.

Include the starting information and the action expected.

Describe observable qualities and decisions, not “high quality.”

Examples: approved policy, product record, calculation, transcript, or subject-matter review.

02Name the failures
Which failures matter for this task?

Select every category the evaluation should challenge. If none are selected, the plan will include general quality checks.

What happens if an output is wrong?

Put one case per line. Use synthetic or anonymized descriptions.

03Define review

Name a role with the knowledge and authority to challenge it.

Include output quality, error patterns, human effort, and the effect on people or operations.

What it considers

Five parts of a useful evaluation.

  1. 01

    Task What real work must the output support?

  2. 02

    Evidence What can a reviewer verify?

  3. 03

    Failure Which errors matter and to whom?

  4. 04

    Coverage Which ordinary and difficult cases must be tested?

  5. 05

    Decision What permits continued use or requires a stop?