Search

Search Cadence Lab

1 published insight

Open full search

AI tooling · AI variance testing

Don't stop at a wrong answer. Find the conditions behind it.

AI outputs change when the request, source material, retrieval, model, tools, or workflow changes. Variance testing connects a bad answer to those conditions so your team can decide what to fix, validate, escalate, or accept.

Visible signals

Calling it a hallucination doesn't explain the failure.

A wrong answer is the end of a path. The useful question is which source, context, model behavior, tool, or review decision allowed it to matter.

01

Repeated runs change important facts

Names, dates, citations, totals, classifications, or recommendations move when the task appears unchanged.

02

A grounded answer goes beyond the source

The response starts with evidence, then adds details or certainty the source doesn't support.

03

A model update changes the result

A new model improves general quality but changes formatting, refusal, classification, or tool behavior.

04

Reviewers fix errors without recording why

People correct facts, tone, or routing, but the reason never becomes evaluation data.

05

One average score defines quality

A summary hides the customers, inputs, or important edge cases where performance fails.

06

A bad result can't be replayed

The prompt remains, but the retrieved context, model version, tool results, and human action are missing.

Trace conditions

Define the conditions before judging the answer.

Output quality depends on the task and evidence used to produce it. A useful trace shows where meaningful variation entered.

01

Task

What decision should the output support?

Define the user, purpose, expected form, unacceptable error, and where variation is allowed.

02

Input

What changed in the request or source?

Trace wording, ambiguity, missing facts, conflicting instructions, format, and provenance.

03

Context

Which evidence reached the model?

Record retrieval, ranking, truncation, history, templates, instructions, and excluded context.

04

Execution

Which model and tools produced the result?

Capture model version, settings, output rules, tool calls, retries, fallbacks, and transformations.

05

Review

How was the output judged or trusted?

Separate automated checks, model grading, expert review, user interpretation, and override.

06

Outcome

What happened after the answer was used?

Connect the output to the customer, decision, correction, delay, risk, or other real consequence.

Trace path

Follow the variation from request to consequence.

Rebuild real runs, compare their conditions, find where they split, and choose the right control.

  1. 01

    Bound

    Which difference matters to the work?

    Task, user, expected behavior, acceptable range, prohibited error, and consequence.

  2. 02

    Capture

    Can the team rebuild each run?

    Input, sources, retrieval, prompt, model, settings, tools, policy, output, and reviewer.

  3. 03

    Compare

    Which condition changed with the result?

    Repeated runs, paired examples, model comparisons, edge cases, and failure groups.

  4. 04

    Explain

    What evidence supports the likely cause?

    Observed pattern, alternatives tested, uncertainty, repeatability, and limits of the trace.

  5. 05

    Control

    What should the workflow do differently?

    Source repair, retrieval changes, validation, fallback, human review, monitoring, and retesting.

Evidence

Keep enough context to explain the difference.

Prompts and answers aren't enough. A useful trace connects model behavior to source evidence, application state, human review, and the final outcome.

01

Representative inputs

Routine cases, important edge cases, unclear requests, incomplete records, and past failures.

02

Source and retrieval evidence

Document versions, provenance, ranking, retrieved passages, missing context, and conflicting sources.

03

Execution settings

Model version, prompts, parameters, schemas, tools, retries, fallbacks, caches, and transformations.

04

Evaluation rules

Reference evidence, rubrics, automated checks, reviewers, thresholds, and documented disagreement.

05

Human review

Corrections, overrides, missed errors, review time, escalation, and the context shown to people.

06

Operating outcomes

Customer effect, workflow result, downstream action, delay, rework, complaints, and recovery.

Trace before certainty

Testing can reveal repeatable conditions and likely causes. It shouldn't claim a precise internal explanation when the evidence supports only a pattern.

Common failure patterns

Evaluation fails when it ignores the real consequence.

A benchmark can look strong while hiding weak sources, poor review, and the cases where one wrong answer matters.

01

One answer becomes the full standard

A rigid reference rejects useful variation without defining the facts and actions that can't vary.

02

The average hides the weak case

Overall performance looks good while one customer group, language, or important decision remains unreliable.

03

A model grader becomes independent truth

A second model scores output without expert review, calibration, or analysis of shared blind spots.

04

Prompt tuning replaces system repair

Teams rewrite instructions while leaving weak sources, retrieval gaps, permissions, and review unchanged.

05

Corrections disappear after delivery

People fix the answer, but the failure, reason, context, and outcome never enter the test set.

06

Evaluation ends at launch

Models, data, tools, and user behavior change while the test set and threshold stay frozen.

Practical outputs

Turn inconsistent answers into a learning loop.

The result makes important variation visible, connects it to operating conditions, and defines what happens when an answer falls outside the safe range.

  1. 01

    End-to-end variance map

    The task, input, evidence, context, model, tools, validation, review, and outcome.

  2. 02

    Evaluation set and scorecard

    Representative cases, failure groups, evidence, rubrics, checks, thresholds, and disagreement.

  3. 03

    Control and escalation rules

    Triggers for validation, source checks, fallback, human review, refusal, or correction.

  4. 04

    Monitoring and revision plan

    Production signals, sampling, corrections, owners, model-change tests, and review cadence.

Engagement fit

Use variance testing when different answers change the work.

This work fits a defined AI workflow with repeatable tasks, available evidence, visible corrections, and outcomes you can evaluate. It won't make a generative model perfectly predictable.

Start a fit check

Related diagnostic paths

Start with the decision the output should support.

AI Service Readiness Review

You need bounded outcomes, evidence, evaluation, human oversight, safeguards, and an implementation decision.

Review the diagnostic

CX Systems Diagnostic

The unreliable output sits inside a customer workflow that crosses teams, systems, and handoffs.

Review the diagnostic

CRM Workflow Audit

The output depends on CRM records, lifecycle signals, fields, automation, reporting, or customer context.

Review the diagnostic