Skip to content
Flovia Documentation
DocsMethodology

Controlled experiments

Isolate a product-surface change and compare agent behavior under a declared experimental design.

Choose the question before the test

Use a controlled comparison to decide whether a change to your documentation or tool descriptions helps agents choose and use your product. Define the question first, then change only the fields needed to answer it.

DesignQuestion it answers
No supplied candidate list (separate design)Which providers does the model propose on its own? The current PromptCase discovery prompt can ask for candidates without a supplied lineup. This measures model-reported candidates, not web retrieval; a search-based evaluation needs a separately configured retrieval path and its execution evidence.
Fixed-candidate comparisonWhich provider does the agent choose among the specified alternatives?
Paired metadata A/BHow does a declared description change affect behavior within the same candidate environment?

These designs answer different questions. A description experiment supplies controlled information to test its effect on decisions. A discovery evaluation examines appearance under its declared retrieval conditions. A fixed-candidate lift is not an organic-discovery result.

Match the method to the evaluation

Flovia uses different analysis paths for diagnosing a task, comparing a description change across the funnel, and estimating a change in provider selection. Choose the path around the decision you need to make.

EvaluationWhat is comparedHow results are summarized
PromptCase funnelAgent decisions and contract checks for each taskStage rates and conditional transitions; supported rate metrics use prompt-cluster bootstrap intervals
PromptCase metadata A/BMatched baseline and variant runs through the funnel; candidate rotations follow the studyStage-by-stage differences, missing observations, and pairs that improved or regressed; candidate orders are pooled with equal weight
Selection-lift studyProvider selection under baseline and variant metadataComplete eligible pairs, the study’s intent or paraphrase-group weighting, and paired cluster bootstrap intervals

The funnel A/B summary and selection-lift estimator are separate analyses. A confirmation study also requires its own recorded sample-size decision and evidence checks. Agree on the analysis path, candidate rotations, interval method, and confirmation requirements when scoping the evaluation.

Pair baseline and variant

In a metadata A/B study, the baseline and variant share the task, requested model, repeat, candidate order, and non-intervened inputs. Only the target fields named in the plan change. Depending on the study, these can be Discovery descriptions, Invocation descriptions, or a declared combination.

Conceptual comparison unit
Shared: task × model × repeat × candidate order × execution path
Baseline: original target description
Variant: revised target description
Fixed: competitor metadata, tool schemas, fixtures, scoring rules
Compare: selection, invocation, completion, and failure stage

When both Discovery and Invocation descriptions change together, the result measures the joint intervention. Attributing the difference to one surface requires a design that isolates that surface.

Control candidate-position effects

Candidate order is an experimental factor. Supported metadata studies use cyclic Latin-square rotations so each provider occupies each candidate position. Baseline and variant are paired within the same order, and summaries can pool with equal weight across the Latin orders.

Pairing holds the surrounding scenario constant; rotation balances the target’s exposure across positions. Together, they make a description comparison easier to interpret when the agent is sensitive to how alternatives are presented.

RotationPosition 1Position 2Position 3
ATargetAlternative 1Alternative 2
BAlternative 1Alternative 2Target
CAlternative 2TargetAlternative 1

This example illustrates position balancing, not a customer experiment. Run order and arm-first balancing follow the selected study plan; they are not assumed for every evaluation.

Interpret the contrast at the tested scope

A paired difference estimates how the tested behavior changed between baseline and variant in the stated tasks, models, candidate environment, and window. Read it alongside pair completeness, missing observations, model-level results, and the estimator actually used.

Descriptive funnel A/B findings do not by themselves establish statistical significance or a general causal effect. A confirmation study has its own pre-specified evidence and qualification requirements. Claims about production impact need evidence beyond a controlled simulated comparison.

© 2026 FloviaAgent Experience Optimizer

Searches all pages and sections · Tab to a result, Enter to open