Controlled experiments
Copy page
Isolate a product-surface change and compare agent behavior under a declared experimental design.
Choose the question before the test
Use a controlled comparison to decide whether a change to your documentation or tool descriptions helps agents choose and use your product. Define the question first, then change only the fields needed to answer it.
| Design | Question it answers |
|---|---|
| No supplied candidate list (separate design) | Which providers does the model propose on its own? The current PromptCase discovery prompt can ask for candidates without a supplied lineup. This measures model-reported candidates, not web retrieval; a search-based evaluation needs a separately configured retrieval path and its execution evidence. |
| Fixed-candidate comparison | Which provider does the agent choose among the specified alternatives? |
| Paired metadata A/B | How does a declared description change affect behavior within the same candidate environment? |
These designs answer different questions. A description experiment supplies controlled information to test its effect on decisions. A discovery evaluation examines appearance under its declared retrieval conditions. A fixed-candidate lift is not an organic-discovery result.
Match the method to the evaluation
Flovia uses different analysis paths for diagnosing a task, comparing a description change across the funnel, and estimating a change in provider selection. Choose the path around the decision you need to make.
| Evaluation | What is compared | How results are summarized |
|---|---|---|
| PromptCase funnel | Agent decisions and contract checks for each task | Stage rates and conditional transitions; supported rate metrics use prompt-cluster bootstrap intervals |
| PromptCase metadata A/B | Matched baseline and variant runs through the funnel; candidate rotations follow the study | Stage-by-stage differences, missing observations, and pairs that improved or regressed; candidate orders are pooled with equal weight |
| Selection-lift study | Provider selection under baseline and variant metadata | Complete eligible pairs, the study’s intent or paraphrase-group weighting, and paired cluster bootstrap intervals |
The funnel A/B summary and selection-lift estimator are separate analyses. A confirmation study also requires its own recorded sample-size decision and evidence checks. Agree on the analysis path, candidate rotations, interval method, and confirmation requirements when scoping the evaluation.
Pair baseline and variant
In a metadata A/B study, the baseline and variant share the task, requested model, repeat, candidate order, and non-intervened inputs. Only the target fields named in the plan change. Depending on the study, these can be Discovery descriptions, Invocation descriptions, or a declared combination.
Shared: task × model × repeat × candidate order × execution path
Baseline: original target description
Variant: revised target description
Fixed: competitor metadata, tool schemas, fixtures, scoring rules
Compare: selection, invocation, completion, and failure stageWhen both Discovery and Invocation descriptions change together, the result measures the joint intervention. Attributing the difference to one surface requires a design that isolates that surface.
Control candidate-position effects
Candidate order is an experimental factor. Supported metadata studies use cyclic Latin-square rotations so each provider occupies each candidate position. Baseline and variant are paired within the same order, and summaries can pool with equal weight across the Latin orders.
Pairing holds the surrounding scenario constant; rotation balances the target’s exposure across positions. Together, they make a description comparison easier to interpret when the agent is sensitive to how alternatives are presented.
| Rotation | Position 1 | Position 2 | Position 3 |
|---|---|---|---|
| A | Target | Alternative 1 | Alternative 2 |
| B | Alternative 1 | Alternative 2 | Target |
| C | Alternative 2 | Target | Alternative 1 |
This example illustrates position balancing, not a customer experiment. Run order and arm-first balancing follow the selected study plan; they are not assumed for every evaluation.
Interpret the contrast at the tested scope
A paired difference estimates how the tested behavior changed between baseline and variant in the stated tasks, models, candidate environment, and window. Read it alongside pair completeness, missing observations, model-level results, and the estimator actually used.
Descriptive funnel A/B findings do not by themselves establish statistical significance or a general causal effect. A confirmation study has its own pre-specified evidence and qualification requirements. Claims about production impact need evidence beyond a controlled simulated comparison.