Skip to content
Flovia Documentation
DocsUse the results

Metrics and evidence

Read every result with its population, uncertainty, execution conditions, and supporting trace.

Start with the denominator

Use the results to decide which change to keep and which task to investigate next. Start with the affected tasks and models, compare the outcomes, then open the underlying runs to see what changed.

To read a rate, start with what was counted. Flovia’s metric records keep the estimate, the numerator and denominator where applicable, the eligible run count, the missing count, and the interval method when one is available. Cumulative stage rates and conditional transitions use different populations.

ViewHow to read it
Cumulative stage rateHow many runs in the declared population reached this stage? Earlier drop-off affects later reach.
Conditional transition rateOf the runs eligible at the preceding gate, how many passed the next one?
End-to-end completionHow many eligible tasks completed across the whole evaluated journey?
Paired liftHow did the outcome differ between matched baseline and variant observations?
Illustrative arithmetic · not a customer result
# Illustrative figures only — not measured customer results
100 eligible tasks → 60 selections → 45 valid invocations
Selection reach: 60 / 100 = 60%
Invocation reach: 45 / 100 = 45%
Selection → invocation: 45 / 60 = 75%

Baseline completion: 30%
Variant completion: 42%
Difference: +12 percentage points

Infrastructure errors, unobserved stages, and missing Judge results mean different things. The selected metric contract sets the exclusions and missing-data treatment. Some paired estimators use only complete eligible pairs; planned-population summaries can keep missing observations in the denominator. The policy declared in the report takes precedence over any generic formula.

Account for repeated tasks

Repeated runs of the same prompt are related observations. Where the selected metric supports it, the engine uses prompt-cluster or intent-cluster bootstrap intervals: it resamples task clusters together to preserve that grouping. Metadata-lift estimators in the corresponding evaluation path support paired cluster bootstrap.

An interval is reported only when the selected method’s requirements are met. A missing interval does not imply zero uncertainty. Read model-level and task-family results alongside the aggregate, especially when models behave differently or the task set is small.

In evaluation paths that support confirmation studies, the engine checks whether coverage, sample size, precision, and traceability to source runs meet the predeclared requirements. This distinguishes a promising pilot from a result that meets confirmation criteria. Descriptive funnel A/B summaries use their own declared estimator.

Trace a finding back to a run

Evaluation workflow Conceptual
  1. 01PlanScope + input versions
  2. 02RunDecisions + results
  3. 03ScoreChecks + rubric
  4. 04FindingMetric + diagnosis
  • Evaluation definition: tasks, prompts, models, candidate conditions, execution paths, and measurement window.
  • Coverage: planned and completed runs, missing observations, exclusions, and failure reasons.
  • Run evidence: candidate and selection outputs, tool calls, argument validation, execution result, and answer where reached.
  • Scoring provenance: deterministic checks, Judge status, criteria, and the selected scoring version.
  • Comparison evidence: baseline and variant identities, changed fields, pairing, estimator, and any applicable qualification.

Scope: what a result can support

An evaluation is tied to its tasks, models, candidate conditions, execution environment, and measurement window. Reading that scope helps you decide whether to adopt a change or collect more evidence.

EvidenceDecision it supportsWhat needs separate evidence
Fixed-candidate comparisonChanges in selection and use among supplied alternativesDiscovery without supplied candidates or actual web retrieval
Fixture-backed executionInvocation and answer correctness against a declared response contractLive API availability, authentication, and latency
Controlled retestThe stage that improved and the failures that remain in the tested scopeEffects on production usage, conversion, or revenue

The requested model identity is recorded; confirming which model actually served the request also needs provider-side evidence. Repeats and related-task evaluations can extend the evidence beyond an initial result.

© 2026 FloviaAgent Experience Optimizer

Searches all pages and sections · Tab to a result, Enter to open