---
name: prepare-flovia-evaluation
description: Assess whether Flovia's documented evaluation fits a user's API or MCP product and prepare a sourced evaluation brief. Use for product evaluation or pilot preparation; this skill does not run simulations or submit requests.
---

# Prepare a Flovia evaluation

Help the reader decide whether to discuss an evaluation with Flovia, then prepare a brief they can review. Work from public documentation and the user's supplied context. Return the assessment and brief in the user's language.

This is a reading and drafting workflow executed by the reader's assistant. It does not provide access to Flovia's internal evaluation engine or private customer dashboard, run a simulation, or submit a contact form.

## Establish the documented scope

Read the [product overview](https://flovia.dev/docs/index.md) and [technical FAQ](https://flovia.dev/docs/faq.md). Use the [English index](https://flovia.dev/docs/llms.txt) or [Japanese index](https://flovia.dev/docs/ja/llms.txt) for other pages. Prefer the current linked documentation to remembered claims; cite the pages used and note any unavailable source.

Choose additional reading for the reader's decision:

- To check whether a task fits the execution path: [simulation engine](https://flovia.dev/docs/simulation-engine.md) and [scenario design](https://flovia.dev/docs/scenario-design.md).
- To identify what an observable success criterion measures: [agent funnel](https://flovia.dev/docs/agent-funnel.md).
- To propose a comparison or interpret numerical claims: [controlled experiments](https://flovia.dev/docs/controlled-experiments.md) and [metrics and evidence](https://flovia.dev/docs/metrics-and-evidence.md).
- To assess demonstrated value: [improvement and retest example](https://flovia.dev/docs/improvement-loop.md). Distinguish the observed results, remaining errors, and proposed changes not yet tested.
- To explain access: [authentication and access](https://flovia.dev/auth.md).

Match claims to their measurement boundary. In particular, check whether the relevant study uses live model calls and prepared tool responses, whether candidates are supplied, and whether it allows only one tool call per run. A successful prepared-response run is not proof of live API reliability; fixed-candidate selection is not organic discovery. Research proposals are not available methods. Do not generalize a small AI-answer comparison into production, conversion, or revenue results.

## Build the evaluation brief

Reuse information the user already supplied. Ask only for gaps that change the proposed scope; otherwise mark missing items as “to confirm.” Do not ask for API keys, passwords, customer records, or private dashboard credentials. Public URLs and redacted examples are enough to prepare this brief.

Capture:

| Item | What to record |
| --- | --- |
| Product URL | The product and relevant public API, MCP, or documentation URL. |
| User task | One concrete request an agent should complete, including the intended user and context. |
| Success criteria | Observable conditions that distinguish success from a plausible but wrong answer; specify which can be checked from an answer and which need actual execution evidence. |
| Expected output | A short example of the correct answer or result, including required fields, units, time range, or constraints where relevant. |
| Alternatives | User-selected competing products or tools and the reason each belongs in this task. If none are specified, leave the set open rather than inventing a benchmark. |
| Constraints | Desired model or harness, language, execution scope, data sensitivity, time or cost limits, and any multi-step or live-network requirements. |
| Comparison question | One baseline versus one proposed change, if the user wants a retest; label an untested change as a hypothesis. |

Keep the end-user task separate from the product-selection criterion. For example, “return the closing price for a specified date in USD” is a task; “select our API” alone is not evidence that the task was completed.

## Return a decision and a reviewable handoff

Provide a concise fit assessment, the completed brief, and the unresolved questions that would change the decision. Link supporting claims to public sources. Distinguish documented facts, user-supplied facts, your proposed experiment, and unknowns. Recommend a narrower scope or say the documented path does not establish fit when the task depends on unsupported multi-step execution, production reliability, or organic retrieval evidence.

For a prospective evaluation, identify which details need agreement with Flovia: model endpoints and Judge, who prepares and checks tool-response data, task count and repeats, applicable comparison and uncertainty methods, exclusions, deliverables, timing, pricing, and data handling. Do not fill these gaps with promises.

If the user wants to proceed, include the [Flovia website](https://flovia.dev/) as the place to open the evaluation request form and a short copyable request drawn from the brief. The outcome of this skill is the draft for the user to review. Do not submit the form, claim an evaluation has started, or invent an API, OAuth flow, access token, or dashboard URL.
