Designing datasets - Langfuse

Designing datasets for AI applications

A dataset is a repeatable set of examples that represent the scope of your application, and that you use to measure and improve your system. By running your application against the same inputs over time, you can track quality with metrics, compare changes, and catch regressions before they affect production users.

Haven't determined what is worth evaluating in a repeatable manner? Check out

The Datasets academy section explains the basic building blocks: input, expected output, and metadata. This guide focuses on the design work that happens before and while you create datasets and dataset items.

Dataset design is iterative. A good starting point is a minimally complete dataset: about 15-30 rows that can run through your application, cover the most important input slices, and have an evaluator or review rubric. Run that version early, fix the schema and evaluator, then expand into the gaps you see from those runs or from production input.

Your application will very likely have more than one evaluation dataset. Datasets are often scoped to a specific part of the system or to one sub-step the agent takes.

Iterate and expand

01Define goalscope / boundary
02Inspect sourcessample / patterns
03Choose distributionslices / roles
04Choose eval stylereference / free
05Design schemainput / output
06Build datasetrows / gaps
07Run experimentfailures / iterate

1. Start with the goal of the dataset

Before writing rows, define the smallest useful goal the dataset should support. This keeps the dataset from becoming a vague bucket of interesting examples.

For example:

A good early starting point for a dataset is looking at the most common examples end to end. This gives you the foundation and the general understanding from which you can expand.

Over time, teams add datasets for single steps, adversarial cases, red teaming, or specific sub-use cases in their application. We have even seen datasets for checking company name spelling.

The goal you set for your dataset defines boundary and job: end-to-end datasets use the payload the application receives, while step-level datasets use the structured state a step sees in production.

If two jobs need different inputs, evaluators, or release decisions, split them. A stable regression dataset and an adversarial-input dataset can both be useful, but combining them without a clear split makes aggregate scores harder to interpret.

2. Inspect available sources

Before selecting examples or writing dataset items, inspect a small sample of the material you could use. The goal is to understand what exists and how it's different from what you expected.

Start with three source types:

For each source, review input topics, input and output shapes, and failure modes. You might learn that production traces lack the retrieval context your evaluator needs, that tickets from legacy systems expose better failure labels than traces, or that there is already a good set of end-to-end examples, because support created a 'FAQ' at some point.

3. Choose the input distribution

For a first dataset, keep the input distribution simple. Start with the few slices that tell you what to do after a run:

For a support-routing dataset, the first version might use a simple scenario-by-difficulty matrix:

Input distribution is the deliberate mix of cases in the dataset: which scenarios appear, how difficult they are, and why each row is included. It is the coverage plan that helps you interpret experiment results by slice.

The distribution does not have to mirror production frequency exactly. A regression dataset may intentionally overrepresent failures, edge cases, or high-value paths. What counts is that the mix is deliberate and visible in metadata.

Add more dimensions only when they change behavior or help interpret results. Channel, language, customer segment, region, context availability, and product area can be useful metadata, but they should not all become balancing constraints on day one.

4. Decide how evaluation will work

Choose the evaluation style before you write expected outputs. This determines what the dataset item needs to contain and how useful the first results will be.

Use reference-based evaluation when each item has a known target: a correct label, expected tool call, required fact, structured output, reference answer, or expected next action. This is the best fit for regression tests and CI gates because failures are easier to inspect. The trade-off is that references take work to write and can become brittle if they over-specify wording instead of behavior.

Use reference-free evaluation when there is no stable expected output, but every item can be judged against the same rule or rubric. This works for checks such as valid JSON, language match, grounding in provided context, safety, or tone. The trade-off is that the evaluator or rubric carries more weight, so ambiguous rubrics produce ambiguous results.

For example, a docs chatbot dataset can use either approach:

Approach input shape expectedOutput shape Evaluator approach
Reference-free { question, retrievedContext[] } Omit expectedOutput or set it to null LLM-as-a-judge checks whether the answer is grounded in the retrieved documentation, answers the question, and avoids unsupported facts.
Reference-based { question } { requiredFacts[], requiredSources[], mustNot[] } A code or LLM evaluator compares the answer against the required facts and sources. This takes more preparation, but failures are easier to inspect.

Choose the cheapest evaluator that captures the requirement: code evaluators for deterministic checks, LLM-as-a-judge for language-quality judgments, and manual evaluation while you are still learning what good and bad outputs look like.

At this stage, you only need to decide how each dataset item will be reviewed: by an automated evaluator, manual annotation, or both. For the deeper work of designing the evaluator itself, see the Academy page on evaluation.

5. Design the item schema

The three dataset item fields are flexible JSON. Here, define them as a concrete contract that your experiment runner, evaluators, and reviewers can all consume.

Define:

Once you have decided on an item schema, enforce it so your team can effectively collaborate on adding items in the right structure.

For a support-routing dataset, one row could look like this:

{
  "input": {
    "message": "I was charged twice for invoice 4831. Can someone fix this?",
    "channel": "support_chat",
    "customer_tier": "business"
  },
  "expectedOutput": {
    "route": "billing_support",
    "required_actions": ["acknowledge_duplicate_charge", "ask_for_invoice_id"],
    "must_not": ["promise_refund_without_review"]
  },
  "metadata": {
    "source": "expert",
    "scenario_type": "billing",
    "difficulty": "medium",
    "dataset_role": "regression",
    "failure_mode": "wrong_route"
  }
}

The input keeps only the context the router needs. The expectedOutput states the behavior to check: route to billing, ask for the invoice ID, and do not promise a refund before review. The metadata records the row's source, scenario, difficulty, and role.

Keep the schema stable before collecting rows in bulk. Preserve behavior-shaping fields such as conversation history, retrieved context, tool state, routing metadata, or user attributes, but avoid arbitrary per-row fields or natural-language summaries of structured context.

6. Draft a first version

Once the goal, distribution, evaluation method, and schema are concrete, start turning selected source examples into dataset items. Use the sources you inspected earlier, and add them in a format that matches the defined schema.

For a minimally complete first version, choose enough rows to test the whole contract along your input distribution:

Do not wait until the dataset feels complete. Run the first coherent version in an experiment before adding more rows.

7. Run the first experiment and expand deliberately

Use the first run to check whether the input shape works with the real application path, if the evaluator outcomes make sense, and if the expected output shape works for the intended purpose.

After the first run, expand and iterate until you arrive at scope and shape that helps you feel confident about shipping changes.

How datasets evolve over time

Datasets have to evolve with your application. By monitoring production and frequently reviewing data through structured error analysis, your datasets can evolve over time to represent the production scope of your system. There are three useful expansion patterns:

Put it into practice

A set of Golden rules to check your dataset against: