Evaluate with Datasets - Langfuse

Evaluate with Datasets

This guide walks you through how to set up a dataset, run an experiment on it, and evaluate the results. This is useful for evaluating changes before you deploy them to production. If you don't yet know what to evaluate, Choosing what to evaluate will help you further. For how datasets and experiments fit together, see Datasets and Experiments.

Agentic installation

Install the Langfuse Agent Skill to let your coding agent access all Langfuse features.

Ask your coding agent

Install the Langfuse Agent Skill from github.com/langfuse/skills
and use it to create a first dataset for this application
with Langfuse.

Agent instruction

Install the Langfuse Agent Skill from github.com/langfuse/skills
and use it to create a first dataset for this application
with Langfuse.

Install via npm ( skills CLI ):

npx skills add langfuse/skills --skill "langfuse"

Alternatively you can manually clone the skill

  1. Clone repo somewhere stable
git clone https://github.com/langfuse/skills.git /path/to/langfuse-skills
  1. Make sure your agent's skills dir exists
mkdir -p /path/to/<agent-skill-root>/skills
  1. Symlink the skill folder
ln -s /path/to/langfuse-skills/skills/langfuse /path/to/<agent-skill-root>/skills/langfuse

Then prompt your agent:

Agent instruction

Set up offline evaluation for this application with Langfuse.

Manual setup

Running experiments to test the performance of your system has three aspects:

This guide uses the current Langfuse SDKs: Python SDK v4 and JS/TS SDK v5. Both use Langfuse's OpenTelemetry-based tracing and the current experiment runner. If you use an older SDK, see the Python v3 → v4 or JS/TS v4 → v5 migration guide.

Create a project and get API keys

  1. Create a Langfuse account or self-host Langfuse.
  2. Create a project and open Settings → API Keys.
  3. Create an API key for your model provider. This example uses OpenAI.

Set the keys as environment variables:

export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"
export OPENAI_API_KEY="sk-..."

Use the base URL for your Langfuse Cloud data region or self-hosted deployment.

Install the SDKs

pip install langfuse openai
# pnpm
pnpm add @langfuse/client @langfuse/openai @langfuse/otel @opentelemetry/sdk-node openai tsx

# npm
npm install @langfuse/client @langfuse/openai @langfuse/otel @opentelemetry/sdk-node openai tsx

Create the dataset

A dataset is a collection of test cases. Each item has an input and, optionally, an expected output that evaluators can compare against. A first dataset might have a handful of representative cases; later you can grow it as you learn what fails.

In the example below we add five questions about San Francisco tourist sites.

Python SDK

from langfuse import get_client

langfuse = get_client()
dataset_name = "san-francisco-sites"

langfuse.create_dataset(
    name=dataset_name,
    description="Questions and expected answers about sites in San Francisco",
)

items = [
    {
        "id": "evaluation-quickstart-sf-golden-gate-bridge",
        "input": {
            "question": "Which red-orange suspension bridge connects San Francisco with Marin County?"
        },
        "expected_output": "Golden Gate Bridge",
    },
    {
        "id": "evaluation-quickstart-sf-alcatraz-island",
        "input": {
            "question": "Which island in San Francisco Bay is home to a former federal prison?"
        },
        "expected_output": "Alcatraz Island",
    },
    {
        "id": "evaluation-quickstart-sf-palace-of-fine-arts",
        "input": {
            "question": "Which Beaux-Arts landmark in the Marina District features a rotunda beside a lagoon?"
        },
        "expected_output": "Palace of Fine Arts",
    },
    {
        "id": "evaluation-quickstart-sf-coit-tower",
        "input": {
            "question": "Which Art Deco tower stands on Telegraph Hill?"
        },
        "expected_output": "Coit Tower",
    },
    {
        "id": "evaluation-quickstart-sf-lombard-street",
        "input": {
            "question": "Which San Francisco street is famous for a steep block with eight hairpin turns?"
        },
        "expected_output": "Lombard Street",
    },
]

for item in items:
    langfuse.create_dataset_item(dataset_name=dataset_name, **item)

print(f"Created {dataset_name} with {len(items)} items")

Run the setup script:

python seed_dataset.py

Run an experiment

The SDK runs your existing application against each dataset item. Your application stays in its own environment, with access to its tools, retrieval logic, and dependencies. A task function maps each item to your application's inputs and returns its output for evaluation. If you only need to test a prompt + model combination, you can also run prompt experiments in the UI.

Example of Python SDK

from langfuse import Evaluation, get_client
from langfuse.openai import OpenAI

langfuse = get_client()
client = OpenAI()

def answer_question(question: str):
    response = client.responses.create(
        model="gpt-5-mini",
        input=[
            {
                "role": "system",
                "content": "Answer the question about a site in San Francisco.",
            },
            {"role": "user", "content": question},
        ],
    )
    return response.output_text

def application_task(*, item, **kwargs):
    return answer_question(item.input["question"])

def exact_match(*, output, expected_output, **kwargs):
    return Evaluation(
        name="exact_match",
        value=1.0 if output == expected_output else 0.0,
    )
try:
    dataset = langfuse.get_dataset("san-francisco-sites")
    result = dataset.run_experiment(
        name="San Francisco sites",
        run_name="San Francisco sites v1",
        description="First prompt for answering questions about San Francisco sites",
        task=application_task,
        evaluators=[exact_match],
    )

print(result.format())
finally:
    langfuse.flush()

Run the experiment:

python sf_sites.py

View the results

The formatted terminal output includes the experiment summary and a link to the dataset run. You can also open Experiments in Langfuse.