# Evaluate with Datasets

This guide walks you through how to set up a dataset, run an experiment on it, and evaluate the results. This is useful for evaluating changes before you deploy them to production. If you don't yet know what to evaluate, [Choosing what to evaluate](/content/academy/evaluate/choosing-what-to-evaluate/index.html) will help you further. For how datasets and experiments fit together, see [Datasets](/content/academy/datasets/index.html) and [Experiments](/content/academy/experiments/index.html).

## Agentic installation

Install the [Langfuse Agent Skill](https://github.com/langfuse/skills) to let your coding agent access all Langfuse features.

### Ask your coding agent
```
Install the Langfuse Agent Skill from github.com/langfuse/skills
and use it to create a first dataset for this application
with Langfuse.
```

### Agent instruction

```bash
Install the Langfuse Agent Skill from github.com/langfuse/skills
and use it to create a first dataset for this application
with Langfuse.
```

### Install via npm ( skills CLI ):
```
npx skills add langfuse/skills --skill "langfuse"
```

### Alternatively you can manually clone the skill

1. Clone repo somewhere stable

```bash
git clone https://github.com/langfuse/skills.git /path/to/langfuse-skills
```

2. Make sure your agent's skills dir exists

```bash
mkdir -p /path/to/<agent-skill-root>/skills
```

3. Symlink the skill folder

```bash
ln -s /path/to/langfuse-skills/skills/langfuse /path/to/<agent-skill-root>/skills/langfuse
```

Then prompt your agent:

### Agent instruction

```
Set up offline evaluation for this application with Langfuse.
```

## Manual setup

Running experiments to test the performance of your system has three aspects:

- **Dataset**: Test cases with inputs and expected outputs
- **Experiment condition**: The application variation or model call you want to test
- **Evaluators**: Functions that score the output

This guide uses the current Langfuse SDKs: **Python SDK v4** and **JS/TS SDK v5**. Both use Langfuse's OpenTelemetry-based tracing and the current experiment runner. If you use an older SDK, see the [Python v3 → v4](/content/docs/observability/sdk/upgrade-path/python-v3-to-v4/index.html) or [JS/TS v4 → v5](/content/docs/observability/sdk/upgrade-path/js-v4-to-v5/index.html) migration guide.

### Create a project and get API keys
1. [Create a Langfuse account](/content/cloud/index.html) or [self-host Langfuse](/content/self-hosting/index.html).
2. Create a project and open **Settings → API Keys**.
3. Create an API key for your model provider. This example uses [OpenAI](https://platform.openai.com/api-keys).

Set the keys as environment variables:

```bash
export LANGFUSE_PUBLIC_KEY="pk-lf-..."
export LANGFUSE_SECRET_KEY="sk-lf-..."
export LANGFUSE_BASE_URL="https://cloud.langfuse.com"
export OPENAI_API_KEY="sk-..."
```

Use the base URL for your [Langfuse Cloud data region](/content/security/data-regions/index.html) or self-hosted deployment.

### Install the SDKs
- Python SDK

```bash
pip install langfuse openai
```

- JS/TS SDK

```bash
# pnpm
pnpm add @langfuse/client @langfuse/openai @langfuse/otel @opentelemetry/sdk-node openai tsx

# npm
npm install @langfuse/client @langfuse/openai @langfuse/otel @opentelemetry/sdk-node openai tsx
```

### Create the dataset
A dataset is a collection of test cases. Each item has an input and, optionally, an expected output that evaluators can compare against. A first dataset might have a handful of representative cases; later you can grow it as you learn what fails.

In the example below we add five questions about San Francisco tourist sites.

**Python SDK**
```python
from langfuse import get_client

langfuse = get_client()
dataset_name = "san-francisco-sites"

langfuse.create_dataset(
    name=dataset_name,
    description="Questions and expected answers about sites in San Francisco",
)

items = [
    {
        "id": "evaluation-quickstart-sf-golden-gate-bridge",
        "input": {
            "question": "Which red-orange suspension bridge connects San Francisco with Marin County?"
        },
        "expected_output": "Golden Gate Bridge",
    },
    {
        "id": "evaluation-quickstart-sf-alcatraz-island",
        "input": {
            "question": "Which island in San Francisco Bay is home to a former federal prison?"
        },
        "expected_output": "Alcatraz Island",
    },
    {
        "id": "evaluation-quickstart-sf-palace-of-fine-arts",
        "input": {
            "question": "Which Beaux-Arts landmark in the Marina District features a rotunda beside a lagoon?"
        },
        "expected_output": "Palace of Fine Arts",
    },
    {
        "id": "evaluation-quickstart-sf-coit-tower",
        "input": {
            "question": "Which Art Deco tower stands on Telegraph Hill?"
        },
        "expected_output": "Coit Tower",
    },
    {
        "id": "evaluation-quickstart-sf-lombard-street",
        "input": {
            "question": "Which San Francisco street is famous for a steep block with eight hairpin turns?"
        },
        "expected_output": "Lombard Street",
    },
]

for item in items:
    langfuse.create_dataset_item(dataset_name=dataset_name, **item)

print(f"Created {dataset_name} with {len(items)} items")
```

**Run the setup script:**
```
python seed_dataset.py
```

## Run an experiment
The SDK runs your existing application against each dataset item. Your application stays in its own environment, with access to its tools, retrieval logic, and dependencies. A task function maps each item to your application's inputs and returns its output for evaluation. If you only need to test a prompt + model combination, you can also run [prompt experiments in the UI](/content/docs/evaluation/experiments/experiments-via-ui/index.html).

**Example of Python SDK**
```python
from langfuse import Evaluation, get_client
from langfuse.openai import OpenAI

langfuse = get_client()
client = OpenAI()

def answer_question(question: str):
    response = client.responses.create(
        model="gpt-5-mini",
        input=[
            {
                "role": "system",
                "content": "Answer the question about a site in San Francisco.",
            },
            {"role": "user", "content": question},
        ],
    )
    return response.output_text

def application_task(*, item, **kwargs):
    return answer_question(item.input["question"])

def exact_match(*, output, expected_output, **kwargs):
    return Evaluation(
        name="exact_match",
        value=1.0 if output == expected_output else 0.0,
    )
try:
    dataset = langfuse.get_dataset("san-francisco-sites")
    result = dataset.run_experiment(
        name="San Francisco sites",
        run_name="San Francisco sites v1",
        description="First prompt for answering questions about San Francisco sites",
        task=application_task,
        evaluators=[exact_match],
    )

print(result.format())
finally:
    langfuse.flush()
```

**Run the experiment:**
```
python sf_sites.py
```

### View the results
The formatted terminal output includes the experiment summary and a link to the dataset run. You can also open [Experiments](https://cloud.langfuse.com/project/~/experiments) in Langfuse.
