# Experiments via SDK

Experiments via SDK are used to programmatically loop your applications or prompts through a dataset and optionally apply Evaluation Methods to the results. You can use a dataset hosted on Langfuse or a local dataset as the foundation for your experiment.

See also the [JS/TS SDK reference](https://js.reference.langfuse.com/classes/_langfuse_client.ExperimentManager.html) and the [Python SDK reference](https://python.reference.langfuse.com/langfuse#Langfuse.run_experiment) for more details on running experiments via the SDK.

## Why use Experiments via SDK?

- Full flexibility to use your own application logic
- Use custom scoring functions to evaluate the outputs of a single item and the full run
- Use [code evaluators](/content/docs/evaluation/evaluation-methods/code-evaluators/index.html) when you want to author deterministic Python or TypeScript checks in the Langfuse UI and reuse them across observations and experiments
- Run multiple experiments on the same dataset in parallel
- Easy to integrate with your existing evaluation infrastructure
- [Run your experiments in CI/CD](/content/docs/evaluation/experiments/experiments-ci-cd/index.html) to catch regressions before they ship

## Experiment runner SDK

Both the Python and JS/TS SDKs provide a high-level abstraction for running an experiment on a dataset. The dataset can be both local or hosted on Langfuse. Using the Experiment runner is the recommended way to run an experiment on a dataset with our SDK.

The experiment runner automatically handles:
- **Concurrent execution** of tasks with configurable limits
- **Automatic tracing** of all executions for observability
- **Flexible evaluation** with both item-level and run-level evaluators
- **Error isolation** so individual failures don't stop the experiment
- **Dataset integration** for easy comparison and tracking

The experiment runner SDK supports both datasets hosted on Langfuse and datasets hosted locally. If you are using a dataset hosted on Langfuse for your experiment, the SDK will automatically create a dataset run for you that you can inspect and compare in the Langfuse UI. For locally hosted datasets not on Langfuse, only traces and scores (if evaluations are used) are tracked in Langfuse.

### Basic Usage

Start with the simplest possible experiment to test your task function on local data. If you already have a dataset in Langfuse, [see here](/content/docs/evaluation/experiments/experiments-via-sdk#usage-with-langfuse-datasets/index.html).

Python SDK

```
from langfuse import get_client
from langfuse.openai import OpenAI

# Initialize client
langfuse = get_client()

# Define your task function
def my_task(*, item, **kwargs):
    question = item["input"]
    response = OpenAI().chat.completions.create(
        model="gpt-4.1", messages=[{"role": "user", "content": question}]
    )

return response.choices[0].message.content

# Run experiment on local data
local_data = [\
    {"input": "What is the capital of France?", "expected_output": "Paris"},\
    {"input": "What is the capital of Germany?", "expected_output": "Berlin"},\
]

result = langfuse.run_experiment(
    name="Geography Quiz",
    description="Testing basic functionality",
    data=local_data,
    task=my_task,
)

# Use format method to display results
print(result.format())
```

Make sure that OpenTelemetry is properly set up for traces to be delivered to Langfuse. See the [tracing setup documentation](/content/docs/observability/sdk/overview#initialize-tracing/index.html) for configuration details. Always flush the span processor at the end of execution to ensure all traces are sent.

### JS/TS SDK

```javascript
import { OpenAI } from "openai";
import { NodeSDK } from "@opentelemetry/sdk-node";

import {
  LangfuseClient,
  ExperimentTask,
  ExperimentItem,
} from "@langfuse/client";
import { observeOpenAI } from "@langfuse/openai";
import { LangfuseSpanProcessor } from "@langfuse/otel";

// Initialize OpenTelemetry
const otelSdk = new NodeSDK({ spanProcessors: [new LangfuseSpanProcessor()] });
otelSdk.start();

// Initialize client
const langfuse = new LangfuseClient();

// Run experiment on local data
const localData: ExperimentItem[] = [\
  { input: "What is the capital of France?", expectedOutput: "Paris" },\
  { input: "What is the capital of Germany?", expectedOutput: "Berlin" },\
];

// Define your task function
const myTask: ExperimentTask = async (item) => {
  const question = item.input;

const response = await observeOpenAI(new OpenAI()).chat.completions.create({
    model: "gpt-4.1",
    messages: [\
      {\
        role: "user",\
        content: question,\
      },\
    ],
  });

return response;
};

// Run the experiment
const result = await langfuse.experiment.run({
  name: "Geography Quiz",
  description: "Testing basic functionality",
  data: localData,
  task: myTask,
});

// Print formatted result
console.log(await result.format());

// Important: shut down OTEL SDK to deliver traces
await otelSdk.shutdown();
```

When running experiments on local data, only traces are created in Langfuse - no dataset runs are generated. Each task execution creates an individual trace for observability and debugging.
