LLM regression testing: fail CI before regressions ship - Langfuse

LLM Regression Testing: Fail CI Before Regressions Ship

A prompt tweak that fixes one complaint can quietly break ten other answers. LLM regression testing catches that break before it merges: every change to a prompt, model, or retrieval component runs against a fixed set of test cases, gets scored, and fails the CI pipeline when a score drops below a threshold. This guide shows the complete setup with Langfuse, from the gate script to the GitHub Actions workflow that blocks the pull request.

TL;DR: An LLM regression test runs your application against a golden dataset, scores every output with evaluators (deterministic code checks plus an LLM judge), and raises RegressionError when an aggregate score misses your threshold. In GitHub Actions, langfuse/experiment-action runs that script against a Langfuse dataset, posts the scores as a pull request comment, and fails the job on regression.

What is LLM Regression Testing?

LLM regression testing is the practice of verifying that a change to an LLM application did not degrade output quality on inputs that used to work. Use exact-match or deterministic assertions for structured outputs and business rules, and semantic evaluation when valid answers can differ in wording. Test the application function or workflow that produces the behavior you care about.

A regression gate should run whenever any component that shapes outputs changes:

This page covers the offline gate in depth. For where the gate sits inside a complete evaluation program, with quality dimensions, production monitoring, and human review, see the broader guide on building an LLM evaluation strategy.

The Minimal Regression Gate: Golden Dataset, Experiment, Threshold

Three pieces make a regression gate:

  1. A golden dataset of representative inputs with expected outputs, stored as a Langfuse dataset so it is versioned and shared.
  2. An experiment that runs your application against every dataset item and scores the outputs with evaluators.
  3. A threshold check that raises RegressionError when an aggregate score is too low, which is what fails the pipeline.
import json
import math

from langfuse import Evaluation, RegressionError, RunnerContext
from langfuse.openai import OpenAI

client = OpenAI()

THRESHOLDS = {
    "avg_contains_answer": 0.9,
    "avg_reference_correctness": 0.8,
}

# Define task
def support_agent_task(*, item, **kwargs):
    response = client.chat.completions.create(
        model="gpt-4.1",
        messages=[
            {"role": "system", "content": "Answer using only the provided context."},
            {"role": "user", "content": f"Context: {item.input['context']}\nQuestion: {item.input['question']}"},
        ],
    )
    return response.choices[0].message.content

# Define evaluators
def contains_answer(*, output, expected_output, **kwargs):
    passed = bool(expected_output) and expected_output.lower() in (output or "").lower()
    return Evaluation(name="contains_answer", value=1.0 if passed else 0.0)

# Run regression experiment
def experiment(context: RunnerContext):
    result = context.run_experiment(
        name="PR gate: support agent",
        task=support_agent_task,
        evaluators=[contains_answer],
    )
    scores = {e.name: e.value for e in result.run_evaluations}
    for metric, threshold in THRESHOLDS.items():
        value = scores.get(metric)
        if not isinstance(value, (int, float)) or not math.isfinite(value) or value < threshold:
            raise RegressionError(
                result=result,
                metric=metric,
                value=float(value) if isinstance(value, (int, float)) else 0.0,
                threshold=threshold,
            )
    return result

Fail CI With GitHub Actions: langfuse/experiment-action

langfuse/experiment-action is the official GitHub Action for running Langfuse experiments in CI. Given the script above, it does four things:

  1. Loads the dataset named in the workflow, optionally pinned to a dataset_version timestamp for reproducible runs.
  2. Executes every experiment script found at experiment_path (a file, directory, or glob; Python, TypeScript, and JavaScript are supported).
  3. Posts or updates a pull request comment with pass/regression/error status per script, run-level scores, an item-level results table, and a link to the experiment in Langfuse.
  4. Fails the job when a script raises RegressionError (should_fail_on_regression, default true) or crashes (should_fail_on_script_error, default true).
name: LLM regression tests

on:
  pull_request:

permissions:
  contents: read # check out the repository
  pull-requests: write # post the experiment result comment
  actions: read # optional: link results to this job's logs

jobs:
  regression-gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v6
      - uses: actions/setup-python@v6
        with:
          python-version: "3.14"
      - uses: langfuse/experiment-action@v1.0.6
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
        with:
          langfuse_public_key: ${{ secrets.LANGFUSE_PUBLIC_KEY }}
          langfuse_secret_key: ${{ secrets.LANGFUSE_SECRET_KEY }}
          langfuse_base_url: https://cloud.langfuse.com
          experiment_path: experiments/regression-gate.py
          dataset_name: support-agent-golden-set
          dataset_version: "2026-07-01T00:00:00Z"
          github_token: ${{ github.token }}

Comparing Regression Runs Over Time

Every gate run against a Langfuse dataset is recorded as a dataset run, so the history accumulates automatically. The experiment comparison view shows runs side by side: aggregate scores per run, per-item outputs, and which specific items regressed between two runs. When the PR comment says avg_reference_correctness dropped from 0.86 to 0.71, the comparison view answers the follow-up question of which ten items broke and what the outputs looked like.

FAQ

How many dataset items does a CI regression gate need?

Enough to cover your known failure modes, and small enough to run on every pull request: tens to low hundreds of items is the practical range for a PR gate.

How do I keep LLM judges from making CI flaky?

Make the judge as deterministic as you can: pin the judge model version, use a binary pass/fail verdict with an explicit rubric instead of a 1-to-10 scale, and request structured JSON output.

Can I migrate from Promptfoo to Langfuse regression gates?

Yes. The concepts map directly: Promptfoo test cases become Langfuse dataset items, assertions become evaluator functions, and the CI invocation becomes langfuse/experiment-action.