LLM regression testing: fail CI before regressions ship - Langfuse
LLM Regression Testing: Fail CI Before Regressions Ship
A prompt tweak that fixes one complaint can quietly break ten other answers. LLM regression testing catches that break before it merges: every change to a prompt, model, or retrieval component runs against a fixed set of test cases, gets scored, and fails the CI pipeline when a score drops below a threshold. This guide shows the complete setup with Langfuse, from the gate script to the GitHub Actions workflow that blocks the pull request.
TL;DR: An LLM regression test runs your application against a golden dataset, scores every output with evaluators (deterministic code checks plus an LLM judge), and raises RegressionError when an aggregate score misses your threshold. In GitHub Actions, langfuse/experiment-action runs that script against a Langfuse dataset, posts the scores as a pull request comment, and fails the job on regression.
What is LLM Regression Testing?
LLM regression testing is the practice of verifying that a change to an LLM application did not degrade output quality on inputs that used to work. Use exact-match or deterministic assertions for structured outputs and business rules, and semantic evaluation when valid answers can differ in wording. Test the application function or workflow that produces the behavior you care about.
A regression gate should run whenever any component that shapes outputs changes:
- Prompt edits, including a new version of a prompt managed outside the codebase.
- Model swaps and provider model upgrades, where the same prompt can behave differently.
- RAG and retrieval changes, such as a new chunking strategy or embedding model.
- Tool and agent changes, where a reworded tool description shifts an agent's behavior.
This page covers the offline gate in depth. For where the gate sits inside a complete evaluation program, with quality dimensions, production monitoring, and human review, see the broader guide on building an LLM evaluation strategy.
The Minimal Regression Gate: Golden Dataset, Experiment, Threshold
Three pieces make a regression gate:
- A golden dataset of representative inputs with expected outputs, stored as a Langfuse dataset so it is versioned and shared.
- An experiment that runs your application against every dataset item and scores the outputs with evaluators.
- A threshold check that raises
RegressionErrorwhen an aggregate score is too low, which is what fails the pipeline.
import json
import math
from langfuse import Evaluation, RegressionError, RunnerContext
from langfuse.openai import OpenAI
client = OpenAI()
THRESHOLDS = {
"avg_contains_answer": 0.9,
"avg_reference_correctness": 0.8,
}
# Define task
def support_agent_task(*, item, **kwargs):
response = client.chat.completions.create(
model="gpt-4.1",
messages=[
{"role": "system", "content": "Answer using only the provided context."},
{"role": "user", "content": f"Context: {item.input['context']}\nQuestion: {item.input['question']}"},
],
)
return response.choices[0].message.content
# Define evaluators
def contains_answer(*, output, expected_output, **kwargs):
passed = bool(expected_output) and expected_output.lower() in (output or "").lower()
return Evaluation(name="contains_answer", value=1.0 if passed else 0.0)
# Run regression experiment
def experiment(context: RunnerContext):
result = context.run_experiment(
name="PR gate: support agent",
task=support_agent_task,
evaluators=[contains_answer],
)
scores = {e.name: e.value for e in result.run_evaluations}
for metric, threshold in THRESHOLDS.items():
value = scores.get(metric)
if not isinstance(value, (int, float)) or not math.isfinite(value) or value < threshold:
raise RegressionError(
result=result,
metric=metric,
value=float(value) if isinstance(value, (int, float)) else 0.0,
threshold=threshold,
)
return result
Fail CI With GitHub Actions: langfuse/experiment-action
langfuse/experiment-action is the official GitHub Action for running Langfuse experiments in CI. Given the script above, it does four things:
- Loads the dataset named in the workflow, optionally pinned to a
dataset_versiontimestamp for reproducible runs. - Executes every experiment script found at
experiment_path(a file, directory, or glob; Python, TypeScript, and JavaScript are supported). - Posts or updates a pull request comment with pass/regression/error status per script, run-level scores, an item-level results table, and a link to the experiment in Langfuse.
- Fails the job when a script raises
RegressionError(should_fail_on_regression, defaulttrue) or crashes (should_fail_on_script_error, defaulttrue).
name: LLM regression tests
on:
pull_request:
permissions:
contents: read # check out the repository
pull-requests: write # post the experiment result comment
actions: read # optional: link results to this job's logs
jobs:
regression-gate:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v6
- uses: actions/setup-python@v6
with:
python-version: "3.14"
- uses: langfuse/experiment-action@v1.0.6
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
with:
langfuse_public_key: ${{ secrets.LANGFUSE_PUBLIC_KEY }}
langfuse_secret_key: ${{ secrets.LANGFUSE_SECRET_KEY }}
langfuse_base_url: https://cloud.langfuse.com
experiment_path: experiments/regression-gate.py
dataset_name: support-agent-golden-set
dataset_version: "2026-07-01T00:00:00Z"
github_token: ${{ github.token }}
Comparing Regression Runs Over Time
Every gate run against a Langfuse dataset is recorded as a dataset run, so the history accumulates automatically. The experiment comparison view shows runs side by side: aggregate scores per run, per-item outputs, and which specific items regressed between two runs. When the PR comment says avg_reference_correctness dropped from 0.86 to 0.71, the comparison view answers the follow-up question of which ten items broke and what the outputs looked like.
FAQ
How many dataset items does a CI regression gate need?
Enough to cover your known failure modes, and small enough to run on every pull request: tens to low hundreds of items is the practical range for a PR gate.
How do I keep LLM judges from making CI flaky?
Make the judge as deterministic as you can: pin the judge model version, use a binary pass/fail verdict with an explicit rubric instead of a 1-to-10 scale, and request structured JSON output.
Can I migrate from Promptfoo to Langfuse regression gates?
Yes. The concepts map directly: Promptfoo test cases become Langfuse dataset items, assertions become evaluator functions, and the CI invocation becomes langfuse/experiment-action.