Troubleshooting and FAQ - Langfuse

We value your privacy

We use cookies to enhance your browsing experience, serve personalised ads or content, and analyse our traffic. By clicking "Accept All", you consent to our use of cookies.

CustomiseReject AllAccept All

Customise Consent Preferences

We use cookies to help you navigate efficiently and perform certain functions. You will find detailed information about all cookies under each consent category below.

The cookies that are categorised as "Necessary" are stored on your browser as they are essential for enabling the basic functionalities of the site. ... Show more

NecessaryAlways Active

Necessary cookies are required to enable the basic features of this site, such as providing secure log-in or adjusting your consent preferences. These cookies do not store any personally identifiable data.

__Secure-next-auth.csrf-token.EU

session

Description is currently not available.

__Secure-next-auth.callback-url.EU

session

Description is currently not available.

__Secure-next-auth.csrf-token.US

session

Description is currently not available.

__Secure-next-auth.callback-url.US

session

Description is currently not available.

__cf_bm

1 hour

This cookie, set by Cloudflare, is used to support Cloudflare Bot Management.

__hssrc

session

This cookie is set by Hubspot whenever it changes the session cookie. The __hssrc cookie set to 1 indicates that the user has restarted the browser, and if the cookie does not exist, it is assumed to be a new session.

__hssc

1 hour

HubSpot sets this cookie to keep track of sessions and to determine if HubSpot should increment the session number and timestamps in the __hstc cookie.

__Secure-next-auth.csrf-token.HIPAA

session

Description is currently not available.

__Secure-next-auth.callback-url.HIPAA

session

Description is currently not available.

__Secure-next-auth.csrf-token.JP

session

Description is currently not available.

__Secure-next-auth.callback-url.JP

session

Description is currently not available.

theme

Never Expires

No description available.

cookieyes-*

1 year

CookieYes sets this cookie for consent solution management.

Functional

Functional cookies help perform certain functionalities like sharing the content of the website on social media platforms, collecting feedback, and other third-party features.

inkeepUsagePreferences_userId

1 year

Description is currently not available.

_octo

1 year

No description available.

logged_in

1 year

No description available.

Analytics

Analytical cookies are used to understand how visitors interact with the website. These cookies help provide information on metrics such as the number of visitors, bounce rate, traffic source, etc.

ph_phc_zkMwFajk8ehObUlMth0D7DtPItFnxETi3lmSvyQDrwB_posthog

1 year

Description is currently not available.

__hstc

6 months

Hubspot set this main cookie for tracking visitors. It contains the domain, initial timestamp (first visit), last timestamp (last visit), current timestamp (this visit), and session number (increments for each subsequent session).

hubspotutk

6 months

HubSpot sets this cookie to keep track of the visitors to the website. This cookie is passed to HubSpot on form submission and used when deduplicating contacts.

_gh_sess

session

GitHub sets this cookie for temporary application and framework state between pages like what step the user is on in a multiple step form.

Performance

Performance cookies are used to understand and analyse the key performance indexes of the website which helps in delivering a better user experience for the visitors.

__sdcfduid

5 years

Stores a unique browser identifier to help Discord detect malicious activity, enforce security protections, manage traffic, and maintain service performance.

__dcfduid

5 years

Enables Discord to uniquely identify your device, enhance account security, prevent spam and abuse, and support diagnostics across browsing sessions.

Advertisement

Advertisement cookies are used to provide visitors with customised advertisements based on the pages you visited previously and to analyse the effectiveness of the ad campaigns.

No cookies to display.

Uncategorised

Other uncategorised cookies are those that are being analysed and have not been classified into a category as yet.

No cookies to display.

Reject AllSave My PreferencesAccept All

DocsTroubleshooting and FAQ

Docs Evaluation Troubleshooting and FAQ

Copy page

Troubleshooting and FAQ

This page addresses frequently asked questions and common troubleshooting topics for Langfuse Evaluation.

If you don't find a solution to your issue here, try using Ask AI for instant answers. For bug reports, please open a ticket on GitHub Issues. For general questions or support, visit our support page.

FAQ

GitHub Discussions

Playground Experiment with HttpStreamable MCP Support Dataset Selection search on Run experiment modal Add result column on the Annotation queue items list Support different types of dataset, like agentic dataset gennerate a friendly dataset item ids feat: Rule-Based Evaluators (Contribution Proposal) feat: Public API for Experiment Runs + GitHub Action (Contribution Proposal) Dataset couldn't move/copy directly from one project to another Feature Request: Display Full Session Context When Annotating Individual Observations in Annotation Queue Feature Request: Support Full Session Context in Annotation Queue for Individual Traces/Observations Add pattern-based trace selection for LLM-as-a-Judge evaluator targets (probably can be added everywhere where this component is used) Can I assign users at trace level in human annotation queue to allow seamless delegation within a annotation queue. Dataset Retrieval Is Not Streamed / Paginated Support for unresolved/literal prompt variables in UI-based Prompt Experiments for custom LLM wrapper APIs Transpose view for dataset run comparison table [Feature Request] Support linking Ground Truth from Datasets in Live Tracing Evaluators & Registering external evaluators in UI The corrected data cannot be imported into the dataset. feat(cookbook): Proposal for "AI Social Engineering" Evaluation Cookbook (based on CPF Framework) LLM-as-a-Judge: Add support for filtering traces by input/output content API/SDK: add bulk dataset item insert to improve versioning Add annotation queue dropdown to traces in session view Track "Added By" and "Added At" for annotation queue items Add annotation and annotation queue filters to trace list view Attaching Guardrail policy to the LLM connection which can be attached to the evaluator in LLM-as-a-Judge. Display "Expected Output" within Annotation Queues for Side-by-Side Comparison Allow for Bulk Deletions inside Datasets Items view Fetching all traces along with Score data Integrate BLOOM evaluation framework for LLM safety and capability assessment Feature Request: Add API endpoint to fetch all dataset items with metadata filtering support Evaluators Running Twice on Parent Trace and Child Observations - Creating Duplicates Save/Pre-Populate Add To Dataset Form Values UX Suggestions in Experiments Flow Edit dataset item IDs via UI or add description field Add functionality to add a regex filter on output of LLM when creating LLM as a judge Multi-turn Conversation Evaluation Make the output_schema stored in the prompt config in the playground and in experiments Gemini 3 thinking_level Support Aggregated Evaluation Scores at Conversation / Agent Level Add datasetItemsUpdate method to SDK Support dataset splits (train/dev/test) in the Experiment Runner SDK Create a custom evaluator which runs on coding log Better handling of score -1 Support for Statistical Aggregation (mean, variance...) across dataset runs Support Dataset edit and deletion via API Dataset Item Creation Form Batch/Multi-Trace LLM Evaluators for Pattern Discovery AI-Generated Summary of Evaluator Reasoning at Dataset Level LLM-as-a-Judge on observation Make datasets.upsertRemoteExperiment API available in SDK LLM as a Judge Sequential Chaining Resolve @@@langfuseMedia:...@@@ placeholders as base64 Improve the UX the comment section for LLMaJ evaluators. Enhanced Filtering in LLM-as-a-Judge Score Tables Have LLM as a judge Evaluator in folder structure ( using / ) Filter & Sort by non-score columns in Compare View Set a prefix for the Dataset item id Add answer-generation as part of Dataset-/Experiment-UI Apply Score Analytics to Dataset Runs feat(Datasets): Filter Dataset runs by metadata / tags feat(annotation): support defining annotation queue item processing order by reference object ids Allow filtering by environment in score analytics Add Human Annotation Queue via sdk API to Fetch Custom LLM-as-a-Judge Evaluator Prompts from Evaluator Library Edit output Friendly Allow filtering Scores by trace metadata Add total results count to Scores and Traces views Create dataset with variables from traces Support multiple metrics in a single LLM Judge output Automatically add traces to dataset based on score Datasets - excepted output can't parse new lines feat(dataset): Add a dataset prefix as the dataset item id Feedback on the prompt experiments and eval settings UX Add ability to add arbitrary code to run evaluations feat(datasets): support tags on datasets Run level score charts not visible Side-by-Side Comparison in Annotation Queues Feature Request: Show Diff Between Expected and Actual Output in Dataset Run Feature Request: Custom Logic for Prompt Experiment Scoring Feature Request: Enhanced Dataset and Dataset Item Operations Separate datasets in folders Dataset Runs: Show run-level metrics in sidebar Provide renaming functionality for datasets in Langfuse public API feat(llm-completion): support Anthropic-specific message roles in ChatML Create a Run in Langfuse even when an evaluation fails before run_experiment() Flow: tag and annotate traces and then find what you annotated. Add Sessions to Datasets Extract human annotation queues to csv Request feature: allow Scores to be able to filter with trace's "Release" feat(datasets): keep dataset items created from traces in sync with trace state feat(experiment-compare): highlighting or ability to sort by detected differences between scores would be nice for locating significant changes Aggregate boolean scores as success ratio See LLM Generation when using llm-as-a-judge Add support for unwrapping jsonpath array output / support jq Add UPDATE (PATCH) method for Score Configs in Python SDK v3 and DELETE endpoint in API Allow getting evaluator ID from score Object created by LLM-as-a-Judge evaluator feat(datasets): ability to delete evaluator from evaluator library 提供删除整个数据集的API langfuse API lacks a deletion API for Datasets. Feature Request: Export Human Annotation Comments from Human Annotation Page openapi: untyped DatasetItem.input LLM-as-a-Judge Evaluations to Support Complex Filters & Variable Mapping Bulk archive dataset items OR "archive all but current item" option [Feature Request] Dataset with more flexible views Given a sessionId, if we can FETCH the scores ( via api/sdk) associated with that session ( not its traces) Allow triggering Prompt Experiments programmatically (via public API) Search by trace, order by user in the annotation queue and add labels Add UI support for LLM-as-a-Judge evaluation on CSV datasets containing pre-existing input/output pairs without requiring re-generation Allow configuring iteration count per dataset item for dataset runs via SDK Feature request : Enable adding API key for Remote Dataset Runs Dataset Transformations Human-readable LangGraph traces in Langfuse (Formatted Mode) Support for YouTube Video URL Inputs in Datasets Python SDK: Pass ScoreConfig instead of config_id when creating scores Filter Support for Dataset Runs Request batch adding of multiple traces to datasets How to categorize topics and create a pie chart? Simultaneous Session and Trace Annotation for Context-Rich Feedback in the same annotation window feat(LLM-as-a-judge): support stratified sampling by trace property Bulk Evaluation Does not Displays the status of the dataset (Like for E.g. processing, evaluated, executed) Annotation Queue Creation API [Langfuse Cloud] Missing SessionId and Author in exported Scores feat(evals): allow canceling a running evaluator with pending evaluation jobs Support walking through inner spans in langfuse SDK for e2e trace evaluation Score Configs: Allow editing the categories of a categorical score Trials for Dataset Runs: Multiple runs on the same item and input to get robust scores Configuring evaluators (LLM-as-a-Judge) via API or SDK (CRUD) UI-LLM as a Jury Ability to download the evaluation logs (LLM as a Judge Logs) Knowledge base (source material) for the LLM-as-a-judge feature Option to run Experiments without traces Auto-generate dataset items feature: support for creation of custom) model adapters Prompt experiment result download. Multiple predictions per-item in a single run Add evaluator(s) to python SDK (FernLangfuse) feat(dataset-runs): return all dataset run item scores in a given run Delete an evaluator from the evaluator library Archive or Delete Evaluators from the Evaluator Library Edit dataset run name, description Dataset UI improvements for easier item management Batch-add traces to datasets Enable Immediate Score Management for User Feedback Export Dataset runs and run items feat(experiments/evaluation): support user defined exponential backoff to avoid hitting LLM tier limits support trajectory LLM-as-judge feat: allow redirect from evaluator logs to filter traces for evaluator score values Annotation Queue for Sessions Filter by scores in session view Support break lines on evaluation run tooltip hint "Create new evaluator" support Qwen model Exporting and sorting evaluators Optional schema definition for datasets feat(datasets): allow deleting via API Versioning for datasets Support new lines when storing / displaying score comments feat(annotation): allow for conditional score configurations Delete multiple dataset items Multi-label scores during human annotation Evaluator: Filter for Scores Re-run LLM-as-a-judge evaluations response_format for experiment runs Code based custom evaluators Support Description on Dataset Items Run prompt experiment with structured output / tool instruction Alta Integration Add run name to columns when looking at a specific dataset item Delete multiple dataset runs Navigation between items in a dataset run is confusing - context of the selected dataset run is lost Enhanced score distribution visualization in experiment analysis Multi-step Prompt Experiments and Playground Option to add trace to new dataset Simplified UI for Scoring More scoring configurations on UI Expose dataset item status (Archived or Active) on the item detail page Expose CRUD API for dataset run items feat [UI] - remember selected charts for Datasets Edit Name of Eval Templates CRUD Evaluation Templates via the API Evals: run an evaluator on a dataset without also running a prompt experiment chart request - mirror score functionality but after filtering within trace or observation panes UI/UX: Scores Analytics - reduce interaction friction LLM as a judge: only execute judge if a specific observation exists Update scores via the UI even if they have been created via the API Trajectory Evaluation - Evaluating path chooses by Agent LLM-as-a-Judge: Categorical and Boolean Scores [Dataset Run] Add filtering functionality for some columns Ability to run evaluations using prompt from prompt editor with content from spans "Add to Dataset" : augment ground truth labelling with custom AI-assistant Allow using User-defined models in Playground & Evals Ability to edit the LLM API keys API for annotation queues to programmatically manage them (enqueue, dequeue, status) Add traces to datasets from the trace table feat(experiments/evals): Regex based evaluators Allow Gemini access via Google AI (not Vertex AI) ui: ability to highlight scores Run prompt experiments via the SDKs/API Custom Non-LLM evaluators/scores through UI Side-by-side comparison in playground Support user input, message history, structured output, and tool calls in prompt experiments Datasets: Add selection of traces to a dataset Multi-user annotation capability in Annotation Queues Multi-turn / session experiments in datasets Enable to use variable of prompt on evaluator. Sessions Table: Scores Column Add new filters for the LLM as a Judge Evaluation (other scores and cost) Export dataset run table feat: support adding trace tags in annotation queue view Feat: De-dupe and show info when adding the same trace/observation to an annotation queue multiple times Diff support for dataset runs view Create Support for gemini models in playground Change AWS access pattern for Bedrock LLM usage, assume role Add ability to export and import evaluators between projects feat: Folder structure for dataset organisation Model-based evaluations triggered by observations Scores: Conditional Annotation Annotation Queues: define optional/mandatory score configs by queue Scores: support for recording multiple choice selection as score value Filter by status in dataset items table Versioning/edit of Evaluator Configuration Support annotation role permission to only access Annotation Queue feature Vector search for datasets for few-shot prompting Filtering dataset items Export Datasets via CSV Only use traces in the dataset to render score columns Source code versioning and automatic statistics caluclation from scores Include metrics and scores in getDatasetRun / `GET /api/public/datasets/{datasetName}/runs/{runName}` API/SDK to create comments on traces Need to retain the old evaluation history results, including input and all ground truth Support for Selecting Specific Paths from JSON Objects in Evaluation Prompts Dataset run description and metadata should not be fixed on the page Optionally set timestamp when creating a score Implement dataset removal method Compare two dataset runs side by side Add Bedrock Guardrails to the LLM Security documentation Set variables in playground from dataset Session-level scores Feature Request: Sidebar Display for Trace Details in Dataset Runs Scoring dataset runs, e.g. precision, recall, f-value Adding userId / author to score (custom metadata) Expand all json-views of Dataset items etc. Add string data type in score config SBS Markdown mode for dataset runs Proposal: Add Support for Uploading Dataset Items via UI Save playground conversation to a dataset Upload datasets via UI API/UI to delete dataset items and runs Support Linking Execution Trace to DatasetItem without Fetching Entire Dataset annotation of traces in langfuse console? What standardized dataset formats are people using? API to delete scores Datasets: Diff of output and expected output Run level confusion matrix via UI Skipping an item in LLM-as-a-judge eval How to debug when evaluations are stalling Cannot access 'corrected_output' from a trace via the SDK or API Experiments via UI for Agent Server? How to filter multiple userids in the evaluator? Inquiry regarding multiple Human Annotations per single trace Inquiry regarding the execution logic and concurrency control of UI-configured Evaluators Fetch Evaluator Library (Eval Templates) via Public API/SDK How can I update datasetItems via the API? (using typescript SDK) Same Evaluator for Live/Dataset Runs but different vars How to emulate same trace format in experiment run in sdk as prompt experiment in ui. Sequence of the query level evaluators Datasets if I run single or batch some questions Will LLM-as-a-Judge Process Uploaded Files in a Multi-LLM Chat Application? How to filter LLM-as-a-Judge evaluator to run only on traces with specific input content? Dataset run items showing identical scores despite unique observation-level scores Dataset Experiment Run - Difficulty Running Judge Prompt for Multiple Questions in a Single Call Score analytics - beta feature review Langfuse Eval workflow Bulk Add Observations to Human Evaluations with Fine-Grained Scoring Aggregated metrics for experiments Metrics & Scores Datasets evaluattion twice Custom tools for LLM-as-a-Judge Evaluation Evaluation Fails When Extracted Generation Parameters Are Not Invoked in Trace Listening to eval scores feat(experiments): Support splitting task JSON response into output and metadata fields in dataset runs My evaluator is not running How to extract dynamic "last user message" from chat history for Evaluation? datasets evaluator filtering options? Best practice for annotating offline human–human conversations in Langfuse Dataset metadata not available during generation in Experiments (Self-Hosted) Retrieving scores updated in the UI via the API Pass the complete table data for aggregate queries 删除整个数据集 Using Langfuse Prompt Management with Bedrock Invoke Model instead of Converse API Add score configs in an specific order in the annotation queue API for existing evaluators How to Evaluate an External API How to evaluate nested spans in dataset.run_experiment() evaluators? Public API does not allow filter by scores Allow defining user ID when creating score via SDK Using both SDK Experiments + Configured LLM-as-a-Judge Together Programatically assigning scores to each item in annotation queue using langfuse sdk Issues when running an evaluator on existing traces ordering dataset items within experiment results by a number that is in the llm response for each item? How do you link Scores to a Prompt Langfuse cloud traces for dataset experiments LLM-as-a-Judge format JSON DataSet Items api doesn't work For depeval evaluation , do we need a real time expected output for evaluating some metrics like g-eval, answer relavancy , hallucination etc. How are evaluator prompts combined and sent to the LLM in LLM-as-a-judge? [question] Custom Model Name + No Actual run + Annotation Is it possible to get judge evaluation result for dataset run item with SDK/REST api ? How to add expected output for Geval? “Body exceeded 1mb limit” error on dataset upload Explicit Feedback - Can a User Researcher add this on behalf of a user? Multi-step metrics from Ragas How to evaluate correct tool calls are being made? How to evaluate existing traces? IMAGE-PROCESSING-USING-GEMINI Ability to View Score Comments in LangFuse UI Evaluation, between two outputs. Dataset run fails with LLM-as-judge against a chat prompt with placeholders 删除整个数据集 Update dataset scores Dataset Experiment via SDK CreateScoreRequest.DatasetRunId not working? how to use Evaluator with problem, answer, and Expected output? How to fetch the traces using custom fields ingested in dashboard? Is it possible for LLM-as-a-Judge to produce a categorical answer? How to link an AI-as-judge eval to a specific observation in Langfuse Cloud Given a sessionId, how can i fetch the scores associated with that session ( not its traces) Dataset run on Composite Prompts SDK支持编辑单个数据 Comparing the LLM as judge prompts on the UI Scores – limitations and improvement suggestions Automating eval runs API - /scores not respecting value when operator '=' Does LangFuse support evaluations on an existing dataset (.csv) deleting evaluators Not able to import my LLM-as-a-Judge evals How to get experiment run scores programmatically? Scoring using span object or using trace id doesn't seem to work in remote runs for dataset evaluations Dataset runs restore from backups Getting scores efficiently via API for analytics purposes run experiment on dataset Experiments on Datasets with Human Annotated Labels? How to Recalculate Total Score on Dashboard After Updating User-Defined Model? How to filter by Categorical Scores in custom dashboard? Running scheduled evals utilising LangFuse Datasets & Evaluators Custom trace_id for traces created with dataset's `item.run()` method how to create langfuse datasetrun in langfuse-java sdk? [Experiment] Send items in parallel. Stop experiment. Focused Mode shows only the final ai.generateObject span—earlier generations are missing Using scores in data sets from the SDK? Results for some data items not present when comparing experiments Deleting Metrics for Langfuse Discrepancies between dataset items found in the UI vs retrieved from the SDK/API Datasets Not Fully Aligned - Dataset Item Runs and Trace Dataset Associated Is there an option to text wrap the output responses while comparing dataset runs? Score function Support for Metric Calculation (Precision@K, Recall@K) and Adding Custom Metrics Use Case Overview Unable to link dataset run items to generation observation when using observe decorator Dataset creates a lot of identical processes that run infinitely. How to deactivate evaluator in llm-as-a-judge programmatically (through SDK or API)? Prompt Experiments Not Generating Traces Templates for LLM-AS-a-judge Take chat history in consideration when running a prompt experiment Cannot use prompt experiments "No dataset item contains any variables" How to update a score? Can I evaluate Span using the External Evaluation Pipeline? Allow injection of generation context into evaluation prompt Weird behaviour of metric values and their reasoning Filter Categorical Score Values How to create an eval config for prompts using python api? Can I use Amazon Bedrock for Langfuse Evals? Configuring Evaluation with "Correctness" Template & Python Code Invocation Self-Host evaluation feature Hey want to change the Eval Templates name, can we do it from UI DataSet Scores are not being displayed How run Langfuse evaluations over specifics spans? Getting all traces logged in a timerange for custom scoring 2 traces generated instead of 1 Evaluations Not Available in Self-Hosted Version? Deleting Duplicate Items in a Dataset Availability of evals when self-hosting How to utilize a dataset w/ typescript and langchain integration Scoring a trace after the LLM chain returns Update/delete score using python sdk Linking dataset run items with existing callback handler Datasets list / by id Run items not appearing when linking to a trace and not a span or a generation Cannot see Add to Dataset button in the UI

GitHubSupportGitHubIdeas

Upvotes

NewGitHub

Discussions last updated: 7/18/2026, 2:42:32 AM (55 hours ago)


Was this page helpful?

Good

Bad

Support


Last edited 3/20/2026

PreviousExperiments in CI/CD NextOverview

On this page

Troubleshooting and FAQ FAQ GitHub Discussions

Actions

Give us feedback Edit this page on GitHub

Contributors

Last edited Mar 20, 2026

Jannik Maierhöfer](https://github.com/jannikmaierhoefer) Marc Klingen](https://github.com/marcklingen)

GitHub X

Ask AI A

A