We value your privacy

We use cookies to enhance your browsing experience, serve personalised ads or content, and analyse our traffic. By clicking "Accept All", you consent to our use of cookies.

CustomiseReject AllAccept All

Customise Consent Preferences

We use cookies to help you navigate efficiently and perform certain functions. You will find detailed information about all cookies under each consent category below.

The cookies that are categorised as "Necessary" are stored on your browser as they are essential for enabling the basic functionalities of the site. ... Show more

NecessaryAlways Active

Necessary cookies are required to enable the basic features of this site, such as providing secure log-in or adjusting your consent preferences. These cookies do not store any personally identifiable data.

- Cookie

\_\_Secure-next-auth.csrf-token.EU

- Duration

session

- Description

Description is currently not available.

- Cookie

\_\_Secure-next-auth.callback-url.EU

- Duration

session

- Description

Description is currently not available.

- Cookie

\_\_Secure-next-auth.csrf-token.US

- Duration

session

- Description

Description is currently not available.

- Cookie

\_\_Secure-next-auth.callback-url.US

- Duration

session

- Description

Description is currently not available.

- Cookie

\_\_cf\_bm

- Duration

1 hour

- Description

This cookie, set by Cloudflare, is used to support Cloudflare Bot Management.

- Cookie

\_\_hssrc

- Duration

session

- Description

This cookie is set by Hubspot whenever it changes the session cookie. The \_\_hssrc cookie set to 1 indicates that the user has restarted the browser, and if the cookie does not exist, it is assumed to be a new session.

- Cookie

\_\_hssc

- Duration

1 hour

- Description

HubSpot sets this cookie to keep track of sessions and to determine if HubSpot should increment the session number and timestamps in the \_\_hstc cookie.

- Cookie

\_\_Secure-next-auth.csrf-token.HIPAA

- Duration

session

- Description

Description is currently not available.

- Cookie

\_\_Secure-next-auth.callback-url.HIPAA

- Duration

session

- Description

Description is currently not available.

- Cookie

\_\_Secure-next-auth.csrf-token.JP

- Duration

session

- Description

Description is currently not available.

- Cookie

\_\_Secure-next-auth.callback-url.JP

- Duration

session

- Description

Description is currently not available.

- Cookie

theme

- Duration

Never Expires

- Description

No description available.

- Cookie

cookieyes-\*

- Duration

1 year

- Description

CookieYes sets this cookie for consent solution management.

Functional

Functional cookies help perform certain functionalities like sharing the content of the website on social media platforms, collecting feedback, and other third-party features.

- Cookie

inkeepUsagePreferences\_userId

- Duration

1 year

- Description

Description is currently not available.

- Cookie

\_octo

- Duration

1 year

- Description

No description available.

- Cookie

logged\_in

- Duration

1 year

- Description

No description available.

Analytics

Analytical cookies are used to understand how visitors interact with the website. These cookies help provide information on metrics such as the number of visitors, bounce rate, traffic source, etc.

- Cookie

ph\_phc\_zkMwFajk8ehObUlMth0D7DtPItFnxETi3lmSvyQDrwB\_posthog

- Duration

1 year

- Description

Description is currently not available.

- Cookie

\_\_hstc

- Duration

6 months

- Description

Hubspot set this main cookie for tracking visitors. It contains the domain, initial timestamp (first visit), last timestamp (last visit), current timestamp (this visit), and session number (increments for each subsequent session).

- Cookie

hubspotutk

- Duration

6 months

- Description

HubSpot sets this cookie to keep track of the visitors to the website. This cookie is passed to HubSpot on form submission and used when deduplicating contacts.

- Cookie

\_gh\_sess

- Duration

session

- Description

GitHub sets this cookie for temporary application and framework state between pages like what step the user is on in a multiple step form.

Performance

Performance cookies are used to understand and analyse the key performance indexes of the website which helps in delivering a better user experience for the visitors.

- Cookie

\_\_sdcfduid

- Duration

5 years

- Description

Stores a unique browser identifier to help Discord detect malicious activity, enforce security protections, manage traffic, and maintain service performance.

- Cookie

\_\_dcfduid

- Duration

5 years

- Description

Enables Discord to uniquely identify your device, enhance account security, prevent spam and abuse, and support diagnostics across browsing sessions.

Advertisement

Advertisement cookies are used to provide visitors with customised advertisements based on the pages you visited previously and to analyse the effectiveness of the ad campaigns.

No cookies to display.

Uncategorised

Other uncategorised cookies are those that are being analysed and have not been classified into a category as yet.

No cookies to display.

Reject AllSave My PreferencesAccept All

DocsTroubleshooting and FAQ

[Docs](/content/docs/index.html) [Evaluation](/content/docs/evaluation/overview/index.html) [Troubleshooting and FAQ](/content/docs/evaluation/troubleshooting-and-faq/index.html)

Copy page

# [Troubleshooting and FAQ](/content/docs/evaluation/troubleshooting-and-faq\#troubleshooting-and-faq/index.html)

This page addresses frequently asked questions and common troubleshooting topics for Langfuse Evaluation.

If you don't find a solution to your issue here, try using [Ask AI](/content/docs/ask-ai/index.html) for instant answers. For bug reports, please open a ticket on [GitHub Issues](/content/issues/index.html). For general questions or support, visit our [support page](/content/support/index.html).

## [FAQ](/content/docs/evaluation/troubleshooting-and-faq\#faq/index.html)

- [How do I upgrade my trace-level evaluators to observation-level evaluators?](/content/faq/all/llm-as-a-judge-migration/index.html)
- [How do I work with observations in Langfuse Fast Preview (Langfuse v4)?](/content/faq/all/explore-observations-in-v4/index.html)
- [How to create and manage Score Configs in Langfuse?](/content/faq/all/manage-score-configs/index.html)
- [How to evaluate sessions/conversations?](/content/faq/all/evaluating-sessions-conversations/index.html)
- [How to retrieve experiment scores via UI or API/SDK?](/content/faq/all/retrieve-experiment-scores/index.html)
- [How to use Langfuse-hosted Evaluators on Dataset Runs?](/content/faq/all/langfuse-evaluators-on-dataset-runs/index.html)
- [I have setup Langfuse, but I do not see any traces in the dashboard. How to solve this?](/content/faq/all/missing-traces/index.html)
- [What are scores in Langfuse and when should I use them?](/content/faq/all/what-are-scores/index.html)
- [Why do observation lookups with the v3 SDK return 404?](/content/faq/all/v3-sdk-observation-lookup-404/index.html)
- [Why is my observation-level evaluator not executing?](/content/faq/all/observation-eval-not-executing/index.html)

## [GitHub Discussions](/content/docs/evaluation/troubleshooting-and-faq\#github-discussions/index.html)

[Playground Experiment with HttpStreamable MCP Support](https://github.com/orgs/langfuse/discussions/12030 "Langfuse Ideas: Playground Experiment with HttpStreamable MCP Support") [Dataset Selection search on Run experiment modal](https://github.com/orgs/langfuse/discussions/12009 "Langfuse Ideas: Dataset Selection search on Run experiment modal") [Add result column on the Annotation queue items list](https://github.com/orgs/langfuse/discussions/11983 "Langfuse Ideas: Add result column on the Annotation queue items list") [Support different types of dataset, like agentic dataset](https://github.com/orgs/langfuse/discussions/11956 "Langfuse Ideas: Support different types of dataset, like agentic dataset") [gennerate a friendly dataset item ids](https://github.com/orgs/langfuse/discussions/11955 "Langfuse Ideas: gennerate a friendly dataset item ids") [feat: Rule-Based Evaluators (Contribution Proposal)](https://github.com/orgs/langfuse/discussions/11951 "Langfuse Ideas: feat: Rule-Based Evaluators (Contribution Proposal)") [feat: Public API for Experiment Runs + GitHub Action (Contribution Proposal)](https://github.com/orgs/langfuse/discussions/11949 "Langfuse Ideas: feat: Public API for Experiment Runs + GitHub Action (Contribution Proposal)") [Dataset couldn't move/copy directly from one project to another](https://github.com/orgs/langfuse/discussions/11932 "Langfuse Ideas: Dataset couldn't move/copy directly from one project to another") [Feature Request: Display Full Session Context When Annotating Individual Observations in Annotation Queue](https://github.com/orgs/langfuse/discussions/11927 "Langfuse Ideas: Feature Request: Display Full Session Context When Annotating Individual Observations in Annotation Queue") [Feature Request: Support Full Session Context in Annotation Queue for Individual Traces/Observations](https://github.com/orgs/langfuse/discussions/11926 "Langfuse Ideas: Feature Request: Support Full Session Context in Annotation Queue for Individual Traces/Observations") [Add pattern-based trace selection for LLM-as-a-Judge evaluator targets (probably can be added everywhere where this component is used)](https://github.com/orgs/langfuse/discussions/11914 "Langfuse Ideas: Add pattern-based trace selection for LLM-as-a-Judge evaluator targets (probably can be added everywhere where this component is used)") [Can I assign users at trace level in human annotation queue to allow seamless delegation within a annotation queue.](https://github.com/orgs/langfuse/discussions/11902 "Langfuse Ideas: Can I assign users at trace level in human annotation queue to allow seamless delegation within a annotation queue.") [Dataset Retrieval Is Not Streamed / Paginated](https://github.com/orgs/langfuse/discussions/11869 "Langfuse Ideas: Dataset Retrieval Is Not Streamed / Paginated") [Support for unresolved/literal prompt variables in UI-based Prompt Experiments for custom LLM wrapper APIs](https://github.com/orgs/langfuse/discussions/11806 "Langfuse Ideas: Support for unresolved/literal prompt variables in UI-based Prompt Experiments for custom LLM wrapper APIs") [Transpose view for dataset run comparison table](https://github.com/orgs/langfuse/discussions/11804 "Langfuse Ideas: Transpose view for dataset run comparison table") [\[Feature Request\] Support linking Ground Truth from Datasets in Live Tracing Evaluators & Registering external evaluators in UI](https://github.com/orgs/langfuse/discussions/11735 "Langfuse Ideas: [Feature Request] Support linking Ground Truth from Datasets in Live Tracing Evaluators & Registering external evaluators in UI") [The corrected data cannot be imported into the dataset.](https://github.com/orgs/langfuse/discussions/11687 "Langfuse Ideas: The corrected data cannot be imported into the dataset.") [feat(cookbook): Proposal for "AI Social Engineering" Evaluation Cookbook (based on CPF Framework)](https://github.com/orgs/langfuse/discussions/11635 "Langfuse Ideas: feat(cookbook): Proposal for \"AI Social Engineering\" Evaluation Cookbook (based on CPF Framework)") [LLM-as-a-Judge: Add support for filtering traces by input/output content](https://github.com/orgs/langfuse/discussions/11592 "Langfuse Ideas: LLM-as-a-Judge: Add support for filtering traces by input/output content") [API/SDK: add bulk dataset item insert to improve versioning](https://github.com/orgs/langfuse/discussions/11567 "Langfuse Ideas: API/SDK: add bulk dataset item insert to improve versioning") [Add annotation queue dropdown to traces in session view](https://github.com/orgs/langfuse/discussions/11545 "Langfuse Ideas: Add annotation queue dropdown to traces in session view") [Track "Added By" and "Added At" for annotation queue items](https://github.com/orgs/langfuse/discussions/11542 "Langfuse Ideas: Track \"Added By\" and \"Added At\" for annotation queue items") [Add annotation and annotation queue filters to trace list view](https://github.com/orgs/langfuse/discussions/11541 "Langfuse Ideas: Add annotation and annotation queue filters to trace list view") [Attaching Guardrail policy to the LLM connection which can be attached to the evaluator in LLM-as-a-Judge.](https://github.com/orgs/langfuse/discussions/11510 "Langfuse Ideas: Attaching Guardrail policy to the LLM connection which can be attached to the evaluator in LLM-as-a-Judge.") [Display "Expected Output" within Annotation Queues for Side-by-Side Comparison](https://github.com/orgs/langfuse/discussions/11452 "Langfuse Ideas: Display \"Expected Output\" within Annotation Queues for Side-by-Side Comparison") [Allow for Bulk Deletions inside Datasets Items view](https://github.com/orgs/langfuse/discussions/11378 "Langfuse Ideas: Allow for Bulk Deletions inside Datasets Items view") [Fetching all traces along with Score data](https://github.com/orgs/langfuse/discussions/11335 "Langfuse Ideas: Fetching all traces along with Score data") [Integrate BLOOM evaluation framework for LLM safety and capability assessment](https://github.com/orgs/langfuse/discussions/11323 "Langfuse Ideas: Integrate BLOOM evaluation framework for LLM safety and capability assessment") [Feature Request: Add API endpoint to fetch all dataset items with metadata filtering support](https://github.com/orgs/langfuse/discussions/11319 "Langfuse Ideas: Feature Request: Add API endpoint to fetch all dataset items with metadata filtering support") [Evaluators Running Twice on Parent Trace and Child Observations - Creating Duplicates](https://github.com/orgs/langfuse/discussions/11318 "Langfuse Ideas: Evaluators Running Twice on Parent Trace and Child Observations - Creating Duplicates") [Save/Pre-Populate Add To Dataset Form Values](https://github.com/orgs/langfuse/discussions/11316 "Langfuse Ideas: Save/Pre-Populate Add To Dataset Form Values") [UX Suggestions in Experiments Flow](https://github.com/orgs/langfuse/discussions/11314 "Langfuse Ideas: UX Suggestions in Experiments Flow") [Edit dataset item IDs via UI or add description field](https://github.com/orgs/langfuse/discussions/11299 "Langfuse Ideas: Edit dataset item IDs via UI or add description field") [Add functionality to add a regex filter on output of LLM when creating LLM as a judge](https://github.com/orgs/langfuse/discussions/11293 "Langfuse Ideas: Add functionality to add a regex filter on output of LLM when creating LLM as a judge") [Multi-turn Conversation Evaluation](https://github.com/orgs/langfuse/discussions/11286 "Langfuse Ideas: Multi-turn Conversation Evaluation") [Make the output\_schema stored in the prompt config in the playground and in experiments](https://github.com/orgs/langfuse/discussions/11285 "Langfuse Ideas: Make the output_schema stored in the prompt config in the playground and in experiments") [Gemini 3 thinking\_level](https://github.com/orgs/langfuse/discussions/11265 "Langfuse Ideas: Gemini 3 thinking_level") [Support Aggregated Evaluation Scores at Conversation / Agent Level](https://github.com/orgs/langfuse/discussions/11219 "Langfuse Ideas: Support Aggregated Evaluation Scores at Conversation / Agent Level") [Add datasetItemsUpdate method to SDK](https://github.com/orgs/langfuse/discussions/11079 "Langfuse Ideas: Add datasetItemsUpdate method to SDK") [Support dataset splits (train/dev/test) in the Experiment Runner SDK](https://github.com/orgs/langfuse/discussions/11064 "Langfuse Ideas: Support dataset splits (train/dev/test) in the Experiment Runner SDK") [Create a custom evaluator which runs on coding log](https://github.com/orgs/langfuse/discussions/11018 "Langfuse Ideas: Create a custom evaluator which runs on coding log") [Better handling of score -1](https://github.com/orgs/langfuse/discussions/10979 "Langfuse Ideas: Better handling of score -1") [Support for Statistical Aggregation (mean, variance...) across dataset runs](https://github.com/orgs/langfuse/discussions/10925 "Langfuse Ideas: Support for Statistical Aggregation (mean, variance...) across dataset runs") [Support Dataset edit and deletion via API](https://github.com/orgs/langfuse/discussions/10919 "Langfuse Ideas: Support Dataset edit and deletion via API") [Dataset Item Creation Form](https://github.com/orgs/langfuse/discussions/10893 "Langfuse Ideas: Dataset Item Creation Form") [Batch/Multi-Trace LLM Evaluators for Pattern Discovery](https://github.com/orgs/langfuse/discussions/10888 "Langfuse Ideas: Batch/Multi-Trace LLM Evaluators for Pattern Discovery") [AI-Generated Summary of Evaluator Reasoning at Dataset Level](https://github.com/orgs/langfuse/discussions/10859 "Langfuse Ideas: AI-Generated Summary of Evaluator Reasoning at Dataset Level") [LLM-as-a-Judge on observation](https://github.com/orgs/langfuse/discussions/10822 "Langfuse Ideas: LLM-as-a-Judge on observation") [Make datasets.upsertRemoteExperiment API available in SDK](https://github.com/orgs/langfuse/discussions/10805 "Langfuse Ideas: Make datasets.upsertRemoteExperiment API available in SDK") [LLM as a Judge Sequential Chaining](https://github.com/orgs/langfuse/discussions/10793 "Langfuse Ideas: LLM as a Judge Sequential Chaining") [Resolve @@@langfuseMedia:...@@@ placeholders as base64](https://github.com/orgs/langfuse/discussions/10821 "Langfuse Ideas: Resolve @@@langfuseMedia:...@@@ placeholders as base64") [Improve the UX the comment section for LLMaJ evaluators.](https://github.com/orgs/langfuse/discussions/10710 "Langfuse Ideas: Improve the UX the comment section for LLMaJ evaluators.") [Enhanced Filtering in LLM-as-a-Judge Score Tables](https://github.com/orgs/langfuse/discussions/10656 "Langfuse Ideas: Enhanced Filtering in LLM-as-a-Judge Score Tables") [Have LLM as a judge Evaluator in folder structure ( using / )](https://github.com/orgs/langfuse/discussions/10655 "Langfuse Ideas: Have LLM as a judge Evaluator in folder structure ( using / )") [Filter & Sort by non-score columns in Compare View](https://github.com/orgs/langfuse/discussions/10475 "Langfuse Ideas: Filter & Sort by non-score columns in Compare View") [Set a prefix for the Dataset item id](https://github.com/orgs/langfuse/discussions/10448 "Langfuse Ideas: Set a prefix for the Dataset item id") [Add answer-generation as part of Dataset-/Experiment-UI](https://github.com/orgs/langfuse/discussions/10432 "Langfuse Ideas: Add answer-generation as part of Dataset-/Experiment-UI") [Apply Score Analytics to Dataset Runs](https://github.com/orgs/langfuse/discussions/10394 "Langfuse Ideas: Apply Score Analytics to Dataset Runs") [feat(Datasets): Filter Dataset runs by metadata / tags](https://github.com/orgs/langfuse/discussions/10392 "Langfuse Ideas: feat(Datasets): Filter Dataset runs by metadata / tags") [feat(annotation): support defining annotation queue item processing order by reference object ids](https://github.com/orgs/langfuse/discussions/10388 "Langfuse Ideas: feat(annotation): support defining annotation queue item processing order by reference object ids") [Allow filtering by environment in score analytics](https://github.com/orgs/langfuse/discussions/10324 "Langfuse Ideas: Allow filtering by environment in score analytics") [Add Human Annotation Queue via sdk](https://github.com/orgs/langfuse/discussions/10313 "Langfuse Ideas: Add Human Annotation Queue via sdk") [API to Fetch Custom LLM-as-a-Judge Evaluator Prompts from Evaluator Library](https://github.com/orgs/langfuse/discussions/10273 "Langfuse Ideas: API to Fetch Custom LLM-as-a-Judge Evaluator Prompts from Evaluator Library") [Edit output Friendly](https://github.com/orgs/langfuse/discussions/10230 "Langfuse Ideas: Edit output Friendly") [Allow filtering Scores by trace metadata](https://github.com/orgs/langfuse/discussions/10173 "Langfuse Ideas: Allow filtering Scores by trace metadata") [Add total results count to Scores and Traces views](https://github.com/orgs/langfuse/discussions/10132 "Langfuse Ideas: Add total results count to Scores and Traces views") [Create dataset with variables from traces](https://github.com/orgs/langfuse/discussions/10096 "Langfuse Ideas: Create dataset with variables from traces") [Support multiple metrics in a single LLM Judge output](https://github.com/orgs/langfuse/discussions/10092 "Langfuse Ideas: Support multiple metrics in a single LLM Judge output") [Automatically add traces to dataset based on score](https://github.com/orgs/langfuse/discussions/10080 "Langfuse Ideas: Automatically add traces to dataset based on score") [Datasets - excepted output can't parse new lines](https://github.com/orgs/langfuse/discussions/10002 "Langfuse Ideas: Datasets - excepted output can't parse new lines") [feat(dataset): Add a dataset prefix as the dataset item id](https://github.com/orgs/langfuse/discussions/9997 "Langfuse Ideas: feat(dataset): Add a dataset prefix as the dataset item id") [Feedback on the prompt experiments and eval settings UX](https://github.com/orgs/langfuse/discussions/9971 "Langfuse Ideas: Feedback on the prompt experiments and eval settings UX") [Add ability to add arbitrary code to run evaluations](https://github.com/orgs/langfuse/discussions/9950 "Langfuse Ideas: Add ability to add arbitrary code to run evaluations") [feat(datasets): support tags on datasets](https://github.com/orgs/langfuse/discussions/9940 "Langfuse Ideas: feat(datasets): support tags on datasets") [Run level score charts not visible](https://github.com/orgs/langfuse/discussions/9923 "Langfuse Ideas: Run level score charts not visible") [Side-by-Side Comparison in Annotation Queues](https://github.com/orgs/langfuse/discussions/9886 "Langfuse Ideas: Side-by-Side Comparison in Annotation Queues") [Feature Request: Show Diff Between Expected and Actual Output in Dataset Run](https://github.com/orgs/langfuse/discussions/9881 "Langfuse Ideas: Feature Request: Show Diff Between Expected and Actual Output in Dataset Run") [Feature Request: Custom Logic for Prompt Experiment Scoring](https://github.com/orgs/langfuse/discussions/9880 "Langfuse Ideas: Feature Request: Custom Logic for Prompt Experiment Scoring") [Feature Request: Enhanced Dataset and Dataset Item Operations](https://github.com/orgs/langfuse/discussions/9879 "Langfuse Ideas: Feature Request: Enhanced Dataset and Dataset Item Operations") [Separate datasets in folders](https://github.com/orgs/langfuse/discussions/9859 "Langfuse Ideas: Separate datasets in folders") [Dataset Runs: Show run-level metrics in sidebar](https://github.com/orgs/langfuse/discussions/9848 "Langfuse Ideas: Dataset Runs: Show run-level metrics in sidebar") [Provide renaming functionality for datasets in Langfuse public API](https://github.com/orgs/langfuse/discussions/9778 "Langfuse Ideas: Provide renaming functionality for datasets in Langfuse public API") [feat(llm-completion): support Anthropic-specific message roles in ChatML](https://github.com/orgs/langfuse/discussions/9722 "Langfuse Ideas: feat(llm-completion): support Anthropic-specific message roles in ChatML") [Create a Run in Langfuse even when an evaluation fails before run\_experiment()](https://github.com/orgs/langfuse/discussions/9713 "Langfuse Ideas: Create a Run in Langfuse even when an evaluation fails before run_experiment()") [Flow: tag and annotate traces and then find what you annotated.](https://github.com/orgs/langfuse/discussions/9698 "Langfuse Ideas: Flow: tag and annotate traces and then find what you annotated.") [Add Sessions to Datasets](https://github.com/orgs/langfuse/discussions/9678 "Langfuse Ideas: Add Sessions to Datasets") [Extract human annotation queues to csv](https://github.com/orgs/langfuse/discussions/9653 "Langfuse Ideas: Extract human annotation queues to csv") [Request feature: allow Scores to be able to filter with trace's "Release"](https://github.com/orgs/langfuse/discussions/9594 "Langfuse Ideas: Request feature: allow Scores to be able to filter with trace's \"Release\"") [feat(datasets): keep dataset items created from traces in sync with trace state](https://github.com/orgs/langfuse/discussions/9591 "Langfuse Ideas: feat(datasets): keep dataset items created from traces in sync with trace state") [feat(experiment-compare): highlighting or ability to sort by detected differences between scores would be nice for locating significant changes](https://github.com/orgs/langfuse/discussions/9574 "Langfuse Ideas: feat(experiment-compare): highlighting or ability to sort by detected differences between scores would be nice for locating significant changes") [Aggregate boolean scores as success ratio](https://github.com/orgs/langfuse/discussions/9543 "Langfuse Ideas: Aggregate boolean scores as success ratio") [See LLM Generation when using llm-as-a-judge](https://github.com/orgs/langfuse/discussions/9533 "Langfuse Ideas: See LLM Generation when using llm-as-a-judge") [Add support for unwrapping jsonpath array output / support jq](https://github.com/orgs/langfuse/discussions/9505 "Langfuse Ideas: Add support for unwrapping jsonpath array output / support jq") [Add UPDATE (PATCH) method for Score Configs in Python SDK v3 and DELETE endpoint in API](https://github.com/orgs/langfuse/discussions/9453 "Langfuse Ideas: Add UPDATE (PATCH) method for Score Configs in Python SDK v3 and DELETE endpoint in API") [Allow getting evaluator ID from score Object created by LLM-as-a-Judge evaluator](https://github.com/orgs/langfuse/discussions/9446 "Langfuse Ideas: Allow getting evaluator ID from score Object created by LLM-as-a-Judge evaluator") [feat(datasets): ability to delete evaluator from evaluator library](https://github.com/orgs/langfuse/discussions/9420 "Langfuse Ideas: feat(datasets): ability to delete evaluator from evaluator library") [提供删除整个数据集的API](https://github.com/orgs/langfuse/discussions/9394 "Langfuse Ideas: 提供删除整个数据集的API") [langfuse API lacks a deletion API for Datasets.](https://github.com/orgs/langfuse/discussions/9369 "Langfuse Ideas: langfuse API lacks a deletion API for Datasets.") [Feature Request: Export Human Annotation Comments from Human Annotation Page](https://github.com/orgs/langfuse/discussions/9366 "Langfuse Ideas: Feature Request: Export Human Annotation Comments from Human Annotation Page") [openapi: untyped DatasetItem.input](https://github.com/orgs/langfuse/discussions/9355 "Langfuse Ideas: openapi: untyped DatasetItem.input") [LLM-as-a-Judge Evaluations to Support Complex Filters & Variable Mapping](https://github.com/orgs/langfuse/discussions/9316 "Langfuse Ideas: LLM-as-a-Judge Evaluations to Support Complex Filters & Variable Mapping") [Bulk archive dataset items OR "archive all but current item" option](https://github.com/orgs/langfuse/discussions/9277 "Langfuse Ideas: Bulk archive dataset items OR \"archive all but current item\" option") [\[Feature Request\] Dataset with more flexible views](https://github.com/orgs/langfuse/discussions/9266 "Langfuse Ideas: [Feature Request] Dataset with more flexible views") [Given a sessionId, if we can FETCH the scores ( via api/sdk) associated with that session ( not its traces)](https://github.com/orgs/langfuse/discussions/9166 "Langfuse Ideas: Given a sessionId, if we can FETCH the scores ( via api/sdk) associated with that session ( not its traces)") [Allow triggering Prompt Experiments programmatically (via public API)](https://github.com/orgs/langfuse/discussions/9131 "Langfuse Ideas: Allow triggering Prompt Experiments programmatically (via public API)") [Search by trace, order by user in the annotation queue and add labels](https://github.com/orgs/langfuse/discussions/9123 "Langfuse Ideas: Search by trace, order by user in the annotation queue and add labels") [Add UI support for LLM-as-a-Judge evaluation on CSV datasets containing pre-existing input/output pairs without requiring re-generation](https://github.com/orgs/langfuse/discussions/9080 "Langfuse Ideas: Add UI support for LLM-as-a-Judge evaluation on CSV datasets containing pre-existing input/output pairs without requiring re-generation") [Allow configuring iteration count per dataset item for dataset runs via SDK](https://github.com/orgs/langfuse/discussions/9056 "Langfuse Ideas: Allow configuring iteration count per dataset item for dataset runs via SDK") [Feature request : Enable adding API key for Remote Dataset Runs](https://github.com/orgs/langfuse/discussions/8992 "Langfuse Ideas: Feature request : Enable adding API key for Remote Dataset Runs") [Dataset Transformations](https://github.com/orgs/langfuse/discussions/8812 "Langfuse Ideas: Dataset Transformations") [Human-readable LangGraph traces in Langfuse (Formatted Mode)](https://github.com/orgs/langfuse/discussions/8653 "Langfuse Ideas: Human-readable LangGraph traces in Langfuse (Formatted Mode)") [Support for YouTube Video URL Inputs in Datasets](https://github.com/orgs/langfuse/discussions/8624 "Langfuse Ideas: Support for YouTube Video URL Inputs in Datasets") [Python SDK: Pass ScoreConfig instead of config\_id when creating scores](https://github.com/orgs/langfuse/discussions/8623 "Langfuse Ideas: Python SDK: Pass ScoreConfig instead of config_id when creating scores") [Filter Support for Dataset Runs](https://github.com/orgs/langfuse/discussions/8622 "Langfuse Ideas: Filter Support for Dataset Runs") [Request batch adding of multiple traces to datasets](https://github.com/orgs/langfuse/discussions/8526 "Langfuse Ideas: Request batch adding of multiple traces to datasets") [How to categorize topics and create a pie chart?](https://github.com/orgs/langfuse/discussions/8512 "Langfuse Ideas: How to categorize topics and create a pie chart?") [Simultaneous Session and Trace Annotation for Context-Rich Feedback in the same annotation window](https://github.com/orgs/langfuse/discussions/8485 "Langfuse Ideas: Simultaneous Session and Trace Annotation for Context-Rich Feedback in the same annotation window") [feat(LLM-as-a-judge): support stratified sampling by trace property](https://github.com/orgs/langfuse/discussions/8480 "Langfuse Ideas: feat(LLM-as-a-judge): support stratified sampling by trace property") [Bulk Evaluation Does not Displays the status of the dataset (Like for E.g. processing, evaluated, executed)](https://github.com/orgs/langfuse/discussions/8410 "Langfuse Ideas: Bulk Evaluation Does not Displays the status of the dataset (Like for E.g. processing, evaluated, executed)") [Annotation Queue Creation API](https://github.com/orgs/langfuse/discussions/8372 "Langfuse Ideas: Annotation Queue Creation API") [\[Langfuse Cloud\] Missing SessionId and Author in exported Scores](https://github.com/orgs/langfuse/discussions/8346 "Langfuse Ideas: [Langfuse Cloud] Missing SessionId and Author in exported Scores") [feat(evals): allow canceling a running evaluator with pending evaluation jobs](https://github.com/orgs/langfuse/discussions/8310 "Langfuse Ideas: feat(evals): allow canceling a running evaluator with pending evaluation jobs") [Support walking through inner spans in langfuse SDK for e2e trace evaluation](https://github.com/orgs/langfuse/discussions/8260 "Langfuse Ideas: Support walking through inner spans in langfuse SDK for e2e trace evaluation") [Score Configs: Allow editing the categories of a categorical score](https://github.com/orgs/langfuse/discussions/8259 "Langfuse Ideas: Score Configs: Allow editing the categories of a categorical score") [Trials for Dataset Runs: Multiple runs on the same item and input to get robust scores](https://github.com/orgs/langfuse/discussions/8258 "Langfuse Ideas: Trials for Dataset Runs: Multiple runs on the same item and input to get robust scores") [Configuring evaluators (LLM-as-a-Judge) via API or SDK (CRUD)](https://github.com/orgs/langfuse/discussions/8241 "Langfuse Ideas: Configuring evaluators (LLM-as-a-Judge) via API or SDK (CRUD)") [UI-LLM as a Jury](https://github.com/orgs/langfuse/discussions/8195 "Langfuse Ideas: UI-LLM as a Jury") [Ability to download the evaluation logs (LLM as a Judge Logs)](https://github.com/orgs/langfuse/discussions/8183 "Langfuse Ideas: Ability to download the evaluation logs (LLM as a Judge Logs)") [Knowledge base (source material) for the LLM-as-a-judge feature](https://github.com/orgs/langfuse/discussions/8172 "Langfuse Ideas: Knowledge base (source material) for the LLM-as-a-judge feature") [Option to run Experiments without traces](https://github.com/orgs/langfuse/discussions/8133 "Langfuse Ideas: Option to run Experiments without traces") [Auto-generate dataset items](https://github.com/orgs/langfuse/discussions/8126 "Langfuse Ideas: Auto-generate dataset items") [feature: support for creation of custom) model adapters](https://github.com/orgs/langfuse/discussions/8123 "Langfuse Ideas: feature: support for creation of custom) model adapters") [Prompt experiment result download.](https://github.com/orgs/langfuse/discussions/8120 "Langfuse Ideas: Prompt experiment result download.") [Multiple predictions per-item in a single run](https://github.com/orgs/langfuse/discussions/8040 "Langfuse Ideas: Multiple predictions per-item in a single run") [Add evaluator(s) to python SDK (FernLangfuse)](https://github.com/orgs/langfuse/discussions/8018 "Langfuse Ideas: Add evaluator(s) to python SDK (FernLangfuse)") [feat(dataset-runs): return all dataset run item scores in a given run](https://github.com/orgs/langfuse/discussions/8011 "Langfuse Ideas: feat(dataset-runs): return all dataset run item scores in a given run") [Delete an evaluator from the evaluator library](https://github.com/orgs/langfuse/discussions/7960 "Langfuse Ideas: Delete an evaluator from the evaluator library") [Archive or Delete Evaluators from the Evaluator Library](https://github.com/orgs/langfuse/discussions/7868 "Langfuse Ideas: Archive or Delete Evaluators from the Evaluator Library") [Edit dataset run name, description](https://github.com/orgs/langfuse/discussions/7814 "Langfuse Ideas: Edit dataset run name, description") [Dataset UI improvements for easier item management](https://github.com/orgs/langfuse/discussions/7733 "Langfuse Ideas: Dataset UI improvements for easier item management") [Batch-add traces to datasets](https://github.com/orgs/langfuse/discussions/7691 "Langfuse Ideas: Batch-add traces to datasets") [Enable Immediate Score Management for User Feedback](https://github.com/orgs/langfuse/discussions/7686 "Langfuse Ideas: Enable Immediate Score Management for User Feedback") [Export Dataset runs and run items](https://github.com/orgs/langfuse/discussions/7601 "Langfuse Ideas: Export Dataset runs and run items") [feat(experiments/evaluation): support user defined exponential backoff to avoid hitting LLM tier limits](https://github.com/orgs/langfuse/discussions/7599 "Langfuse Ideas: feat(experiments/evaluation): support user defined exponential backoff to avoid hitting LLM tier limits") [support trajectory LLM-as-judge](https://github.com/orgs/langfuse/discussions/7586 "Langfuse Ideas: support trajectory LLM-as-judge") [feat: allow redirect from evaluator logs to filter traces for evaluator score values](https://github.com/orgs/langfuse/discussions/7559 "Langfuse Ideas: feat: allow redirect from evaluator logs to filter traces for evaluator score values") [Annotation Queue for Sessions](https://github.com/orgs/langfuse/discussions/7551 "Langfuse Ideas: Annotation Queue for Sessions") [Filter by scores in session view](https://github.com/orgs/langfuse/discussions/7528 "Langfuse Ideas: Filter by scores in session view") [Support break lines on evaluation run tooltip hint](https://github.com/orgs/langfuse/discussions/7452 "Langfuse Ideas: Support break lines on evaluation run tooltip hint") ["Create new evaluator" support Qwen model](https://github.com/orgs/langfuse/discussions/7450 "Langfuse Ideas: \"Create new evaluator\" support Qwen model") [Exporting and sorting evaluators](https://github.com/orgs/langfuse/discussions/6699 "Langfuse Ideas: Exporting and sorting evaluators") [Optional schema definition for datasets](https://github.com/orgs/langfuse/discussions/6670 "Langfuse Ideas: Optional schema definition for datasets") [feat(datasets): allow deleting via API](https://github.com/orgs/langfuse/discussions/6610 "Langfuse Ideas: feat(datasets): allow deleting via API") [Versioning for datasets](https://github.com/orgs/langfuse/discussions/6596 "Langfuse Ideas: Versioning for datasets") [Support new lines when storing / displaying score comments](https://github.com/orgs/langfuse/discussions/6473 "Langfuse Ideas: Support new lines when storing / displaying score comments") [feat(annotation): allow for conditional score configurations](https://github.com/orgs/langfuse/discussions/6362 "Langfuse Ideas: feat(annotation): allow for conditional score configurations") [Delete multiple dataset items](https://github.com/orgs/langfuse/discussions/6332 "Langfuse Ideas: Delete multiple dataset items") [Multi-label scores during human annotation](https://github.com/orgs/langfuse/discussions/6310 "Langfuse Ideas: Multi-label scores during human annotation") [Evaluator: Filter for Scores](https://github.com/orgs/langfuse/discussions/6236 "Langfuse Ideas: Evaluator: Filter for Scores") [Re-run LLM-as-a-judge evaluations](https://github.com/orgs/langfuse/discussions/6098 "Langfuse Ideas: Re-run LLM-as-a-judge evaluations") [response\_format for experiment runs](https://github.com/orgs/langfuse/discussions/6089 "Langfuse Ideas: response_format for experiment runs") [Code based custom evaluators](https://github.com/orgs/langfuse/discussions/6087 "Langfuse Ideas: Code based custom evaluators") [Support Description on Dataset Items](https://github.com/orgs/langfuse/discussions/6011 "Langfuse Ideas: Support Description on Dataset Items") [Run prompt experiment with structured output / tool instruction](https://github.com/orgs/langfuse/discussions/5958 "Langfuse Ideas: Run prompt experiment with structured output / tool instruction") [Alta Integration](https://github.com/orgs/langfuse/discussions/5957 "Langfuse Ideas: Alta Integration") [Add run name to columns when looking at a specific dataset item](https://github.com/orgs/langfuse/discussions/5929 "Langfuse Ideas: Add run name to columns when looking at a specific dataset item") [Delete multiple dataset runs](https://github.com/orgs/langfuse/discussions/5893 "Langfuse Ideas: Delete multiple dataset runs") [Navigation between items in a dataset run is confusing - context of the selected dataset run is lost](https://github.com/orgs/langfuse/discussions/5892 "Langfuse Ideas: Navigation between items in a dataset run is confusing - context of the selected dataset run is lost") [Enhanced score distribution visualization in experiment analysis](https://github.com/orgs/langfuse/discussions/5819 "Langfuse Ideas: Enhanced score distribution visualization in experiment analysis") [Multi-step Prompt Experiments and Playground](https://github.com/orgs/langfuse/discussions/5812 "Langfuse Ideas: Multi-step Prompt Experiments and Playground") [Option to add trace to new dataset](https://github.com/orgs/langfuse/discussions/5756 "Langfuse Ideas: Option to add trace to new dataset") [Simplified UI for Scoring](https://github.com/orgs/langfuse/discussions/5721 "Langfuse Ideas: Simplified UI for Scoring") [More scoring configurations on UI](https://github.com/orgs/langfuse/discussions/5719 "Langfuse Ideas: More scoring configurations on UI") [Expose dataset item status (Archived or Active) on the item detail page](https://github.com/orgs/langfuse/discussions/5689 "Langfuse Ideas: Expose dataset item status (Archived or Active) on the item detail page") [Expose CRUD API for dataset run items](https://github.com/orgs/langfuse/discussions/5670 "Langfuse Ideas: Expose CRUD API for dataset run items") [feat \[UI\] - remember selected charts for Datasets](https://github.com/orgs/langfuse/discussions/5641 "Langfuse Ideas: feat [UI] - remember selected charts for Datasets") [Edit Name of Eval Templates](https://github.com/orgs/langfuse/discussions/5623 "Langfuse Ideas: Edit Name of Eval Templates") [CRUD Evaluation Templates via the API](https://github.com/orgs/langfuse/discussions/5612 "Langfuse Ideas: CRUD Evaluation Templates via the API") [Evals: run an evaluator on a dataset without also running a prompt experiment](https://github.com/orgs/langfuse/discussions/5585 "Langfuse Ideas: Evals: run an evaluator on a dataset without also running a prompt experiment") [chart request - mirror score functionality but after filtering within trace or observation panes](https://github.com/orgs/langfuse/discussions/5426 "Langfuse Ideas: chart request - mirror score functionality but after filtering within trace or observation panes") [UI/UX: Scores Analytics - reduce interaction friction](https://github.com/orgs/langfuse/discussions/5365 "Langfuse Ideas: UI/UX: Scores Analytics - reduce interaction friction") [LLM as a judge: only execute judge if a specific observation exists](https://github.com/orgs/langfuse/discussions/5344 "Langfuse Ideas: LLM as a judge: only execute judge if a specific observation exists") [Update scores via the UI even if they have been created via the API](https://github.com/orgs/langfuse/discussions/5261 "Langfuse Ideas: Update scores via the UI even if they have been created via the API") [Trajectory Evaluation - Evaluating path chooses by Agent](https://github.com/orgs/langfuse/discussions/5206 "Langfuse Ideas: Trajectory Evaluation - Evaluating path chooses by Agent") [LLM-as-a-Judge: Categorical and Boolean Scores](https://github.com/orgs/langfuse/discussions/4965 "Langfuse Ideas: LLM-as-a-Judge: Categorical and Boolean Scores") [\[Dataset Run\] Add filtering functionality for some columns](https://github.com/orgs/langfuse/discussions/4903 "Langfuse Ideas: [Dataset Run] Add filtering functionality for some columns") [Ability to run evaluations using prompt from prompt editor with content from spans](https://github.com/orgs/langfuse/discussions/4851 "Langfuse Ideas: Ability to run evaluations using prompt from prompt editor with content from spans") ["Add to Dataset" : augment ground truth labelling with custom AI-assistant](https://github.com/orgs/langfuse/discussions/4850 "Langfuse Ideas: \"Add to Dataset\" : augment ground truth labelling with custom AI-assistant") [Allow using User-defined models in Playground & Evals](https://github.com/orgs/langfuse/discussions/4789 "Langfuse Ideas: Allow using User-defined models in Playground & Evals") [Ability to edit the LLM API keys](https://github.com/orgs/langfuse/discussions/4726 "Langfuse Ideas: Ability to edit the LLM API keys") [API for annotation queues to programmatically manage them (enqueue, dequeue, status)](https://github.com/orgs/langfuse/discussions/4714 "Langfuse Ideas: API for annotation queues to programmatically manage them (enqueue, dequeue, status)") [Add traces to datasets from the trace table](https://github.com/orgs/langfuse/discussions/4682 "Langfuse Ideas: Add traces to datasets from the trace table") [feat(experiments/evals): Regex based evaluators](https://github.com/orgs/langfuse/discussions/4671 "Langfuse Ideas: feat(experiments/evals): Regex based evaluators") [Allow Gemini access via Google AI (not Vertex AI)](https://github.com/orgs/langfuse/discussions/4669 "Langfuse Ideas: Allow Gemini access via Google AI (not Vertex AI)") [ui: ability to highlight scores](https://github.com/orgs/langfuse/discussions/4617 "Langfuse Ideas: ui: ability to highlight scores") [Run prompt experiments via the SDKs/API](https://github.com/orgs/langfuse/discussions/4587 "Langfuse Ideas: Run prompt experiments via the SDKs/API") [Custom Non-LLM evaluators/scores through UI](https://github.com/orgs/langfuse/discussions/4484 "Langfuse Ideas: Custom Non-LLM evaluators/scores through UI") [Side-by-side comparison in playground](https://github.com/orgs/langfuse/discussions/4483 "Langfuse Ideas: Side-by-side comparison in playground") [Support user input, message history, structured output, and tool calls in prompt experiments](https://github.com/orgs/langfuse/discussions/4454 "Langfuse Ideas: Support user input, message history, structured output, and tool calls in prompt experiments") [Datasets: Add selection of traces to a dataset](https://github.com/orgs/langfuse/discussions/4353 "Langfuse Ideas: Datasets: Add selection of traces to a dataset") [Multi-user annotation capability in Annotation Queues](https://github.com/orgs/langfuse/discussions/4348 "Langfuse Ideas: Multi-user annotation capability in Annotation Queues") [Multi-turn / session experiments in datasets](https://github.com/orgs/langfuse/discussions/4208 "Langfuse Ideas: Multi-turn / session experiments in datasets") [Enable to use variable of prompt on evaluator.](https://github.com/orgs/langfuse/discussions/4121 "Langfuse Ideas: Enable to use variable of prompt on evaluator.") [Sessions Table: Scores Column](https://github.com/orgs/langfuse/discussions/4120 "Langfuse Ideas: Sessions Table: Scores Column") [Add new filters for the LLM as a Judge Evaluation (other scores and cost)](https://github.com/orgs/langfuse/discussions/4106 "Langfuse Ideas: Add new filters for the LLM as a Judge Evaluation (other scores and cost)") [Export dataset run table](https://github.com/orgs/langfuse/discussions/4077 "Langfuse Ideas: Export dataset run table") [feat: support adding trace tags in annotation queue view](https://github.com/orgs/langfuse/discussions/4037 "Langfuse Ideas: feat: support adding trace tags in annotation queue view") [Feat: De-dupe and show info when adding the same trace/observation to an annotation queue multiple times](https://github.com/orgs/langfuse/discussions/4035 "Langfuse Ideas: Feat: De-dupe and show info when adding the same trace/observation to an annotation queue multiple times") [Diff support for dataset runs view](https://github.com/orgs/langfuse/discussions/4025 "Langfuse Ideas: Diff support for dataset runs view") [Create Support for gemini models in playground](https://github.com/orgs/langfuse/discussions/4019 "Langfuse Ideas: Create Support for gemini models in playground") [Change AWS access pattern for Bedrock LLM usage, assume role](https://github.com/orgs/langfuse/discussions/3988 "Langfuse Ideas: Change AWS access pattern for Bedrock LLM usage, assume role") [Add ability to export and import evaluators between projects](https://github.com/orgs/langfuse/discussions/3970 "Langfuse Ideas: Add ability to export and import evaluators between projects") [feat: Folder structure for dataset organisation](https://github.com/orgs/langfuse/discussions/3935 "Langfuse Ideas: feat: Folder structure for dataset organisation") [Model-based evaluations triggered by observations](https://github.com/orgs/langfuse/discussions/3918 "Langfuse Ideas: Model-based evaluations triggered by observations") [Scores: Conditional Annotation](https://github.com/orgs/langfuse/discussions/3842 "Langfuse Ideas: Scores: Conditional Annotation") [Annotation Queues: define optional/mandatory score configs by queue](https://github.com/orgs/langfuse/discussions/3841 "Langfuse Ideas: Annotation Queues: define optional/mandatory score configs by queue") [Scores: support for recording multiple choice selection as score value](https://github.com/orgs/langfuse/discussions/3840 "Langfuse Ideas: Scores: support for recording multiple choice selection as score value") [Filter by status in dataset items table](https://github.com/orgs/langfuse/discussions/3818 "Langfuse Ideas: Filter by status in dataset items table") [Versioning/edit of Evaluator Configuration](https://github.com/orgs/langfuse/discussions/3798 "Langfuse Ideas: Versioning/edit of Evaluator Configuration") [Support annotation role permission to only access Annotation Queue feature](https://github.com/orgs/langfuse/discussions/3772 "Langfuse Ideas: Support annotation role permission to only access Annotation Queue feature") [Vector search for datasets for few-shot prompting](https://github.com/orgs/langfuse/discussions/3538 "Langfuse Ideas: Vector search for datasets for few-shot prompting") [Filtering dataset items](https://github.com/orgs/langfuse/discussions/3439 "Langfuse Ideas: Filtering dataset items") [Export Datasets via CSV](https://github.com/orgs/langfuse/discussions/3438 "Langfuse Ideas: Export Datasets via CSV") [Only use traces in the dataset to render score columns](https://github.com/orgs/langfuse/discussions/3415 "Langfuse Ideas: Only use traces in the dataset to render score columns") [Source code versioning and automatic statistics caluclation from scores](https://github.com/orgs/langfuse/discussions/3412 "Langfuse Ideas: Source code versioning and automatic statistics caluclation from scores") [Include metrics and scores in getDatasetRun / \`GET /api/public/datasets/{datasetName}/runs/{runName}\`](https://github.com/orgs/langfuse/discussions/3275 "Langfuse Ideas: Include metrics and scores in getDatasetRun / `GET /api/public/datasets/{datasetName}/runs/{runName}`") [API/SDK to create comments on traces](https://github.com/orgs/langfuse/discussions/3263 "Langfuse Ideas: API/SDK to create comments on traces") [Need to retain the old evaluation history results, including input and all ground truth](https://github.com/orgs/langfuse/discussions/3252 "Langfuse Ideas: Need to retain the old evaluation history results, including input and all ground truth") [Support for Selecting Specific Paths from JSON Objects in Evaluation Prompts](https://github.com/orgs/langfuse/discussions/3237 "Langfuse Ideas: Support for Selecting Specific Paths from JSON Objects in Evaluation Prompts") [Dataset run description and metadata should not be fixed on the page](https://github.com/orgs/langfuse/discussions/3212 "Langfuse Ideas: Dataset run description and metadata should not be fixed on the page") [Optionally set timestamp when creating a score](https://github.com/orgs/langfuse/discussions/3194 "Langfuse Ideas: Optionally set timestamp when creating a score") [Implement dataset removal method](https://github.com/orgs/langfuse/discussions/3050 "Langfuse Ideas: Implement dataset removal method") [Compare two dataset runs side by side](https://github.com/orgs/langfuse/discussions/2883 "Langfuse Ideas: Compare two dataset runs side by side") [Add Bedrock Guardrails to the LLM Security documentation](https://github.com/orgs/langfuse/discussions/2864 "Langfuse Ideas: Add Bedrock Guardrails to the LLM Security documentation") [Set variables in playground from dataset](https://github.com/orgs/langfuse/discussions/2738 "Langfuse Ideas: Set variables in playground from dataset") [Session-level scores](https://github.com/orgs/langfuse/discussions/2728 "Langfuse Ideas: Session-level scores") [Feature Request: Sidebar Display for Trace Details in Dataset Runs](https://github.com/orgs/langfuse/discussions/2523 "Langfuse Ideas: Feature Request: Sidebar Display for Trace Details in Dataset Runs") [Scoring dataset runs, e.g. precision, recall, f-value](https://github.com/orgs/langfuse/discussions/2511 "Langfuse Ideas: Scoring dataset runs, e.g. precision, recall, f-value") [Adding userId / author to score (custom metadata)](https://github.com/orgs/langfuse/discussions/2469 "Langfuse Ideas: Adding userId / author to score (custom metadata)") [Expand all json-views of Dataset items etc.](https://github.com/orgs/langfuse/discussions/2454 "Langfuse Ideas: Expand all json-views of Dataset items etc.") [Add string data type in score config](https://github.com/orgs/langfuse/discussions/2402 "Langfuse Ideas: Add string data type in score config") [SBS Markdown mode for dataset runs](https://github.com/orgs/langfuse/discussions/2371 "Langfuse Ideas: SBS Markdown mode for dataset runs") [Proposal: Add Support for Uploading Dataset Items via UI](https://github.com/orgs/langfuse/discussions/2124 "Langfuse Ideas: Proposal: Add Support for Uploading Dataset Items via UI") [Save playground conversation to a dataset](https://github.com/orgs/langfuse/discussions/2071 "Langfuse Ideas: Save playground conversation to a dataset") [Upload datasets via UI](https://github.com/orgs/langfuse/discussions/1988 "Langfuse Ideas: Upload datasets via UI") [API/UI to delete dataset items and runs](https://github.com/orgs/langfuse/discussions/1987 "Langfuse Ideas: API/UI to delete dataset items and runs") [Support Linking Execution Trace to DatasetItem without Fetching Entire Dataset](https://github.com/orgs/langfuse/discussions/1862 "Langfuse Ideas: Support Linking Execution Trace to DatasetItem without Fetching Entire Dataset") [annotation of traces in langfuse console?](https://github.com/orgs/langfuse/discussions/1760 "Langfuse Ideas: annotation of traces in langfuse console?") [What standardized dataset formats are people using?](https://github.com/orgs/langfuse/discussions/1398 "Langfuse Ideas: What standardized dataset formats are people using?") [API to delete scores](https://github.com/orgs/langfuse/discussions/1133 "Langfuse Ideas: API to delete scores") [Datasets: Diff of output and expected output](https://github.com/orgs/langfuse/discussions/1008 "Langfuse Ideas: Datasets: Diff of output and expected output") [Run level confusion matrix via UI](https://github.com/orgs/langfuse/discussions/12002 "Langfuse Support: Run level confusion matrix via UI") [Skipping an item in LLM-as-a-judge eval](https://github.com/orgs/langfuse/discussions/11994 "Langfuse Support: Skipping an item in LLM-as-a-judge eval") [How to debug when evaluations are stalling](https://github.com/orgs/langfuse/discussions/11912 "Langfuse Support: How to debug when evaluations are stalling") [Cannot access 'corrected\_output' from a trace via the SDK or API](https://github.com/orgs/langfuse/discussions/11870 "Langfuse Support: Cannot access 'corrected_output' from a trace via the SDK or API") [Experiments via UI for Agent Server?](https://github.com/orgs/langfuse/discussions/11855 "Langfuse Support: Experiments via UI for Agent Server?") [How to filter multiple userids in the evaluator?](https://github.com/orgs/langfuse/discussions/11835 "Langfuse Support: How to filter multiple userids in the evaluator?") [Inquiry regarding multiple Human Annotations per single trace](https://github.com/orgs/langfuse/discussions/11822 "Langfuse Support: Inquiry regarding multiple Human Annotations per single trace") [Inquiry regarding the execution logic and concurrency control of UI-configured Evaluators](https://github.com/orgs/langfuse/discussions/11785 "Langfuse Support: Inquiry regarding the execution logic and concurrency control of UI-configured Evaluators") [Fetch Evaluator Library (Eval Templates) via Public API/SDK](https://github.com/orgs/langfuse/discussions/11754 "Langfuse Support: Fetch Evaluator Library (Eval Templates) via Public API/SDK") [How can I update datasetItems via the API? (using typescript SDK)](https://github.com/orgs/langfuse/discussions/11751 "Langfuse Support: How can I update datasetItems via the API? (using typescript SDK)") [Same Evaluator for Live/Dataset Runs but different vars](https://github.com/orgs/langfuse/discussions/11729 "Langfuse Support: Same Evaluator for Live/Dataset Runs but different vars") [How to emulate same trace format in experiment run in sdk as prompt experiment in ui.](https://github.com/orgs/langfuse/discussions/11700 "Langfuse Support: How to emulate same trace format in experiment run in sdk as prompt experiment in ui.") [Sequence of the query level evaluators](https://github.com/orgs/langfuse/discussions/11699 "Langfuse Support: Sequence of the query level evaluators") [Datasets if I run single or batch some questions](https://github.com/orgs/langfuse/discussions/11697 "Langfuse Support: Datasets if I run single or batch some questions") [Will LLM-as-a-Judge Process Uploaded Files in a Multi-LLM Chat Application?](https://github.com/orgs/langfuse/discussions/11675 "Langfuse Support: Will LLM-as-a-Judge Process Uploaded Files in a Multi-LLM Chat Application?") [How to filter LLM-as-a-Judge evaluator to run only on traces with specific input content?](https://github.com/orgs/langfuse/discussions/11590 "Langfuse Support: How to filter LLM-as-a-Judge evaluator to run only on traces with specific input content?") [Dataset run items showing identical scores despite unique observation-level scores](https://github.com/orgs/langfuse/discussions/11513 "Langfuse Support: Dataset run items showing identical scores despite unique observation-level scores") [Dataset Experiment Run - Difficulty Running Judge Prompt for Multiple Questions in a Single Call](https://github.com/orgs/langfuse/discussions/11499 "Langfuse Support: Dataset Experiment Run - Difficulty Running Judge Prompt for Multiple Questions in a Single Call") [Score analytics - beta feature review](https://github.com/orgs/langfuse/discussions/11444 "Langfuse Support: Score analytics - beta feature review") [Langfuse Eval workflow](https://github.com/orgs/langfuse/discussions/11432 "Langfuse Support: Langfuse Eval workflow") [Bulk Add Observations to Human Evaluations with Fine-Grained Scoring](https://github.com/orgs/langfuse/discussions/11419 "Langfuse Support: Bulk Add Observations to Human Evaluations with Fine-Grained Scoring") [Aggregated metrics for experiments](https://github.com/orgs/langfuse/discussions/11404 "Langfuse Support: Aggregated metrics for experiments") [Metrics & Scores](https://github.com/orgs/langfuse/discussions/11372 "Langfuse Support: Metrics & Scores") [Datasets evaluattion twice](https://github.com/orgs/langfuse/discussions/11324 "Langfuse Support: Datasets evaluattion twice") [Custom tools for LLM-as-a-Judge Evaluation](https://github.com/orgs/langfuse/discussions/11309 "Langfuse Support: Custom tools for LLM-as-a-Judge Evaluation") [Evaluation Fails When Extracted Generation Parameters Are Not Invoked in Trace](https://github.com/orgs/langfuse/discussions/11306 "Langfuse Support: Evaluation Fails When Extracted Generation Parameters Are Not Invoked in Trace") [Listening to eval scores](https://github.com/orgs/langfuse/discussions/11217 "Langfuse Support: Listening to eval scores") [feat(experiments): Support splitting task JSON response into output and metadata fields in dataset runs](https://github.com/orgs/langfuse/discussions/11214 "Langfuse Support: feat(experiments): Support splitting task JSON response into output and metadata fields in dataset runs") [My evaluator is not running](https://github.com/orgs/langfuse/discussions/11213 "Langfuse Support: My evaluator is not running") [How to extract dynamic "last user message" from chat history for Evaluation?](https://github.com/orgs/langfuse/discussions/11023 "Langfuse Support: How to extract dynamic \"last user message\" from chat history for Evaluation?") [datasets evaluator filtering options？](https://github.com/orgs/langfuse/discussions/11017 "Langfuse Support: datasets evaluator filtering options？") [Best practice for annotating offline human–human conversations in Langfuse](https://github.com/orgs/langfuse/discussions/10950 "Langfuse Support: Best practice for annotating offline human–human conversations in Langfuse") [Dataset metadata not available during generation in Experiments (Self-Hosted)](https://github.com/orgs/langfuse/discussions/10837 "Langfuse Support: Dataset metadata not available during generation in Experiments (Self-Hosted)") [Retrieving scores updated in the UI via the API](https://github.com/orgs/langfuse/discussions/10593 "Langfuse Support: Retrieving scores updated in the UI via the API") [Pass the complete table data for aggregate queries](https://github.com/orgs/langfuse/discussions/10557 "Langfuse Support: Pass the complete table data for aggregate queries") [删除整个数据集](https://github.com/orgs/langfuse/discussions/10534 "Langfuse Support: 删除整个数据集") [Using Langfuse Prompt Management with Bedrock Invoke Model instead of Converse API](https://github.com/orgs/langfuse/discussions/10518 "Langfuse Support: Using Langfuse Prompt Management with Bedrock Invoke Model instead of Converse API") [Add score configs in an specific order in the annotation queue](https://github.com/orgs/langfuse/discussions/10486 "Langfuse Support: Add score configs in an specific order in the annotation queue") [API for existing evaluators](https://github.com/orgs/langfuse/discussions/10453 "Langfuse Support: API for existing evaluators") [How to Evaluate an External API](https://github.com/orgs/langfuse/discussions/10341 "Langfuse Support: How to Evaluate an External API") [How to evaluate nested spans in dataset.run\_experiment() evaluators?](https://github.com/orgs/langfuse/discussions/10307 "Langfuse Support: How to evaluate nested spans in dataset.run_experiment() evaluators?") [Public API does not allow filter by scores](https://github.com/orgs/langfuse/discussions/10265 "Langfuse Support: Public API does not allow filter by scores") [Allow defining user ID when creating score via SDK](https://github.com/orgs/langfuse/discussions/10187 "Langfuse Support: Allow defining user ID when creating score via SDK") [Using both SDK Experiments + Configured LLM-as-a-Judge Together](https://github.com/orgs/langfuse/discussions/10133 "Langfuse Support: Using both SDK Experiments + Configured LLM-as-a-Judge Together") [Programatically assigning scores to each item in annotation queue using langfuse sdk](https://github.com/orgs/langfuse/discussions/10124 "Langfuse Support: Programatically assigning scores to each item in annotation queue using langfuse sdk") [Issues when running an evaluator on existing traces](https://github.com/orgs/langfuse/discussions/10093 "Langfuse Support: Issues when running an evaluator on existing traces") [ordering dataset items within experiment results by a number that is in the llm response for each item?](https://github.com/orgs/langfuse/discussions/10087 "Langfuse Support: ordering dataset items within experiment results by a number that is in the llm response for each item?") [How do you link Scores to a Prompt](https://github.com/orgs/langfuse/discussions/10042 "Langfuse Support: How do you link Scores to a Prompt") [Langfuse cloud traces for dataset experiments](https://github.com/orgs/langfuse/discussions/10030 "Langfuse Support: Langfuse cloud traces for dataset experiments") [LLM-as-a-Judge format JSON](https://github.com/orgs/langfuse/discussions/9993 "Langfuse Support: LLM-as-a-Judge format JSON") [DataSet Items api doesn't work](https://github.com/orgs/langfuse/discussions/9991 "Langfuse Support: DataSet Items api doesn't work") [For depeval evaluation , do we need a real time expected output for evaluating some metrics like g-eval, answer relavancy , hallucination etc.](https://github.com/orgs/langfuse/discussions/9973 "Langfuse Support: For depeval evaluation , do we need a real time expected output for evaluating some metrics like g-eval, answer relavancy , hallucination etc.") [How are evaluator prompts combined and sent to the LLM in LLM-as-a-judge?](https://github.com/orgs/langfuse/discussions/9935 "Langfuse Support: How are evaluator prompts combined and sent to the LLM in LLM-as-a-judge?") [\[question\] Custom Model Name + No Actual run + Annotation](https://github.com/orgs/langfuse/discussions/9862 "Langfuse Support: [question] Custom Model Name + No Actual run + Annotation") [Is it possible to get judge evaluation result for dataset run item with SDK/REST api ?](https://github.com/orgs/langfuse/discussions/9834 "Langfuse Support: Is it possible to get judge evaluation result for dataset run item with SDK/REST api ?") [How to add expected output for Geval?](https://github.com/orgs/langfuse/discussions/9755 "Langfuse Support: How to add expected output for Geval?") [“Body exceeded 1mb limit” error on dataset upload](https://github.com/orgs/langfuse/discussions/9751 "Langfuse Support: “Body exceeded 1mb limit” error on dataset upload") [Explicit Feedback - Can a User Researcher add this on behalf of a user?](https://github.com/orgs/langfuse/discussions/9703 "Langfuse Support: Explicit Feedback - Can a User Researcher add this on behalf of a user?") [Multi-step metrics from Ragas](https://github.com/orgs/langfuse/discussions/9687 "Langfuse Support: Multi-step metrics from Ragas") [How to evaluate correct tool calls are being made?](https://github.com/orgs/langfuse/discussions/9554 "Langfuse Support: How to evaluate correct tool calls are being made?") [How to evaluate existing traces?](https://github.com/orgs/langfuse/discussions/9552 "Langfuse Support: How to evaluate existing traces?") [IMAGE-PROCESSING-USING-GEMINI](https://github.com/orgs/langfuse/discussions/9522 "Langfuse Support: IMAGE-PROCESSING-USING-GEMINI") [Ability to View Score Comments in LangFuse UI](https://github.com/orgs/langfuse/discussions/9504 "Langfuse Support: Ability to View Score Comments in LangFuse UI") [Evaluation, between two outputs.](https://github.com/orgs/langfuse/discussions/9499 "Langfuse Support: Evaluation, between two outputs.") [Dataset run fails with LLM-as-judge against a chat prompt with placeholders](https://github.com/orgs/langfuse/discussions/9409 "Langfuse Support: Dataset run fails with LLM-as-judge against a chat prompt with placeholders") [删除整个数据集](https://github.com/orgs/langfuse/discussions/9395 "Langfuse Support: 删除整个数据集") [Update dataset scores](https://github.com/orgs/langfuse/discussions/9377 "Langfuse Support: Update dataset scores") [Dataset Experiment via SDK](https://github.com/orgs/langfuse/discussions/9376 "Langfuse Support: Dataset Experiment via SDK") [CreateScoreRequest.DatasetRunId not working?](https://github.com/orgs/langfuse/discussions/9363 "Langfuse Support: CreateScoreRequest.DatasetRunId not working?") [how to use Evaluator with problem, answer, and Expected output?](https://github.com/orgs/langfuse/discussions/9360 "Langfuse Support: how to use Evaluator with problem, answer, and Expected output?") [How to fetch the traces using custom fields ingested in dashboard?](https://github.com/orgs/langfuse/discussions/9352 "Langfuse Support: How to fetch the traces using custom fields ingested in dashboard?") [Is it possible for LLM-as-a-Judge to produce a categorical answer?](https://github.com/orgs/langfuse/discussions/9325 "Langfuse Support: Is it possible for LLM-as-a-Judge to produce a categorical answer?") [How to link an AI-as-judge eval to a specific observation in Langfuse Cloud](https://github.com/orgs/langfuse/discussions/9248 "Langfuse Support: How to link an AI-as-judge eval to a specific observation in Langfuse Cloud") [Given a sessionId, how can i fetch the scores associated with that session ( not its traces)](https://github.com/orgs/langfuse/discussions/9160 "Langfuse Support: Given a sessionId, how can i fetch the scores associated with that session ( not its traces)") [Dataset run on Composite Prompts](https://github.com/orgs/langfuse/discussions/9155 "Langfuse Support: Dataset run on Composite Prompts") [SDK支持编辑单个数据](https://github.com/orgs/langfuse/discussions/9149 "Langfuse Support: SDK支持编辑单个数据") [Comparing the LLM as judge prompts on the UI](https://github.com/orgs/langfuse/discussions/9138 "Langfuse Support: Comparing the LLM as judge prompts on the UI") [Scores – limitations and improvement suggestions](https://github.com/orgs/langfuse/discussions/9125 "Langfuse Support: Scores – limitations and improvement suggestions") [Automating eval runs](https://github.com/orgs/langfuse/discussions/8898 "Langfuse Support: Automating eval runs") [API - /scores not respecting value when operator '='](https://github.com/orgs/langfuse/discussions/8770 "Langfuse Support: API - /scores not respecting value when operator '='") [Does LangFuse support evaluations on an existing dataset (.csv)](https://github.com/orgs/langfuse/discussions/8665 "Langfuse Support: Does LangFuse support evaluations on an existing dataset (.csv)") [deleting evaluators](https://github.com/orgs/langfuse/discussions/8640 "Langfuse Support: deleting evaluators") [Not able to import my LLM-as-a-Judge evals](https://github.com/orgs/langfuse/discussions/8636 "Langfuse Support: Not able to import my LLM-as-a-Judge evals") [How to get experiment run scores programmatically?](https://github.com/orgs/langfuse/discussions/8590 "Langfuse Support: How to get experiment run scores programmatically?") [Scoring using span object or using trace id doesn't seem to work in remote runs for dataset evaluations](https://github.com/orgs/langfuse/discussions/8556 "Langfuse Support: Scoring using span object or using trace id doesn't seem to work in remote runs for dataset evaluations") [Dataset runs restore from backups](https://github.com/orgs/langfuse/discussions/8534 "Langfuse Support: Dataset runs restore from backups") [Getting scores efficiently via API for analytics purposes](https://github.com/orgs/langfuse/discussions/8520 "Langfuse Support: Getting scores efficiently via API for analytics purposes") [run experiment on dataset](https://github.com/orgs/langfuse/discussions/8433 "Langfuse Support: run experiment on dataset") [Experiments on Datasets with Human Annotated Labels?](https://github.com/orgs/langfuse/discussions/8414 "Langfuse Support: Experiments on Datasets with Human Annotated Labels?") [How to Recalculate Total Score on Dashboard After Updating User-Defined Model?](https://github.com/orgs/langfuse/discussions/8375 "Langfuse Support: How to Recalculate Total Score on Dashboard After Updating User-Defined Model?") [How to filter by Categorical Scores in custom dashboard?](https://github.com/orgs/langfuse/discussions/8356 "Langfuse Support: How to filter by Categorical Scores in custom dashboard?") [Running scheduled evals utilising LangFuse Datasets & Evaluators](https://github.com/orgs/langfuse/discussions/8355 "Langfuse Support: Running scheduled evals utilising LangFuse Datasets & Evaluators") [Custom trace\_id for traces created with dataset's \`item.run()\` method](https://github.com/orgs/langfuse/discussions/8254 "Langfuse Support: Custom trace_id for traces created with dataset's `item.run()` method") [how to create langfuse datasetrun in langfuse-java sdk?](https://github.com/orgs/langfuse/discussions/8184 "Langfuse Support: how to create langfuse datasetrun in langfuse-java sdk?") [\[Experiment\] Send items in parallel. Stop experiment.](https://github.com/orgs/langfuse/discussions/7701 "Langfuse Support: [Experiment] Send items in parallel. Stop experiment.") [Focused Mode shows only the final ai.generateObject span—earlier generations are missing](https://github.com/orgs/langfuse/discussions/6669 "Langfuse Support: Focused Mode shows only the final ai.generateObject span—earlier generations are missing") [Using scores in data sets from the SDK?](https://github.com/orgs/langfuse/discussions/6016 "Langfuse Support: Using scores in data sets from the SDK?") [Results for some data items not present when comparing experiments](https://github.com/orgs/langfuse/discussions/5928 "Langfuse Support: Results for some data items not present when comparing experiments") [Deleting Metrics for Langfuse](https://github.com/orgs/langfuse/discussions/5849 "Langfuse Support: Deleting Metrics for Langfuse") [Discrepancies between dataset items found in the UI vs retrieved from the SDK/API](https://github.com/orgs/langfuse/discussions/5822 "Langfuse Support: Discrepancies between dataset items found in the UI vs retrieved from the SDK/API") [Datasets Not Fully Aligned - Dataset Item Runs and Trace Dataset Associated](https://github.com/orgs/langfuse/discussions/5745 "Langfuse Support: Datasets Not Fully Aligned - Dataset Item Runs and Trace Dataset Associated") [Is there an option to text wrap the output responses while comparing dataset runs?](https://github.com/orgs/langfuse/discussions/5544 "Langfuse Support: Is there an option to text wrap the output responses while comparing dataset runs?") [Score function](https://github.com/orgs/langfuse/discussions/5257 "Langfuse Support: Score function") [Support for Metric Calculation (Precision@K, Recall@K) and Adding Custom Metrics Use Case Overview](https://github.com/orgs/langfuse/discussions/5215 "Langfuse Support: Support for Metric Calculation (Precision@K, Recall@K) and Adding Custom Metrics Use Case Overview") [Unable to link dataset run items to generation observation when using observe decorator](https://github.com/orgs/langfuse/discussions/4845 "Langfuse Support: Unable to link dataset run items to generation observation when using observe decorator") [Dataset creates a lot of identical processes that run infinitely.](https://github.com/orgs/langfuse/discussions/4844 "Langfuse Support: Dataset creates a lot of identical processes that run infinitely.") [How to deactivate evaluator in llm-as-a-judge programmatically (through SDK or API)?](https://github.com/orgs/langfuse/discussions/4560 "Langfuse Support: How to deactivate evaluator in llm-as-a-judge programmatically (through SDK or API)?") [Prompt Experiments Not Generating Traces](https://github.com/orgs/langfuse/discussions/4505 "Langfuse Support: Prompt Experiments Not Generating Traces") [Templates for LLM-AS-a-judge](https://github.com/orgs/langfuse/discussions/4502 "Langfuse Support: Templates for LLM-AS-a-judge") [Take chat history in consideration when running a prompt experiment](https://github.com/orgs/langfuse/discussions/4452 "Langfuse Support: Take chat history in consideration when running a prompt experiment") [Cannot use prompt experiments "No dataset item contains any variables"](https://github.com/orgs/langfuse/discussions/4418 "Langfuse Support: Cannot use prompt experiments \"No dataset item contains any variables\"") [How to update a score?](https://github.com/orgs/langfuse/discussions/4178 "Langfuse Support: How to update a score?") [Can I evaluate Span using the External Evaluation Pipeline?](https://github.com/orgs/langfuse/discussions/4118 "Langfuse Support: Can I evaluate Span using the External Evaluation Pipeline?") [Allow injection of generation context into evaluation prompt](https://github.com/orgs/langfuse/discussions/3905 "Langfuse Support: Allow injection of generation context into evaluation prompt") [Weird behaviour of metric values and their reasoning](https://github.com/orgs/langfuse/discussions/3799 "Langfuse Support: Weird behaviour of metric values and their reasoning") [Filter Categorical Score Values](https://github.com/orgs/langfuse/discussions/3797 "Langfuse Support: Filter Categorical Score Values") [How to create an eval config for prompts using python api?](https://github.com/orgs/langfuse/discussions/3756 "Langfuse Support: How to create an eval config for prompts using python api?") [Can I use Amazon Bedrock for Langfuse Evals?](https://github.com/orgs/langfuse/discussions/3698 "Langfuse Support: Can I use Amazon Bedrock for Langfuse Evals?") [Configuring Evaluation with "Correctness" Template & Python Code Invocation](https://github.com/orgs/langfuse/discussions/3410 "Langfuse Support: Configuring Evaluation with \"Correctness\" Template & Python Code Invocation") [Self-Host evaluation feature](https://github.com/orgs/langfuse/discussions/3393 "Langfuse Support: Self-Host evaluation feature") [Hey want to change the Eval Templates name, can we do it from UI](https://github.com/orgs/langfuse/discussions/3364 "Langfuse Support: Hey want to change the Eval Templates name, can we do it from UI") [DataSet Scores are not being displayed](https://github.com/orgs/langfuse/discussions/3307 "Langfuse Support: DataSet Scores are not being displayed") [How run Langfuse evaluations over specifics spans?](https://github.com/orgs/langfuse/discussions/2852 "Langfuse Support: How run Langfuse evaluations over specifics spans?") [Getting all traces logged in a timerange for custom scoring](https://github.com/orgs/langfuse/discussions/2481 "Langfuse Support: Getting all traces logged in a timerange for custom scoring") [2 traces generated instead of 1](https://github.com/orgs/langfuse/discussions/2244 "Langfuse Support: 2 traces generated instead of 1") [Evaluations Not Available in Self-Hosted Version?](https://github.com/orgs/langfuse/discussions/2130 "Langfuse Support: Evaluations Not Available in Self-Hosted Version?") [Deleting Duplicate Items in a Dataset](https://github.com/orgs/langfuse/discussions/2099 "Langfuse Support: Deleting Duplicate Items in a Dataset") [Availability of evals when self-hosting](https://github.com/orgs/langfuse/discussions/2042 "Langfuse Support: Availability of evals when self-hosting") [How to utilize a dataset w/ typescript and langchain integration](https://github.com/orgs/langfuse/discussions/1969 "Langfuse Support: How to utilize a dataset w/ typescript and langchain integration") [Scoring a trace after the LLM chain returns](https://github.com/orgs/langfuse/discussions/1610 "Langfuse Support: Scoring a trace after the LLM chain returns") [Update/delete score using python sdk](https://github.com/orgs/langfuse/discussions/1486 "Langfuse Support: Update/delete score using python sdk") [Linking dataset run items with existing callback handler](https://github.com/orgs/langfuse/discussions/1445 "Langfuse Support: Linking dataset run items with existing callback handler") [Datasets list / by id](https://github.com/orgs/langfuse/discussions/1420 "Langfuse Support: Datasets list / by id") [Run items not appearing when linking to a trace and not a span or a generation](https://github.com/orgs/langfuse/discussions/1357 "Langfuse Support: Run items not appearing when linking to a trace and not a span or a generation") [Cannot see Add to Dataset button in the UI](https://github.com/orgs/langfuse/discussions/866 "Langfuse Support: Cannot see Add to Dataset button in the UI")

GitHubSupportGitHubIdeas

Upvotes

[NewGitHub](https://github.com/orgs/langfuse/discussions/new/choose)

- [9votes\\
\\
Update/delete score using python sdk\\
\\
msanand•3/25/2024•\\
\\
2Resolved](https://github.com/orgs/langfuse/discussions/1486)
- [7votes\\
\\
Allow defining user ID when creating score via SDK\\
\\
marcjaner•11/4/2025•\\
\\
2](https://github.com/orgs/langfuse/discussions/10187)
- [5votes\\
\\
How to get experiment run scores programmatically?\\
\\
anuras•8/18/2025•\\
\\
1Resolved](https://github.com/orgs/langfuse/discussions/8590)
- [5votes\\
\\
Is there an option to text wrap the output responses while comparing dataset runs?\\
\\
adityadev11•2/14/2025•\\
\\
2Resolved](https://github.com/orgs/langfuse/discussions/5544)
- [5votes\\
\\
Cannot use prompt experiments "No dataset item contains any variables"\\
\\
j10sanders•11/25/2024•\\
\\
5Resolved](https://github.com/orgs/langfuse/discussions/4418)
- [5votes\\
\\
Filter Categorical Score Values\\
\\
alabrashJr•10/17/2024•\\
\\
3Resolved](https://github.com/orgs/langfuse/discussions/3797)
- [4votes\\
\\
How to evaluate correct tool calls are being made?\\
\\
sesprit•10/6/2025•\\
\\
3](https://github.com/orgs/langfuse/discussions/9554)

Discussions last updated: 7/18/2026, 2:42:32 AM (55 hours ago)

* * *

Was this page helpful?

Good

Bad

[Support](/content/support/index.html)

* * *

Last edited 3/20/2026

[PreviousExperiments in CI/CD](/content/docs/evaluation/experiments/experiments-ci-cd/index.html) [NextOverview](/content/docs/metrics/overview/index.html)

### On this page

[Troubleshooting and FAQ](/content/docs/evaluation/troubleshooting-and-faq#troubleshooting-and-faq/index.html) [FAQ](/content/docs/evaluation/troubleshooting-and-faq#faq/index.html) [GitHub Discussions](/content/docs/evaluation/troubleshooting-and-faq#github-discussions/index.html)

Actions

[Give us feedback](https://github.com/langfuse/langfuse-docs/issues/new?title=Feedback+for+%22Troubleshooting+and+FAQ%22&labels=feedback) [Edit this page on GitHub](https://github.com/langfuse/langfuse-docs/edit/main/content/docs/evaluation/troubleshooting-and-faq.mdx)

Contributors

Last edited Mar 20, 2026

Jannik Maierhöfer](https://github.com/jannikmaierhoefer) Marc Klingen](https://github.com/marcklingen)

[GitHub](https://github.com/langfuse/langfuse) [X](https://x.com/langfuse)

`Ask AI`  `A`

`A`
