# How to set up a user feedback loop for your application

Once traces are coming in, capturing user feedback is a great next step. Your users judge every response through what they click, edit, retry, or say. Capturing that on your traces gives you a quality signal grounded in real usage, where each score points at a trace you can open and learn from.

A lot of teams ask _"I have traces coming in, how do I now get started with evals?"_. Often, these teams are better off starting with user feedback first.

This feedback gathering step generally has a very high ROI, in the AI Engineering process. We often write about this process in the [Langfuse Academy](/content/academy/index.html), where we walk through the [AI engineering loop](/content/academy/ai-engineering-loop/index.html) in more detail. Gathering user feedback happens in the monitoring stage highlighted below.

## Prerequisites

You already have traces coming into Langfuse. See [tracing](/content/docs/observability/get-started/index.html) if you have not set this up yet.

## Walkthrough

### Choose your feedback signals

Feedback comes in two forms. **Explicit feedback** is a rating the user gives on purpose, like a thumbs up or down or a star rating: unambiguous, but rare and skewed toward unhappy users. **Implicit feedback** is derived from what users do, like retrying a query or editing a draft: you have data on every trace, but it needs interpretation.

The best signals depend heavily on your use case. Some examples as inspiration:

| Signal | What it tells you | Typically captured via |
| --- | --- | --- |
| Thumbs up or down on a response | Direct rating of that response | Browser SDK |
| A user rephrasing the same question | The previous answer didn't land | LLM-as-a-Judge |
| A user asking to speak to a human | The user stopped trusting the assistant | LLM-as-a-Judge |
| A response copied by the user | The output was good enough to reuse | Browser SDK |
| A drafted reply edited before it is sent | What was wrong or missing in the draft | SDK / API |
| A copilot suggestion accepted or dismissed | Whether the suggestion fit the context | Browser SDK |
| An extracted field corrected in a review step | Which field was extracted incorrectly | SDK / API |
| A search query reformulated without clicking a result | The results missed the intent | SDK / API |
| A recommended item skipped right after it starts | The pick missed the user's taste | SDK / API |

Tips:
- signals combine well, no need to implement only one
- capture free-form text wherever it fits (can be as a [comment on the score](/content/docs/evaluation/scores/data-model/index.html)), as much of this analysis is now done by agents, and they work better with more context
- the [academy examples](/content/academy/examples/index.html) show how signals are chosen for a specific application and change over its lifecycle

### Capture the signals as scores

Every signal ends up as a [score](/content/docs/evaluation/scores/data-model/index.html) on the trace that produced the output. A score has a name, a value with a data type (boolean, numeric, categorical or free text), and an optional `comment`. Evaluators also store their results as scores, so your user feedback can land in the same filters, dashboards, and analytics.

### Act on the feedback

#### Monitor how quality is trending

[Score analytics](/content/docs/evaluation/scores/score-analytics/index.html) and [custom dashboards](/content/docs/metrics/features/custom-dashboards/index.html) chart how each signal moves over time. One caveat: not every signal is a clean quality metric, e.g. thumbs feedback skews toward unhappy users.

#### Surface interesting traces

Filter traces by score, for example `response_rating = 0`, to get a concrete list of traces to investigate. [Error analysis](/content/academy/monitoring/error-analysis/index.html) is a structured way to read and cluster them, and an [annotation queue](/content/docs/evaluation/evaluation-methods/annotation-queues/index.html) brings in more reviewers.

#### Use it as input for structured improvement

Add the failing cases to a [dataset](/content/docs/evaluation/experiments/datasets/index.html), test a fix with an [experiment](/content/docs/evaluation/experiments/experiments-via-ui/index.html) before shipping, and turn a recurring failure mode into an [automated evaluator](/content/docs/evaluation/core-concepts#evaluation-methods/index.html) that scores every production trace from then on.

**Much of this can also be handed to an agent**, see [Let an agent act on the feedback](/content/guides/user-feedback-loop#agents/index.html) below.

## Let an agent act on the feedback

Langfuse is built for agent access through the [Langfuse CLI](/content/docs/api-and-data-platform/features/cli/index.html), the [Agent Skill](/content/docs/api-and-data-platform/features/agent-skill/index.html), and the [MCP Server](/content/docs/api-and-data-platform/features/mcp-server/index.html): point one at your project and it can interpret and act on the behavior that gets surfaced via user feedback. An example task:

```
Fetch the traces from the last 7 days with a response_rating score of 0.
Read the comments and outputs, cluster the failures into categories,
and propose a prompt change for the two biggest ones.
```

Some reading material as inspiration:
- [Using Agent Skills to Automatically Improve your Prompts](/content/blog/2026-02-16-prompt-improvement-claude-skills/index.html)
- [AI is eating the AI engineering loop](/content/blog/2026-06-09-ai-is-eating-ai-engineering/index.html)
