Blog - Langfuse

Langfuse Blog

The latest updates from Langfuse. See Changelog for more product updates.

Highlights

engineering
Designing the runtime for Langfuse code evaluators
Code evaluators let you score traces with your own Python or TypeScript code. A look at the execution model behind them: the requirements, the options we rejected, and the security stance we adopted.
4 Weeks Ago
·
Tobias

engineering
AI is eating the AI engineering loop
The full AI engineering loop can technically be automated now. But that doesn't mean it should. Here is what we think you should hand to agents, and what you should keep doing yourself.
June 9, 2026
·
Lotte

engineering
The Rage Clicks of LLM apps: High-Signal Production Monitoring for AI Customer Support Agents
How to use LLM-as-a-judge to detect when your users say "f**k".
April 1, 2026
·
Annabell

All Posts

All [73]

engineering
How we use agents to review production infrastructure
How repo-owned agent workflows help us review incidents, infra cost, security findings, and bugs in production.
June 5, 2026
·
Max

update
Langfuse May Update
Code Evaluators, full-text search, Langfuse MCP, Experiments in CI/CD and more
May 31, 2026
·
Marc

update
Langfuse April Update
Japan Cloud Region, new Experiments, Langfuse Academy, LLM-as-a-Judge API and more
April 30, 2026
·
Marc

announcement
Langfuse Cloud 日本リージョンを開始しました
Langfuse Cloud 日本リージョンを公開しました。LLM のトレースや評価データを日本国内に保管したいチーム向けの専用クラウドリージョンです。
April 27, 2026
·
Marc, Clemens, Max

announcement
Langfuse Cloud Japan Region
Langfuse Cloud Japan is live. A dedicated cloud region hosted in Japan for teams that need their LLM observability data to stay in Japan.
April 27, 2026
·
Marc, Clemens, Max

engineering
Classifying User Intent with Categorical LLM-as-a-Judge
A guide on how to set up a categorical LLM-as-a-judge evaluator to classify user intent. Follow along with how we applied this to our demo application.
April 14, 2026
·
Lotte

update
Langfuse March Update
Agent Skill, Langfuse CLI, boolean and categorical LLM-as-a-Judge scores, Kiro integration, and more
March 31, 2026
·
Marc

engineering
We Used Autoresearch on Our AI Skill, It Taught Us to Write Better Tests
We applied Karpathy's autoresearch to optimize our Langfuse prompt migration skill — and got a lesson in why the target function matters more than the optimizer.
March 24, 2026
·
Lotte

engineering
How We Built an Agent Skill to Synthesize what Langfuse Users want
We built an AI agent skill that synthesizes GitHub issues, support tickets, and meeting notes into a weekly digest, and used Langfuse to monitor and improve it.
March 13, 2026
·
Lotte

engineering, architecture
Simplifying Langfuse for Scale
A deep dive into how we moved Langfuse to an observations-first data model.
March 10, 2026
·
Steffen, Valeriy, Max

update
Langfuse February Update
Observation-centric data model, faster UI, Observations API v2 and Metrics API v2 out of beta, faster evaluation workflows
February 28, 2026
·
Marc

engineering
Evaluating AI Agent Skills
How we used Langfuse datasets, tracing, and the cloud agent SDK to iteratively evaluate and improve our AI agent skill.
February 26, 2026
·
Lotte

guide
Using Agent Skills to Automatically Improve your Prompts
Use the Langfuse skill for Claude Code to analyze trace feedback and iteratively improve your prompts.
February 16, 2026
·
Lotte

engineering
Will you be my CLI? Making Agents fall in love with Langfuse.
How we optimized Langfuse for AI agents with CLI, Skills, markdown endpoints, public RAG endpoint, llms.txt, and MCP servers. A love letter to agents that need great developer tools.
February 14, 2026
·
Felix, Lotte, Nimar, Marc

update
Langfuse January Update
Langfuse joins ClickHouse, Claude Code integration, corrected outputs, Langfuse v4 Beta, and new integrations
January 31, 2026
·
Marc

Load 15 more