This is an early release preview. You may encounter bugs.
Langfuse logo
Unclaimed

Tool monitoring qa

Langfuse

Open-source platform for tracing, evaluating, and debugging LLM applications, self-hosted or cloud

A+ 92/100 GitHub score ? This grade is derived from GitHub signals, not user votes. Open for the full breakdown.
No votes yet

01 / About

What Langfuse is.

Langfuse is a platform for developing, monitoring, evaluating, and debugging LLM applications. You instrument an application to send traces of LLM calls, retrieval steps, embeddings, and agent actions, then inspect those traces and user sessions in the web interface. It is built on the ClickHouse database and can be run as a managed cloud service or self-hosted.

Beyond tracing, the platform covers prompt management (central versioning with server- and client-side caching), evaluations (LLM-as-a-judge, code evaluators, user feedback, manual labelling, and custom pipelines through the API), datasets for test sets and benchmarks, and a playground for iterating on prompts and model settings. An OpenAPI spec, a Postman collection, and typed Python and JS/TS SDKs expose every building block for custom LLMOps workflows.

Self-hosting paths include Docker Compose on a local machine or a single VM, a Helm chart for Kubernetes, and Terraform templates for AWS, Azure, and GCP.

Integration Languages Method
SDK Python, JS/TS Manual instrumentation
OpenAI Python, JS/TS Drop-in replacement of the OpenAI SDK
LangChain Python, JS/TS Callback handler
LlamaIndex Python Callback system
Haystack Python Content tracing system
LiteLLM Python, JS/TS (proxy) Proxy for 100+ LLMs
Vercel AI SDK JS/TS Toolkit integration
Mastra JS/TS Framework integration

Features

  • Tracing: captures LLM calls, retrieval, embedding, and agent steps as nested traces with session views
  • Prompt management: versions prompts centrally and caches them so prompt iteration adds no request latency
  • Evaluations: LLM-as-a-judge, code evaluators, user feedback collection, manual labelling, and custom evaluation pipelines
  • Datasets: test sets and benchmarks for pre-deployment testing and structured experiments
  • Playground: tests prompts and model configurations, reachable directly from a trace
  • API and SDKs: OpenAPI spec, Postman collection, and typed Python and JS/TS SDKs
  • Deployment options: managed cloud, Docker Compose, Kubernetes with Helm, and Terraform templates for AWS, Azure, and GCP
  • Further integrations: Instructor, DSPy, Mirascope, Ollama, Amazon Bedrock, AutoGen, Flowise, Langflow, Dify, OpenWebUI, and Promptfoo

02 / Discussion CREDIBILITY-GATED

Discussion

Reading is open to everyone. Posting and voting need a verified identity or a GitHub grade of B or higher.

  • No discussions yet.

03 / Build

Build with Langfuse.

Browse the catalogue for frameworks, tools, and harnesses, each scored on real GitHub credibility.

Get Langfuse →

Browse the catalogue