This is an early release preview. You may encounter bugs.

Catalogue

Submit a tool

Harnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.

48 tools · showing 1–24

#evaluation ×
Tool

Langfuse

MonitoringQA

langfuse

Open-source platform for tracing, evaluating, and debugging LLM applications, self-hosted or cloud

observabilityevaluationpromptlocal-first

A+
Framework

NeMo Agent Toolkit

CodingMonitoringQA

NVIDIA

Library for connecting, profiling and optimizing teams of agents across frameworks.

profilingobservabilityevaluationmulti-agent

A+
Tool

OpenLIT

MonitoringSecurity

openlit

Open-source AI engineering platform with OpenTelemetry-native LLM observability, evaluations, and prompt management

telemetryevaluationpromptguardrails

A+
Tool

Opik

Monitoring

comet-ml

Open-source platform for tracing, evaluating, and monitoring LLM and agentic applications

observabilityevaluationpromptlangchain

A+
Framework

Agent Development Kit (ADK)

CodingQADeployment

google

Google's open-source SDK for building, evaluating and deploying multi-agent systems.

multi-agentadk-webevaluation

A+
Tool

OpenSRE

DeploymentMonitoring

Tracer-Cloud

Build your own AI SRE agents. The open source toolkit for the AI era.

sreincident-responsedebuggingobservability

A+
Tool

Giskard

QASecurity

Giskard-AI

Python library for testing agentic systems — evals with LLM-as-judge checks, red-teaming scans, and RAG quality evaluation

red-teamevaluationragprompt-injection

A+
Tool

Latitude

Monitoring

latitude-dev

Latitude traces your agent in production, finds the failures, and dispatches your coding agent to fix them.

observabilityissue-detectiontelemetryevaluation

A+
Tool

LM Evaluation Harness

Coding

EleutherAI

Framework for evaluating language models across 60+ academic benchmarks through a tokenization-agnostic, multi-backend interface

evaluationhuggingfacevllmtransformers

A+
Framework

Verifiers

Monitoring

PrimeIntellect-ai

Library of shareable RL environments and rubric-based verifiers for agents.

reinforcement-learningevaluation

A+
Tool

Hindsight

Memory

vectorize-io

Agent memory system that stores and retrieves long-term memories so agents learn across sessions, not just recall history

long-term-memorydockerevaluationrag

A
Tool

iFixAi

MonitoringSecurity

ifixai-ai

Catch your AI's mistakes and blind spots before your customers or regulators do..

governancered-teamevaluationcli

A
Tool

Phoenix

Monitoring

Arize-ai

Open-source AI observability platform for tracing, evaluating, and troubleshooting LLM and agent applications

observabilityevaluationllmopsrag

A
Tool

Promptfoo

QASecurity

promptfoo

CLI and library for evaluating and red-teaming LLM apps, with side-by-side model comparison and CI/CD checks

eval-harnessred-teamci-cdvulnerability-scanner

A
Tool

MLflow

Monitoring

mlflow

Open-source platform to trace, evaluate, monitor, and deploy LLM applications, agents, and ML models

observabilityevaluationllmopsmlops

A
Tool

Langsmith SDK

Coding

langchain-ai

Python and JavaScript SDKs for tracing, evaluating and monitoring LLM apps on the LangSmith platform

observabilityevaluationlangchain

A
Platform

Harbor

QADeployment

harbor-framework

Framework and infrastructure for running arbitrary agents (Claude Code, OpenHands, Codex CLI) in thousands of parallel sandboxed environments for evaluation and RL rollout generation.

evaluationsandboxreinforcement-learningcli

A
Tool

Inspector

CodingDeploymentSecurity

MCPJam

Development and testing platform to debug, chat with, inspect, and run evals against MCP servers, MCP apps, and ChatGPT apps

oauthevaluationdebuggingobservability

A
Tool

mini-SWE-agent

Coding

SWE-agent

The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified

coding-agentevaluationlightweightlitellm

A
Harness

Ouroboros

Coding

Q00

Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop.

coding-agentevaluationspec-drivenhuman-in-the-loop

A
Tool

DeepEval

Coding

confident-ai

Open-source framework for unit-testing and evaluating LLM apps with ready-made metrics that run locally

evaluationragpytest

A
Tool

Weave

MonitoringQA

wandb

Tracing, evaluation and LLM-as-judge scoring for agent applications.

observabilityevaluationwandb

A
Platform

Bisheng

Monitoring

dataelement

BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI wo

ragorchestrationfine-tunemodel-management

A
Tool

rLLM

TrainingQA

rllm-org

rLLM bolts RL onto agents you already wrote. verl, trlx, OpenRLHF make you rewrite the agent into their pipeline. A deco

reinforcement-learningagent-trainingsandboxdistributed-training

A