Catalogue
Submit a toolHarnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.
48 tools · showing 1–24
Langfuse
MonitoringQA
langfuse
Open-source platform for tracing, evaluating, and debugging LLM applications, self-hosted or cloud
observabilityevaluationpromptlocal-first
NeMo Agent Toolkit
CodingMonitoringQA
NVIDIA
Library for connecting, profiling and optimizing teams of agents across frameworks.
profilingobservabilityevaluationmulti-agent
ToolOpenLIT
MonitoringSecurity
openlit
Open-source AI engineering platform with OpenTelemetry-native LLM observability, evaluations, and prompt management
telemetryevaluationpromptguardrails
ToolOpik
Monitoring
comet-ml
Open-source platform for tracing, evaluating, and monitoring LLM and agentic applications
observabilityevaluationpromptlangchain
Agent Development Kit (ADK)
CodingQADeployment
Google's open-source SDK for building, evaluating and deploying multi-agent systems.
multi-agentadk-webevaluation
ToolOpenSRE
DeploymentMonitoring
Tracer-Cloud
Build your own AI SRE agents. The open source toolkit for the AI era.
sreincident-responsedebuggingobservability
ToolGiskard
QASecurity
Giskard-AI
Python library for testing agentic systems — evals with LLM-as-judge checks, red-teaming scans, and RAG quality evaluation
red-teamevaluationragprompt-injection
ToolLatitude
Monitoring
latitude-dev
Latitude traces your agent in production, finds the failures, and dispatches your coding agent to fix them.
observabilityissue-detectiontelemetryevaluation
ToolLM Evaluation Harness
Coding
EleutherAI
Framework for evaluating language models across 60+ academic benchmarks through a tokenization-agnostic, multi-backend interface
evaluationhuggingfacevllmtransformers
FrameworkVerifiers
Monitoring
PrimeIntellect-ai
Library of shareable RL environments and rubric-based verifiers for agents.
reinforcement-learningevaluation
ToolHindsight
Memory
vectorize-io
Agent memory system that stores and retrieves long-term memories so agents learn across sessions, not just recall history
long-term-memorydockerevaluationrag
TooliFixAi
MonitoringSecurity
ifixai-ai
Catch your AI's mistakes and blind spots before your customers or regulators do..
governancered-teamevaluationcli
ToolPhoenix
Monitoring
Arize-ai
Open-source AI observability platform for tracing, evaluating, and troubleshooting LLM and agent applications
observabilityevaluationllmopsrag
ToolPromptfoo
QASecurity
promptfoo
CLI and library for evaluating and red-teaming LLM apps, with side-by-side model comparison and CI/CD checks
eval-harnessred-teamci-cdvulnerability-scanner
ToolMLflow
Monitoring
mlflow
Open-source platform to trace, evaluate, monitor, and deploy LLM applications, agents, and ML models
observabilityevaluationllmopsmlops
ToolLangsmith SDK
Coding
langchain-ai
Python and JavaScript SDKs for tracing, evaluating and monitoring LLM apps on the LangSmith platform
observabilityevaluationlangchain
PlatformHarbor
QADeployment
harbor-framework
Framework and infrastructure for running arbitrary agents (Claude Code, OpenHands, Codex CLI) in thousands of parallel sandboxed environments for evaluation and RL rollout generation.
evaluationsandboxreinforcement-learningcli
ToolInspector
CodingDeploymentSecurity
MCPJam
Development and testing platform to debug, chat with, inspect, and run evals against MCP servers, MCP apps, and ChatGPT apps
oauthevaluationdebuggingobservability
Toolmini-SWE-agent
Coding
SWE-agent
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified
coding-agentevaluationlightweightlitellm
HarnessOuroboros
Coding
Q00
Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop.
coding-agentevaluationspec-drivenhuman-in-the-loop
ToolDeepEval
Coding
confident-ai
Open-source framework for unit-testing and evaluating LLM apps with ready-made metrics that run locally
evaluationragpytest
ToolWeave
MonitoringQA
wandb
Tracing, evaluation and LLM-as-judge scoring for agent applications.
observabilityevaluationwandb
PlatformBisheng
Monitoring
dataelement
BISHENG is an open LLM devops platform for next generation Enterprise AI applications. Powerful and comprehensive features include: GenAI wo
ragorchestrationfine-tunemodel-management
ToolrLLM
TrainingQA
rllm-org
rLLM bolts RL onto agents you already wrote. verl, trlx, OpenRLHF make you rewrite the agent into their pipeline. A deco
reinforcement-learningagent-trainingsandboxdistributed-training