Catalogue
Submit a toolHarnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.
51 tools · showing 25–48
ToolGodot MCP
CodingQA
satelliteoflove
Give your AI assistant eyes and hands in the Godot editor: scene editing, input injection, deterministic playtesting, and live game state fo
godotgame-developmentplaytestingscene-editing
ToolAgents CLI
CodingDeploymentQA
The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.
adkgcpcliskill
PlatformBraintrust
QAMonitoring
braintrustdata
Eval, logging and prompt-playground platform for LLM and agent applications.
evaluationobservabilityprompt
ToolHarness Score
CodingQA
paladini
Your AI coding agent is only as reliable as the harness around it. Measure that harness in seconds with harness-score.
clicoding-agentstatic-analysismaturity-model
ToolRaindrop Workshop
QAMonitoringCoding
raindrop-ai
Open-source tool that lets a coding agent write and run agent evals locally
observabilityevaluationdebugginglocal-first
ToolSpacedock
CodingQA
clkao
Pi ships four tools ... read, write, edit, bash ... and calls it done. No MCP, no sub-agents, no permission dialogs, no
human-in-the-loopmulti-agentorchestrationgovernance
HarnessAIDE ML
CodingQA
WecoAI
LLM-driven agent that writes, evaluates and improves machine-learning code.
clitree-searchautomlevaluation
ToolAutocontext
CodingQA
greyhaven-ai
Agents that start cold every run can't compound what worked. autocontext splits improvement into five roles per cycle …
token-optimizationself-improvementplaybooksevaluation-loop
MCP Server Tauri
CodingQA
hypothesi
A Model Context Protocol (MCP) server and plugin for Tauri v2 development
tauricomputer-usedebuggingscreenshot
ToolTerminal-Bench
QACoding
laude-institute
Benchmark and harness for evaluating AI agents on end-to-end tasks in real terminal environments.
evaluationcliterminal-tasks
ToolChrome Agent
InterfaceQA
captivus
Open-source CLI that controls Chrome over the DevTools Protocol so agents can browse with a sense-act-verify loop instead of an MCP browser server.
browser-automationchromeclidebugging
ToolLooper
CodingQA
ksimback
Claude Code skill for designing review-gated agent loops - goal, plan, review, deliver, judge - then emitting a runnable spec
evaluationyamlworkflow-design
ToolScorable
QAMonitoring
root-signals
Evaluation platform (formerly Root Signals) exposing judges and evaluators to agents via SDK and MCP for in-loop quality scoring.
evaluationsdk
ToolTracely
MonitoringQADeployment
Jwuthri
You fix the agent bug and write the regression test. Then you try to reproduce the run: t...
observabilityci-cdevaluationregression-testing
ToolEvidently
QAMonitoring
evidentlyai
Open-source Python framework to evaluate, test, and monitor ML and LLM systems from experiments to production
data-driftobservabilityevaluationllmops
ToolOpen RAG Eval
QA
vectara
Open-source framework for evaluating RAG pipelines without golden answers
ragevaluationobservabilityvectara
ToolGorilla
QATraining
ShishirPatil
Fine-tuned models, datasets, and the Berkeley leaderboard for training and evaluating LLM API and function calling
tool-callingevaluationhuggingfacedataset
SimWorld
QA
SimWorld-AI
SimWorld now lets you drop Clawbot agents into a virtual city and watch what happens. They wake up, commute, run errands
simulationunreal-engineroboticsevaluation
Hermes Agent Self-Evolution
CodingQA
NousResearch
Evolves Hermes Agent skills and prompts with DSPy and GEPA, gated by tests and human PR review
dspygepapromptpull-request
ToolMultiTown
CodingQA
w1u2d3i4
Code-only runtime toolkit for cost-aware multi-agent organization and control
multi-agentorchestrationsimulationreplay
PlatformVivaria
QAResearch
METR
Self-hostable web app, server and CLI for running agents on tasks in sandboxed environments and analysing the runs.
clisandboxtask-environments
Adaptive Engine
TrainingInferenceQA
RLOps platform for continuously fine-tuning, evaluating and serving specialised LLMs from production feedback (company acquired by Datadog).
rlopsfine-tuneproduction-feedback
Contextual AI Platform
MemoryQA
Managed platform for building and hosting grounded RAG agents over enterprise documents, with API + GUI and built-in evals.
ragdocument-processingknowledge-managementevaluation
MutagenT
QAMonitoring
Agent-engineering platform: offline build-and-eval loops, trace diagnosis and automated optimization; official Claude Code skills at github.com/mutagent-io.
evaluationobservabilitytrace-diagnosisagent-lifecycle