This is an early release preview. You may encounter bugs.

Catalogue

Submit a tool

Harnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.

51 tools · showing 25–48

qa ×
Tool

Godot MCP

CodingQA

satelliteoflove

Give your AI assistant eyes and hands in the Godot editor: scene editing, input injection, deterministic playtesting, and live game state fo

godotgame-developmentplaytestingscene-editing

A
Tool

Agents CLI

CodingDeploymentQA

google

The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.

adkgcpcliskill

A
Platform

Braintrust

QAMonitoring

braintrustdata

Eval, logging and prompt-playground platform for LLM and agent applications.

evaluationobservabilityprompt

B
Tool

Harness Score

CodingQA

paladini

Your AI coding agent is only as reliable as the harness around it. Measure that harness in seconds with harness-score.

clicoding-agentstatic-analysismaturity-model

B
Tool

Raindrop Workshop

QAMonitoringCoding

raindrop-ai

Open-source tool that lets a coding agent write and run agent evals locally

observabilityevaluationdebugginglocal-first

B
Tool

Spacedock

CodingQA

clkao

Pi ships four tools ... read, write, edit, bash ... and calls it done. No MCP, no sub-agents, no permission dialogs, no

human-in-the-loopmulti-agentorchestrationgovernance

B
Harness

AIDE ML

CodingQA

WecoAI

LLM-driven agent that writes, evaluates and improves machine-learning code.

clitree-searchautomlevaluation

B
Tool

Autocontext

CodingQA

greyhaven-ai

Agents that start cold every run can't compound what worked. autocontext splits improvement into five roles per cycle …

token-optimizationself-improvementplaybooksevaluation-loop

B
Tool

MCP Server Tauri

CodingQA

hypothesi

A Model Context Protocol (MCP) server and plugin for Tauri v2 development

tauricomputer-usedebuggingscreenshot

B
Tool

Terminal-Bench

QACoding

laude-institute

Benchmark and harness for evaluating AI agents on end-to-end tasks in real terminal environments.

evaluationcliterminal-tasks

B
Tool

Chrome Agent

InterfaceQA

captivus

Open-source CLI that controls Chrome over the DevTools Protocol so agents can browse with a sense-act-verify loop instead of an MCP browser server.

browser-automationchromeclidebugging

B
Tool

Looper

CodingQA

ksimback

Claude Code skill for designing review-gated agent loops - goal, plan, review, deliver, judge - then emitting a runnable spec

evaluationyamlworkflow-design

B
Tool

Scorable

QAMonitoring

root-signals

Evaluation platform (formerly Root Signals) exposing judges and evaluators to agents via SDK and MCP for in-loop quality scoring.

evaluationsdk

B
Tool

Tracely

MonitoringQADeployment

Jwuthri

You fix the agent bug and write the regression test. Then you try to reproduce the run: t...

observabilityci-cdevaluationregression-testing

B
Tool

Evidently

QAMonitoring

evidentlyai

Open-source Python framework to evaluate, test, and monitor ML and LLM systems from experiments to production

data-driftobservabilityevaluationllmops

B
Tool

Open RAG Eval

QA

vectara

Open-source framework for evaluating RAG pipelines without golden answers

ragevaluationobservabilityvectara

C
Tool

Gorilla

QATraining

ShishirPatil

Fine-tuned models, datasets, and the Berkeley leaderboard for training and evaluating LLM API and function calling

tool-callingevaluationhuggingfacedataset

C
Tool

SimWorld

QA

SimWorld-AI

SimWorld now lets you drop Clawbot agents into a virtual city and watch what happens. They wake up, commute, run errands

simulationunreal-engineroboticsevaluation

C
Tool

Hermes Agent Self-Evolution

CodingQA

NousResearch

Evolves Hermes Agent skills and prompts with DSPy and GEPA, gated by tests and human PR review

dspygepapromptpull-request

C
Tool

MultiTown

CodingQA

w1u2d3i4

Code-only runtime toolkit for cost-aware multi-agent organization and control

multi-agentorchestrationsimulationreplay

C
Platform

Vivaria

QAResearch

METR

Self-hostable web app, server and CLI for running agents on tasks in sandboxed environments and analysing the runs.

clisandboxtask-environments

C
Platform

Adaptive Engine

TrainingInferenceQA

RLOps platform for continuously fine-tuning, evaluating and serving specialised LLMs from production feedback (company acquired by Datadog).

rlopsfine-tuneproduction-feedback

Score unavailable
Platform

Contextual AI Platform

MemoryQA

Managed platform for building and hosting grounded RAG agents over enterprise documents, with API + GUI and built-in evals.

ragdocument-processingknowledge-managementevaluation

Score unavailable
Platform

MutagenT

QAMonitoring

Agent-engineering platform: offline build-and-eval loops, trace diagnosis and automated optimization; official Claude Code skills at github.com/mutagent-io.

evaluationobservabilitytrace-diagnosisagent-lifecycle

Score unavailable