Catalogue
Submit a toolHarnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.
51 tools · showing 1–24
ToolChrome DevTools MCP
CodingInterfaceQA
ChromeDevTools
Chrome DevTools for coding agents
browser-automationmcp-serverchromedebugging
Playwright
InterfaceQA
microsoft
Drives Chromium, Firefox, and WebKit through one API for e2e tests, scripts, and AI-agent automation via MCP or CLI
browser-automationchromefirefox
ToolPuppeteer
InterfaceQA
puppeteer
JavaScript library that drives Chrome or Firefox over the DevTools Protocol or WebDriver BiDi, headless by default
browser-automationheadless-chromewebdriver-bidiscrape
PlatformInsForge
CodingMemoryQA
InsForge
The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting,
backend-as-a-servicepostgrespgvectorauth
Langfuse
MonitoringQA
langfuse
Open-source platform for tracing, evaluating, and debugging LLM applications, self-hosted or cloud
observabilityevaluationpromptlocal-first
Mastra
CodingMemoryQA
mastra-ai
TypeScript framework for AI agents and apps, with model routing, graph-based workflows, memory, and built-in evals
orchestrationragnextjs
Midscene
QAInterface
web-infra-dev
Vision-driven UI automation that acts from screenshots and natural-language steps across web, mobile, and desktop
playwrightcomputer-usemobile-automationvision-language-model
NeMo Agent Toolkit
CodingMonitoringQA
NVIDIA
Library for connecting, profiling and optimizing teams of agents across frameworks.
profilingobservabilityevaluationmulti-agent
Testsprite CLI
InterfaceQA
TestSprite
The verification layer for the agentic coding era. AI ships code in minutes — verifying it hasn't. testsprite opens your live app, uses it like a real user, and shows your coding agent exactly what broke.
playwrightbrowser-automationclicoding-agent
Agent Development Kit (ADK)
CodingQADeployment
Google's open-source SDK for building, evaluating and deploying multi-agent systems.
multi-agentadk-webevaluation
ToolGiskard
QASecurity
Giskard-AI
Python library for testing agentic systems — evals with LLM-as-judge checks, red-teaming scans, and RAG quality evaluation
red-teamevaluationragprompt-injection
ToolPromptfoo
QASecurity
promptfoo
CLI and library for evaluating and red-teaming LLM apps, with side-by-side model comparison and CI/CD checks
eval-harnessred-teamci-cdvulnerability-scanner
PlatformHarbor
QADeployment
harbor-framework
Framework and infrastructure for running arbitrary agents (Claude Code, OpenHands, Codex CLI) in thousands of parallel sandboxed environments for evaluation and RL rollout generation.
evaluationsandboxreinforcement-learningcli
FrameworkKitaru
QAMonitoring
zenml-io
Durable execution runtime for Python agents: checkpointed flows, replay and overrides - 'agent traces you can run, not just read'.
replayregression-detectioncheckpointingdurable-execution
ToolWeave
MonitoringQA
wandb
Tracing, evaluation and LLM-as-judge scoring for agent applications.
observabilityevaluationwandb
ToolAgent Device
QAInterface
callstack
CLI to control iOS and Android devices for AI agents
mobile-automationiosandroidaccessibility-snapshot
AgentInspect
MonitoringQA
rajudandigam
Local execution trees for TypeScript AI agents. agent-inspect helps you understand what happened inside an AI agent run — locally. It turns
execution-tracetrajectory-testingci-cdlocal-first
ToolrLLM
TrainingQA
rllm-org
rLLM bolts RL onto agents you already wrote. verl, trlx, OpenRLHF make you rewrite the agent into their pipeline. A deco
reinforcement-learningagent-trainingsandboxdistributed-training
ToolAgenta
MonitoringQA
Agenta-AI
Workspace for building agents through chat and running them in the background, with swappable harnesses and models
llmopsevaluationobservabilityagent-builder
ToolLLM Space
CodingQA
deer-flow
A desktop app to prototype agent ideas, inspect every harness step, replay failures, and evaluate performance, all in one place. Local-first
desktopobservabilitydebuggingprompt
ToolSunpeak
QAInterface
Alignbase
Server-agnostic MCP testing framework and full-stack MCP App framework for ChatGPT Apps, Claude Connectors, and more.
chatgpt-appsmcp-uihot-reloadvisual-testing
ToolFuture AGI
MonitoringQASecurity
future-agi
OpenTelemetry-based tracing and evaluation instrumentation for voice-AI applications.
telemetryevaluationguardrailsobservability
ToolArgent
QAInterface
software-mansion
An agentic toolkit to control, debug, and profile iOS and Android apps. Made by Software Mansion.
mobile-automationiosandroidmobile
ToolCodegraff
CodingQA
justrach
graff — a fast agentic coding harness in Zig: multi-provider, MCP, workflows, DGM evolution loop, TS/Python SDKs