This is an early release preview. You may encounter bugs.

Catalogue

Submit a tool

Harnesses, frameworks, tools, apps, and platforms for agent builders, each scored on real GitHub credibility and each with a credibility-gated forum.

51 tools · showing 1–24

qa ×
Tool

Chrome DevTools MCP

CodingInterfaceQA

ChromeDevTools

Chrome DevTools for coding agents

browser-automationmcp-serverchromedebugging

A+
Tool

Playwright

InterfaceQA

microsoft

Drives Chromium, Firefox, and WebKit through one API for e2e tests, scripts, and AI-agent automation via MCP or CLI

browser-automationchromefirefox

A+
Tool

Puppeteer

InterfaceQA

puppeteer

JavaScript library that drives Chrome or Firefox over the DevTools Protocol or WebDriver BiDi, headless by default

browser-automationheadless-chromewebdriver-bidiscrape

A+
Platform

InsForge

CodingMemoryQA

InsForge

The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting,

backend-as-a-servicepostgrespgvectorauth

A+
Tool

Langfuse

MonitoringQA

langfuse

Open-source platform for tracing, evaluating, and debugging LLM applications, self-hosted or cloud

observabilityevaluationpromptlocal-first

A+
Framework

Mastra

CodingMemoryQA

mastra-ai

TypeScript framework for AI agents and apps, with model routing, graph-based workflows, memory, and built-in evals

orchestrationragnextjs

A+
Tool

Midscene

QAInterface

web-infra-dev

Vision-driven UI automation that acts from screenshots and natural-language steps across web, mobile, and desktop

playwrightcomputer-usemobile-automationvision-language-model

A+
Framework

NeMo Agent Toolkit

CodingMonitoringQA

NVIDIA

Library for connecting, profiling and optimizing teams of agents across frameworks.

profilingobservabilityevaluationmulti-agent

A+
Tool

Testsprite CLI

InterfaceQA

TestSprite

The verification layer for the agentic coding era. AI ships code in minutes — verifying it hasn't. testsprite opens your live app, uses it like a real user, and shows your coding agent exactly what broke.

playwrightbrowser-automationclicoding-agent

A+
Framework

Agent Development Kit (ADK)

CodingQADeployment

google

Google's open-source SDK for building, evaluating and deploying multi-agent systems.

multi-agentadk-webevaluation

A+
Tool

Giskard

QASecurity

Giskard-AI

Python library for testing agentic systems — evals with LLM-as-judge checks, red-teaming scans, and RAG quality evaluation

red-teamevaluationragprompt-injection

A+
Tool

Promptfoo

QASecurity

promptfoo

CLI and library for evaluating and red-teaming LLM apps, with side-by-side model comparison and CI/CD checks

eval-harnessred-teamci-cdvulnerability-scanner

A
Platform

Harbor

QADeployment

harbor-framework

Framework and infrastructure for running arbitrary agents (Claude Code, OpenHands, Codex CLI) in thousands of parallel sandboxed environments for evaluation and RL rollout generation.

evaluationsandboxreinforcement-learningcli

A
Framework

Kitaru

QAMonitoring

zenml-io

Durable execution runtime for Python agents: checkpointed flows, replay and overrides - 'agent traces you can run, not just read'.

replayregression-detectioncheckpointingdurable-execution

A
Tool

Weave

MonitoringQA

wandb

Tracing, evaluation and LLM-as-judge scoring for agent applications.

observabilityevaluationwandb

A
Tool

Agent Device

QAInterface

callstack

CLI to control iOS and Android devices for AI agents

mobile-automationiosandroidaccessibility-snapshot

A
Tool

AgentInspect

MonitoringQA

rajudandigam

Local execution trees for TypeScript AI agents. agent-inspect helps you understand what happened inside an AI agent run — locally. It turns

execution-tracetrajectory-testingci-cdlocal-first

A
Tool

rLLM

TrainingQA

rllm-org

rLLM bolts RL onto agents you already wrote. verl, trlx, OpenRLHF make you rewrite the agent into their pipeline. A deco

reinforcement-learningagent-trainingsandboxdistributed-training

A
Tool

Agenta

MonitoringQA

Agenta-AI

Workspace for building agents through chat and running them in the background, with swappable harnesses and models

llmopsevaluationobservabilityagent-builder

A
Tool

LLM Space

CodingQA

deer-flow

A desktop app to prototype agent ideas, inspect every harness step, replay failures, and evaluate performance, all in one place. Local-first

desktopobservabilitydebuggingprompt

A
Tool

Sunpeak

QAInterface

Alignbase

Server-agnostic MCP testing framework and full-stack MCP App framework for ChatGPT Apps, Claude Connectors, and more.

chatgpt-appsmcp-uihot-reloadvisual-testing

A
Tool

Future AGI

MonitoringQASecurity

future-agi

OpenTelemetry-based tracing and evaluation instrumentation for voice-AI applications.

telemetryevaluationguardrailsobservability

A
Tool

Argent

QAInterface

software-mansion

An agentic toolkit to control, debug, and profile iOS and Android apps. Made by Software Mansion.

mobile-automationiosandroidmobile

A
Tool

Codegraff

CodingQA

justrach

graff — a fast agentic coding harness in Zig: multi-provider, MCP, workflows, DGM evolution loop, TS/Python SDKs

A