Collection
Code with an agent
Agents and libraries that read, write and review code in your repo.
38 tools · showing 1–24
ToolChrome DevTools MCP
CodingInterfaceQA
ChromeDevTools
Chrome DevTools for coding agents
browser-automationmcp-serverchromedebugging
Agent Development Kit (ADK)
CodingQADeployment
Google's open-source SDK for building, evaluating and deploying multi-agent systems.
multi-agentadk-webevaluation
PlatformInsForge
CodingMemoryQA
InsForge
The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting,
backend-as-a-servicepostgrespgvectorauth
Mastra
CodingMemoryQA
mastra-ai
TypeScript framework for AI agents and apps, with model routing, graph-based workflows, memory, and built-in evals
orchestrationragnextjsevaluation
NeMo Agent Toolkit
CodingMonitoringQA
NVIDIA
Library for connecting, profiling and optimizing teams of agents across frameworks.
profilingobservabilityevaluationmulti-agent
Testsprite CLI
InterfaceQA
TestSprite
The verification layer for the agentic coding era. AI ships code in minutes — verifying it hasn't. testsprite opens your live app, uses it like a real user, and shows your coding agent exactly what broke.
playwrightbrowser-automationclicoding-agent
ToolLatitude
Monitoring
latitude-dev
Latitude traces your agent in production, finds the failures, and dispatches your coding agent to fix them.
observabilityissue-detectiontelemetryevaluation
ToolLM Evaluation Harness
Coding
EleutherAI
Framework for evaluating language models across 60+ academic benchmarks through a tokenization-agnostic, multi-backend interface
evaluationhuggingfacevllmtransformers
ToolLangsmith SDK
Coding
langchain-ai
Python and JavaScript SDKs for tracing, evaluating and monitoring LLM apps on the LangSmith platform
observabilityevaluationlangchain
ToolVoltAgent
CodingInterfaceMonitoring
VoltAgent
TypeScript framework and console for building, observing, and operating AI agents
ragobservabilityorchestrationmcp-server
ToolFailproofAI
CodingMonitoringQA
FailproofAI
Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement.
local-firstcloudobservabilityclaude
ToolInspector
CodingDeploymentSecurity
MCPJam
Development and testing platform to debug, chat with, inspect, and run evals against MCP servers, MCP apps, and ChatGPT apps
oauthevaluationdebuggingobservability
Toolmini-SWE-agent
Coding
SWE-agent
The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified
coding-agentevaluationlightweightlitellm
HarnessOuroboros
Coding
Q00
Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop.
coding-agentevaluationspec-drivenhuman-in-the-loop
ToolRemote Factory
Coding
akashgit
Domain-agnostic multi-agent software design and evolution harness
multi-agentorchestrationcoding-agentcli
ToolDeepEval
Coding
confident-ai
Open-source framework for unit-testing and evaluating LLM apps with ready-made metrics that run locally
evaluationragpytest
ToolRaven
CodingMemoryMonitoring
EverMind-AI
The memory-first, self-improving agent harness built on EverOS, with MiroThinker-powered deep research and reasoning.
self-improvementlocal-firstobservabilityskill
ToolLLM Space
CodingQA
deer-flow
A desktop app to prototype agent ideas, inspect every harness step, replay failures, and evaluate performance, all in one place. Local-first
desktopobservabilitydebuggingprompt
SkillOpt
Coding
microsoft
SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, valid
promptskillself-improvementreflection
ToolCodeGraff
CodingQA
justrach
graff — a fast agentic coding harness in Zig: multi-provider, MCP, workflows, DGM evolution loop, TS/Python SDKs
coding-agentmulti-agentgit-worktreeevaluation
ToolAgents CLI
CodingDeploymentQA
The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.
adkgcpcliskill
ToolClaude Cookbooks
Coding
anthropics
Recipe collection of code examples and guides for building with the Claude API, from tool use to evaluations
claudetool-callingragevaluation
ToolGodot MCP
CodingQA
satelliteoflove
Give your AI assistant eyes and hands in the Godot editor: scene editing, input injection, deterministic playtesting, and live game state fo
godotgame-developmentplaytestingscene-editing
ToolHarness Score
CodingQA
paladini
Your AI coding agent is only as reliable as the harness around it. Measure that harness in seconds with harness-score.
clicoding-agentstatic-analysismaturity-model