This is an early release preview. You may encounter bugs.

Code with an agent

Agents and libraries that read, write and review code in your repo.

38 tools · showing 1–24

Tool

Chrome DevTools MCP

CodingInterfaceQA

ChromeDevTools

Chrome DevTools for coding agents

browser-automationmcp-serverchromedebugging

A+
Framework

Agent Development Kit (ADK)

CodingQADeployment

google

Google's open-source SDK for building, evaluating and deploying multi-agent systems.

multi-agentadk-webevaluation

A+
Platform

InsForge

CodingMemoryQA

InsForge

The all-in-one, open-source backend platform for agentic coding. InsForge gives your coding agent database, auth, storage, compute, hosting,

backend-as-a-servicepostgrespgvectorauth

A+
Framework

Mastra

CodingMemoryQA

mastra-ai

TypeScript framework for AI agents and apps, with model routing, graph-based workflows, memory, and built-in evals

orchestrationragnextjsevaluation

A+
Framework

NeMo Agent Toolkit

CodingMonitoringQA

NVIDIA

Library for connecting, profiling and optimizing teams of agents across frameworks.

profilingobservabilityevaluationmulti-agent

A+
Tool

Testsprite CLI

InterfaceQA

TestSprite

The verification layer for the agentic coding era. AI ships code in minutes — verifying it hasn't. testsprite opens your live app, uses it like a real user, and shows your coding agent exactly what broke.

playwrightbrowser-automationclicoding-agent

A+
Tool

Latitude

Monitoring

latitude-dev

Latitude traces your agent in production, finds the failures, and dispatches your coding agent to fix them.

observabilityissue-detectiontelemetryevaluation

A+
Tool

LM Evaluation Harness

Coding

EleutherAI

Framework for evaluating language models across 60+ academic benchmarks through a tokenization-agnostic, multi-backend interface

evaluationhuggingfacevllmtransformers

A+
Tool

Langsmith SDK

Coding

langchain-ai

Python and JavaScript SDKs for tracing, evaluating and monitoring LLM apps on the LangSmith platform

observabilityevaluationlangchain

A
Tool

VoltAgent

CodingInterfaceMonitoring

VoltAgent

TypeScript framework and console for building, observing, and operating AI agents

ragobservabilityorchestrationmcp-server

A
Tool

FailproofAI

CodingMonitoringQA

FailproofAI

Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement.

local-firstcloudobservabilityclaude

A
Tool

Inspector

CodingDeploymentSecurity

MCPJam

Development and testing platform to debug, chat with, inspect, and run evals against MCP servers, MCP apps, and ChatGPT apps

oauthevaluationdebuggingobservability

A
Tool

mini-SWE-agent

Coding

SWE-agent

The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified

coding-agentevaluationlightweightlitellm

A
Harness

Ouroboros

Coding

Q00

Agent OS: the agent gets smarter on its own. We just hold the line: the grading command and expected result never make it into the success contract we hand it. Interview-gated, staged evaluation, budgeted evolution loop.

coding-agentevaluationspec-drivenhuman-in-the-loop

A
Tool

Remote Factory

Coding

akashgit

Domain-agnostic multi-agent software design and evolution harness

multi-agentorchestrationcoding-agentcli

A
Tool

DeepEval

Coding

confident-ai

Open-source framework for unit-testing and evaluating LLM apps with ready-made metrics that run locally

evaluationragpytest

A
Tool

Raven

CodingMemoryMonitoring

EverMind-AI

The memory-first, self-improving agent harness built on EverOS, with MiroThinker-powered deep research and reasoning.

self-improvementlocal-firstobservabilityskill

A
Tool

LLM Space

CodingQA

deer-flow

A desktop app to prototype agent ideas, inspect every harness step, replay failures, and evaluate performance, all in one place. Local-first

desktopobservabilitydebuggingprompt

A
Tool

SkillOpt

Coding

microsoft

SkillOpt is a text-space optimizer that trains reusable natural-language skills for frozen LLM agents through trajectory-driven edits, valid

promptskillself-improvementreflection

A
Tool

CodeGraff

CodingQA

justrach

graff — a fast agentic coding harness in Zig: multi-provider, MCP, workflows, DGM evolution loop, TS/Python SDKs

coding-agentmulti-agentgit-worktreeevaluation

A
Tool

Agents CLI

CodingDeploymentQA

google

The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.

adkgcpcliskill

B
Tool

Claude Cookbooks

Coding

anthropics

Recipe collection of code examples and guides for building with the Claude API, from tool use to evaluations

claudetool-callingragevaluation

B
Tool

Godot MCP

CodingQA

satelliteoflove

Give your AI assistant eyes and hands in the Godot editor: scene editing, input injection, deterministic playtesting, and live game state fo

godotgame-developmentplaytestingscene-editing

B
Tool

Harness Score

CodingQA

paladini

Your AI coding agent is only as reliable as the harness around it. Measure that harness in seconds with harness-score.

clicoding-agentstatic-analysismaturity-model

B

More ways in