This is an early release preview. You may encounter bugs.

Code with an agent

Agents and libraries that read, write and review code in your repo.

38 tools · showing 25–38

Tool

Agent SOP

Coding

strands-agents

Natural language workflows that enable AI agents to perform complex, multi-step tasks with consistency and reliability.

orchestrationpromptstrands-agentsevaluation

B
Harness

SWE-agent

CodingSecurity

SWE-agent

Lets a language model autonomously use tools to fix GitHub issues and CTF challenges; superseded by mini-swe-agent

evaluationgithubctf

B
Tool

Raindrop Workshop

QAMonitoringCoding

raindrop-ai

Open-source tool that lets a coding agent write and run agent evals locally

observabilityevaluationdebugginglocal-first

B
Tool

Spacedock

CodingQA

spacedock-dev

Pi ships four tools ... read, write, edit, bash ... and calls it done. No MCP, no sub-agents, no permission dialogs, no

human-in-the-loopmulti-agentorchestrationgovernance

B
Harness

AIDE ML

CodingQA

WecoAI

LLM-driven agent that writes, evaluates and improves machine-learning code.

clitree-searchautomlevaluation

B
Tool

PandaProbe

CodingMonitoring

chirpz-ai

Open-source platform for tracing, evaluating, and monitoring AI agents, with integrations for LangGraph, CrewAI, and agent SDKs

observabilityevaluationlangchaincrewai

B
Tool

Autocontext

CodingQA

greyhaven-ai

Agents that start cold every run can't compound what worked. autocontext splits improvement into five roles per cycle …

token-optimizationself-improvementplaybooksevaluation-loop

B
Tool

MCP Server Tauri

CodingQA

hypothesi

A Model Context Protocol (MCP) server and plugin for Tauri v2 development

tauricomputer-usedebuggingscreenshot

B
Tool

Terminal-Bench

QACoding

harbor-framework

Benchmark and harness for evaluating AI agents on end-to-end tasks in real terminal environments.

evaluationcliterminal-taskssandbox

B
Tool

Easy Dataset

Coding

ConardLi

Build LLM fine-tuning and evaluation datasets from documents with parsing, chunking, QA generation, and multi-format export

dataset-generationfine-tuneragevaluation

B
Tool

Looper

CodingQA

ksimback

Claude Code skill for designing review-gated agent loops - goal, plan, review, deliver, judge - then emitting a runnable spec

evaluationyamlworkflow-design

B
Tool

Evals

Coding

openai

Framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

evaluationmodel-gradedgpt

B
Tool

Hermes Agent Self-Evolution

CodingQA

NousResearch

Evolves Hermes Agent skills and prompts with DSPy and GEPA, gated by tests and human PR review

dspygepapromptpull-request

C
Tool

MultiTown

CodingQA

w1u2d3i4

Code-only runtime toolkit for cost-aware multi-agent organization and control

multi-agentorchestrationsimulationreplay

C

More ways in