Collection
Code with an agent
Agents and libraries that read, write and review code in your repo.
38 tools · showing 25–38
ToolAgent SOP
Coding
strands-agents
Natural language workflows that enable AI agents to perform complex, multi-step tasks with consistency and reliability.
orchestrationpromptstrands-agentsevaluation
HarnessSWE-agent
CodingSecurity
SWE-agent
Lets a language model autonomously use tools to fix GitHub issues and CTF challenges; superseded by mini-swe-agent
evaluationgithubctf
ToolRaindrop Workshop
QAMonitoringCoding
raindrop-ai
Open-source tool that lets a coding agent write and run agent evals locally
observabilityevaluationdebugginglocal-first
ToolSpacedock
CodingQA
spacedock-dev
Pi ships four tools ... read, write, edit, bash ... and calls it done. No MCP, no sub-agents, no permission dialogs, no
human-in-the-loopmulti-agentorchestrationgovernance
HarnessAIDE ML
CodingQA
WecoAI
LLM-driven agent that writes, evaluates and improves machine-learning code.
clitree-searchautomlevaluation
ToolPandaProbe
CodingMonitoring
chirpz-ai
Open-source platform for tracing, evaluating, and monitoring AI agents, with integrations for LangGraph, CrewAI, and agent SDKs
observabilityevaluationlangchaincrewai
ToolAutocontext
CodingQA
greyhaven-ai
Agents that start cold every run can't compound what worked. autocontext splits improvement into five roles per cycle …
token-optimizationself-improvementplaybooksevaluation-loop
MCP Server Tauri
CodingQA
hypothesi
A Model Context Protocol (MCP) server and plugin for Tauri v2 development
tauricomputer-usedebuggingscreenshot
ToolTerminal-Bench
QACoding
harbor-framework
Benchmark and harness for evaluating AI agents on end-to-end tasks in real terminal environments.
evaluationcliterminal-taskssandbox
ToolEasy Dataset
Coding
ConardLi
Build LLM fine-tuning and evaluation datasets from documents with parsing, chunking, QA generation, and multi-format export
dataset-generationfine-tuneragevaluation
ToolLooper
CodingQA
ksimback
Claude Code skill for designing review-gated agent loops - goal, plan, review, deliver, judge - then emitting a runnable spec
evaluationyamlworkflow-design
Evals
Coding
openai
Framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
evaluationmodel-gradedgpt
Hermes Agent Self-Evolution
CodingQA
NousResearch
Evolves Hermes Agent skills and prompts with DSPy and GEPA, gated by tests and human PR review
dspygepapromptpull-request
ToolMultiTown
CodingQA
w1u2d3i4
Code-only runtime toolkit for cost-aware multi-agent organization and control
multi-agentorchestrationsimulationreplay