Evaluated on FrontierHarness tasks. Graff includes a reproducible FrontierHarness evaluation runner, with recorded outcomes and explicit protocol differences. See how we measure it for the public task suite, grading, and the limits of comparisons with the published board.

Are you the maintainer?
Claim this page →Codegraff
graff — a fast agentic coding harness in Zig: multi-provider, MCP, workflows, DGM evolution loop, TS/Python SDKs
01 / About
What Codegraff is.
02 / Discussion CREDIBILITY-GATED
Discussion
Reading is open to everyone. Posting and voting need a verified identity or a GitHub grade of B or higher.
- No discussions yet.
03 / Related
More around Codegraff.
Similar tools
ToolAgents CLI
CodingDeploymentQA
The CLI and skills that turn any coding assistant into an expert at creating, evaluating, and deploying AI agents on Google Cloud.
adkgcpcliskill
ToolAutocontext
CodingQA
greyhaven-ai
Agents that start cold every run can't compound what worked. autocontext splits improvement into five roles per cycle …
token-optimizationself-improvementplaybooksevaluation-loop
ToolLooper
CodingQA
ksimback
Claude Code skill for designing review-gated agent loops - goal, plan, review, deliver, judge - then emitting a runnable spec
evaluationyamlworkflow-design
ToolRaindrop Workshop
QAMonitoringCoding
raindrop-ai
Open-source tool that lets a coding agent write and run agent evals locally
observabilityevaluationdebugginglocal-first
ToolSpacedock
CodingQA
clkao
Pi ships four tools ... read, write, edit, bash ... and calls it done. No MCP, no sub-agents, no permission dialogs, no
human-in-the-loopmulti-agentorchestrationgovernance
ToolMultiTown
CodingQA
w1u2d3i4
Code-only runtime toolkit for cost-aware multi-agent organization and control
multi-agentorchestrationsimulationreplay
04 / Build
Build with Codegraff.
Browse the catalogue for frameworks, tools, and harnesses, each scored on real GitHub credibility.
Get Codegraff →