This is an early release preview. You may encounter bugs.
Unclaimed

Platform qa monitoring

MutagenT

Agent-engineering platform: offline build-and-eval loops, trace diagnosis and automated optimization; official Claude Code skills at github.com/mutagent-io.

No votes yet

01 / About

What MutagenT is.

Mutagent is a platform for working on AI agents with other agents: it plans, builds, evaluates, diagnoses, and improves an agent across its development lifecycle. Rather than observing and scoring like a telemetry product, it acts on what observability shows, through a set of agents named Spec, Build, Diagnose, Benchmark, Optimize, Validate, Deploy, and Watch that can be run individually or as an end-to-end loop.

The build path turns your existing context into a buildable specification, scaffolds the agent together with its evaluation suite, and produces a signed spec for review before anything is generated. Evaluation datasets and evals are derived from real production traces, and an LLM-as-a-judge is calibrated against your own experts so scores stay anchored; coverage grows as new failures are found.

Diagnosis reads production traces, root-causes failures, and turns each one into a fix, an eval, or a guardrail, raised as a pull request for review. The optimisation loop takes a goal — for example a target accuracy for a support agent — and runs rounds of diagnose, fix, verify, and re-score against the eval suite, with each change approved before it ships.

Mutagent connects to tracing sources including Datadog and raw JSONL, and generates for frameworks such as the Vercel AI SDK, LangChain, Mastra, and DeepAgents. It runs inside the coding agent you already use, on your own machine, so traces, prompts, and datasets stay on your infrastructure. A published case study covers a cost teardown of Clera's outreach agent and the change that reversed the trend.

Features

  • Spec and build: a buildable spec is generated from your context, scaffolded together with an eval suite, and signed off before generation
  • Trace-derived evals: datasets and evals come from production traces, with an LLM judge calibrated against your experts
  • Failure diagnosis: production traces are root-caused automatically and each failure becomes a fix, an eval, or a guardrail
  • Pull-request output: changes arrive as reviewable pull requests rather than opaque edits
  • Goal-based optimisation: rounds of diagnose, fix, verify, and re-score run against the eval suite until the goal is met
  • Tracing sources: connectors for Datadog, raw JSONL, and other providers
  • Framework targets: generation for the Vercel AI SDK, LangChain, Mastra, DeepAgents, and other stacks
  • Local execution: the agents run in your existing coding agent and your data stays on your infrastructure

02 / Discussion CREDIBILITY-GATED

Discussion

Reading is open to everyone. Posting and voting need a verified identity or a GitHub grade of B or higher.

  • No discussions yet.

04 / Build

Build with MutagenT.

Browse the catalogue for frameworks, tools, and harnesses, each scored on real GitHub credibility.

Get MutagenT →

Browse the catalogue