Problem
Agent failures are difficult to reproduce, measure, and repair without introducing regressions.
An agent that watches other agents fail—and verifies the repair.
An 8-stage autonomous meta-agent that supervises production LLM agents through Phoenix telemetry, detecting hallucinations, prompt drift, and tool-call failures in under 10 seconds via LLM-as-judge.
Interactive system walkthrough · simulated frontend data · not live telemetry
Agent failures are difficult to reproduce, measure, and repair without introducing regressions.
Inspect Phoenix traces, classify failures, generate adversarial tests, patch prompts, compare pass-rate, cost, and latency, then replay the original failure.
Baseline and patched prompts are evaluated against synthesized adversarial cases before a change is accepted.
A plausible patch can improve one example while degrading broader behavior.
Treat prompt changes like code changes: test, compare, replay, and only then resolve.
Failure detection in under 10 seconds; 5 supervision tools exposed through cassandra-mcp.