Sirjan Singh
CASE STUDY / Autonomous LLM Agent Supervision

Cassandra

An agent that watches other agents fail—and verifies the repair.

PythonGoogle ADKGeminiArize Phoenix MCP

An 8-stage autonomous meta-agent that supervises production LLM agents through Phoenix telemetry, detecting hallucinations, prompt drift, and tool-call failures in under 10 seconds via LLM-as-judge.

WHAT THIS PROVES

Evaluation / Observability / Regression-safe patching / MCP tool authoring

SUPERVISION_TRACEINTERACTIVE
NODE_1 / AGENT

Interactive system walkthrough · simulated frontend data · not live telemetry

01THE PRESSURE

Problem

Agent failures are difficult to reproduce, measure, and repair without introducing regressions.

02THE SYSTEM

Approach

Inspect Phoenix traces, classify failures, generate adversarial tests, patch prompts, compare pass-rate, cost, and latency, then replay the original failure.

03THE MEASURE

Experiments

Baseline and patched prompts are evaluated against synthesized adversarial cases before a change is accepted.

04THE BREAK

Failure mode

A plausible patch can improve one example while degrading broader behavior.

05THE JUDGMENT

Engineering decision

Treat prompt changes like code changes: test, compare, replay, and only then resolve.

06THE EVIDENCE

Measured result

Failure detection in under 10 seconds; 5 supervision tools exposed through cassandra-mcp.

Return to all selected workVERIFY THE CLAIMS. INSPECT THE WORK.