Sirjan Singh

Projects

Built, broken,
rebuilt.

Agents, evaluation, retrieval, reinforcement learning and developer tools. Each one written up the same way: what I figured out, how it broke, and what I built instead.

6 of 6
traceplantool.calltool.resultanswerflagged by judgepatch+ check tool argsreplaypass

01 Autonomous LLM Agent Supervision

Cassandra

An agent that watches other agents fail—and verifies the repair.

Figured out
Agent failures are difficult to reproduce, measure, and repair without introducing regressions.
Broke
A plausible patch can improve one example while degrading broader behavior.
Built better
Treat prompt changes like code changes: test, compare, replay, and only then resolve.

OutcomeFailure detection in under 10 seconds; 5 supervision tools exposed through cassandra-mcp.

model offersafter guardpublished entitlement7-line guard

02 Agent-Transactable Commerce

Vendable

Sometimes reliability means removing the model from the critical path.

Figured out
Prompt rules can be persuaded away under persistent negotiation, even when the business rule is non-negotiable.
Broke
Pure buyer persistence defeated prompt-only rules.
Built better
Move critical invariants out of probabilistic inference and into deterministic enforcement.

OutcomeAcross 105 recorded negotiations, a 7-line statutory guard held every median on the published entitlement.

modelsearchread_filewrite_drafthuman approvalpublishobservability: every step logged

03 Agent Orchestration & Observability

AgentOS

Agent execution should be inspectable while it is happening.

Figured out
Multi-step agents hide consequential tool calls and make cost or failure difficult to attribute.
Broke
A run-level success label can hide which skill or tool caused cost growth or failure.
Built better
Make observability granular to skills and tools, not only the final response.

OutcomeA real-time interface for agent runs, approvals, skill costs, and outcomes.

Reward collected0.0what the policy optimises
Agents alive4 / 4what we actually wanted
tick 0 / 240

A toy greedy policy, not the trained Qwen2.5-3B agent. It reproduces the failure SurviveCity documented: reward can climb while agents starve.

04 OpenEnv Multi-Agent RL

SurviveCity

Reward design is part of the system—not a score added afterward.

Figured out
Agents optimized the rubric in ways that failed the actual survival objective.
Broke
The apparent policy-collapse signal was actually starvation.
Built better
Debug the environment and reward signals before blaming the policy.

OutcomeAgent survival rate went from 15% to 60%, four times higher; 3 reward-hacking exploits were closed.

BM25densemergetop 1reranked

05 Repo-Aware RAG Assistant

Contextual

Retrieval quality is an evaluation problem before it is a generation problem.

Figured out
Semantic similarity alone misses exact identifiers and repository-specific language.
Broke
Dense retrieval underweighted exact code symbols and project vocabulary.
Built better
Use complementary retrievers and evaluate retrieval independently from generation.

Outcome30% retrieval F1 improvement versus a semantic-only baseline.

observed failuresrecurrence thresholdone scoped rulereplayed against the original task

06 Agentic Development Toolkit

claude-kit

Guardrails should come from observed failure patterns, not intuition alone.

Figured out
Agent coding rules often accumulate as untested opinions rather than responses to measured failure.
Broke
Overbroad rules can make agents less effective while appearing safer.
Built better
Prefer small, evidence-backed constraints over sprawling instruction sets.

OutcomeReusable agent tooling for other developers, released under MIT.