Sirjan Singh
CASE STUDY / OpenEnv Multi-Agent RL

SurviveCity

Reward design is part of the system—not a score added afterward.

PyTorchTRLGRPOLoRAQwen2.5-3B

An OpenEnv-compliant multi-agent survival environment with hidden roles, 14 actions, and 15 reward rubrics. Qwen2.5-3B was fine-tuned with GRPO and LoRA on a V100 DGX cluster.

WHAT THIS PROVES

Multi-agent RL / Reward-hack analysis / Post-training / Custom rollout

LAB_02 / REWARD_HACKINGEDUCATIONAL SIMULATION
REWARD63
SURVIVAL73
BEHAVIORBALANCED POLICY

Interactive system walkthrough · simulated frontend data · not live telemetry

01THE PRESSURE

Problem

Agents optimized the rubric in ways that failed the actual survival objective.

02THE SYSTEM

Approach

Build a custom multi-action rollout and patch specific reward rubrics after reproducing exploits.

03THE MEASURE

Experiments

Tracked survival, actions, hidden roles, and 15 rubrics through GRPO + LoRA training.

04THE BREAK

Failure mode

The apparent policy-collapse signal was actually starvation.

05THE JUDGMENT

Engineering decision

Debug the environment and reward signals before blaming the policy.

06THE EVIDENCE

Measured result

Agent survival rate went from 15% to 60%, four times higher; 3 reward-hacking exploits were closed.

Return to all selected workVERIFY THE CLAIMS. INSPECT THE WORK.