Problem
Agents optimized the rubric in ways that failed the actual survival objective.
Reward design is part of the system—not a score added afterward.
An OpenEnv-compliant multi-agent survival environment with hidden roles, 14 actions, and 15 reward rubrics. Qwen2.5-3B was fine-tuned with GRPO and LoRA on a V100 DGX cluster.
Interactive system walkthrough · simulated frontend data · not live telemetry
Agents optimized the rubric in ways that failed the actual survival objective.
Build a custom multi-action rollout and patch specific reward rubrics after reproducing exploits.
Tracked survival, actions, hidden roles, and 15 rubrics through GRPO + LoRA training.
The apparent policy-collapse signal was actually starvation.
Debug the environment and reward signals before blaming the policy.
Agent survival rate went from 15% to 60%, four times higher; 3 reward-hacking exploits were closed.