OpenEnv-compatible environment for measuring conversational agent drift in deployable voice agents — termination loops, goal abandonment, instruction violations, unprompted language switches. Four failure modes from production transcripts, turned into composable programmatic detectors.
Pre-registered held-out eval with bootstrap confidence intervals, GRPO training recipe against a frozen adversarial customer simulator, and published eval artifacts on Hugging Face. V10 benchmark falsified the deployment win — training return rose while held-out performance fell.