Anish Adamane

Operator-researcher on AI reliability — building OpenEnv environments that make agent failure modes measurable and falsifiable. Founder of Astrazen, where production voice agents surfaced the role drift problem. Student at Masters Union.

Now

Iterating on Astrazen's voice infrastructure — Pipecat orchestration on Twilio SIP, Deepgram for speech, Groq for inference, with a switchable TTS layer for Indian and global voices. Targeting a ₹7-9 per-minute cost ceiling without making the agent sound like one.

Recent Finding

GRPO on composable drift detectors raised training return to +3.2, but held-out eval falsified the deployment hypothesis: trained checkpoint scored −1.58 vs prompted baseline +0.19 on in-domain scenarios (V10, June 2026). Train/eval disconnect — documented, not hidden.

findings →

Selected Work

Role Drift Environment

OpenEnv-compatible environment for measuring conversational agent drift in deployable voice agents — termination loops, goal abandonment, instruction violations, unprompted language switches. Four failure modes from production transcripts, turned into composable programmatic detectors.

Pre-registered held-out eval with bootstrap confidence intervals, GRPO training recipe against a frozen adversarial customer simulator, and published eval artifacts on Hugging Face. V10 benchmark falsified the deployment win — training return rose while held-out performance fell.

Stack: Python · OpenEnv · Qwen 2.5 · GRPO (TRL) · vLLM · Hugging Face
github.com/GeniusPlums/OpenEnv-Finale →
Astrazen

AI voice agents for high-intent follow-ups in real estate, premium SaaS, education, and D2C — categories where users hesitate, questions shift mid-call, and trust takes time. Started as WhatsApp automation, dropped after the META licensing reality. Pivoted to tech services, then to voice AI after a real estate client conversation made it clear which problem was worth solving.

Iterated through three voice stacks (Vapi → LiveKit → Pipecat) before settling on the current architecture, choosing for cost, latency, and how natural the agent sounds in Hindi and Hinglish.

Stack: Pipecat · Twilio SIP · Deepgram Nova-3 · Groq Llama 3.3 · Sarvam / ElevenLabs Flash · Python · Next.js
astrazen.co →
LeadQualEnv

Text-based reinforcement learning environment for real estate lead qualification, built for Meta and Hugging Face's OpenEnv competition with Palak Dua. Three difficulty levels, deterministic graders, and Pydantic-modeled observation, action, and reward structures.

Two-layer simulator: a deterministic rule engine, with an isolated LLM paraphrase layer on top. Profile signals (problem clarity, red-flag language, scope instability) modeled from real call transcripts.

Stack: Python · Pydantic · OpenAI / Anthropic APIs · Hugging Face Spaces
huggingface.co/spaces/GeniusPlums/scaler_submission →

Background

Founder · Astrazen
Customer discovery, product, and full technical build. Voice infrastructure, agent design, and the full platform around it — campaign management, CRM-lite, transcript review, integrations.
Independent tech services
Two-year run as an independent operator through college. Clients included Asian Paints and Ceigall Industries. Web and automation projects. Pivoted to product after customer pull made it clear voice agents were the next thing to build.
Masters Union
Student. Building Astrazen alongside coursework.

Contact