Pretraining & mid-training
Scaling ladders, data mixtures, curriculum design, run monitoring, trace analysis, failure diagnosis, and large-run decision making.
Frontier model training · Reinforcement learning
Across pretraining, mid-training, post-training, and RL—from scaling experiments to full-scale runs.
I am Zhuoyuan Chen (Levin), a research scientist with hands-on experience across the complete model-training lifecycle. At xAI, I helped build the pretraining scaling ladder and worked in the small core group operating and diagnosing large mid-training runs for Grok. At Google, I developed post-training and RL recipes for Gemini, continuing an RL thread that began with distributed AlphaZero systems at Meta FAIR.
What I do
My work centers on the decisions that make large runs succeed: turning controlled experiments into reliable recipes, finding the source of regressions, and balancing new capabilities with general model quality.
Scaling ladders, data mixtures, curriculum design, run monitoring, trace analysis, failure diagnosis, and large-run decision making.
SFT and RL recipes, policy optimization, reward-driven learning, distributed self-play, evaluation, and training-dynamics analysis.
Targeted supervision, controlled ablations, few-shot and perplexity evaluation, regression prevention, and release-quality feedback loops.
Selected experience
I have moved across model generations and organizations while staying close to the same core problem: how learning systems gain capability reliably at scale.
Member of Technical Staff · Foundation Model Training
Core contributor to successive Grok generations, including Grok 4.5 and Grok 4.6, across pretraining, mid-training, SFT/RL, and evaluation. I helped build the pretraining scaling ladder and served in the small core group operating and debugging large mid-training runs— reading traces, diagnosing failures and regressions, and informing curriculum, data-mixture, and recipe decisions.
Staff Research Engineer & Tech Lead · Vertex AI / Gemini
Developed post-training and RL recipes through billion-scale training data, end-to-end model runs, and controlled ablations. Validated improvements entered Gemini core pipelines. I also built SFT, RL, LoRA adaptation, and release-evaluation workflows and contributed across Gemini 1.5, 2.0, 2.5, and Gemini 3.
Research Engineer · Large-Scale Reinforcement Learning
Built large-scale reinforcement-learning systems spanning distributed self-play, training, inference, and evaluation for ELF OpenGo, an open reproduction and analysis of AlphaZero. Published as an ICML 2019 Oral.
Earlier career
Before frontier foundation models, I led production learning systems at Apple, researched robust learning for safety-critical systems at Uber ATG, and worked on deep learning, RL, and training infrastructure at Baidu Research. These roles built the systems and research foundation I now bring to frontier-scale training.
Selected releases & research
My earlier research spans generative modeling, robust learning, and production machine learning, with publications at ICML, NeurIPS, CVPR, ICCV, and ECCV. See the complete record on Google Scholar ↗.