Frontier model training · Reinforcement learning

I train frontier models.

Across pretraining, mid-training, post-training, and RL—from scaling experiments to full-scale runs.

I am Zhuoyuan Chen (Levin), a research scientist with hands-on experience across the complete model-training lifecycle. At xAI, I helped build the pretraining scaling ladder and worked in the small core group operating and diagnosing large mid-training runs for Grok. At Google, I developed post-training and RL recipes for Gemini, continuing an RL thread that began with distributed AlphaZero systems at Meta FAIR.

Grok 4.5 & Grok 4.6core training contributor
PT → MT → RLfull training lifecycle
AlphaZero → Gemini → Grokcontinuous RL thread
10,000+research citations

What I do

Training judgment
across every stage.

My work centers on the decisions that make large runs succeed: turning controlled experiments into reliable recipes, finding the source of regressions, and balancing new capabilities with general model quality.

01

Pretraining & mid-training

Scaling ladders, data mixtures, curriculum design, run monitoring, trace analysis, failure diagnosis, and large-run decision making.

02

Post-training & RL

SFT and RL recipes, policy optimization, reward-driven learning, distributed self-play, evaluation, and training-dynamics analysis.

03

Capability development

Targeted supervision, controlled ablations, few-shot and perplexity evaluation, regression prevention, and release-quality feedback loops.

Selected experience

Frontier training,
not a single vertical.

I have moved across model generations and organizations while staying close to the same core problem: how learning systems gain capability reliably at scale.

2024–2025Google

Staff Research Engineer & Tech Lead · Vertex AI / Gemini

Gemini: capability development through training

Developed post-training and RL recipes through billion-scale training data, end-to-end model runs, and controlled ablations. Validated improvements entered Gemini core pipelines. I also built SFT, RL, LoRA adaptation, and release-evaluation workflows and contributed across Gemini 1.5, 2.0, 2.5, and Gemini 3.

2017–2019Meta FAIR

Research Engineer · Large-Scale Reinforcement Learning

ELF OpenGo: distributed AlphaZero systems

Built large-scale reinforcement-learning systems spanning distributed self-play, training, inference, and evaluation for ELF OpenGo, an open reproduction and analysis of AlphaZero. Published as an ICML 2019 Oral.

Earlier career

Before frontier foundation models, I led production learning systems at Apple, researched robust learning for safety-critical systems at Uber ATG, and worked on deep learning, RL, and training infrastructure at Baidu Research. These roles built the systems and research foundation I now bring to frontier-scale training.

Selected releases & research

A long arc of
learning systems.

My earlier research spans generative modeling, robust learning, and production machine learning, with publications at ICML, NeurIPS, CVPR, ICCV, and ECCV. See the complete record on Google Scholar ↗.