Frontier foundation models · Pretraining · Post-training · RL

Zhuoyuan Chen (Levin)

I build learning systems that move from research ideas to frontier-scale training and real products.

Research scientist with 15+ years building frontier foundation models and large-scale learning systems. I work across the full model lifecycle—from data and scaling ladders through pretraining, mid-training, SFT, reinforcement learning, evaluation, and deployment. Contributor to Grok, Gemini, and Apple Vision Pro, with particular depth in multimodal learning, generative modeling, and large-scale RL.

15+years in AI research
10,000+Google Scholar citations
PT → RLfull model lifecycle
ICML · NeurIPS · CVPR · ICCV · ECCVresearch track record

Selected work

Models and systems
built across scales.

My work connects frontier-model training, reinforcement learning, generative modeling, and spatial intelligence.

01 Grok

xAI · Frontier Models

Grok

Frontier-model development across scaling ladders, pretraining, mid-training, post-training, evaluation, large-scale data, and successive Grok releases.

PretrainingPost-trainingEvaluation
02 Gemini

Google · Foundation Models

Gemini

Billion-scale data construction, end-to-end training and controlled ablations, model adaptation, and release evaluation integrated with Gemini core training pipelines.

SFT / RLDataModel Adaptation
03 ELF OpenGo

Meta FAIR · Reinforcement Learning

ELF OpenGo

A large-scale open reproduction and analysis of AlphaZero, spanning distributed self-play, training, inference, and evaluation. ICML 2019 Oral.

AlphaZeroDistributed RLICML Oral
04 VoxelCNN

Meta FAIR · Generative Modeling

VoxelCNN

Order-aware autoregressive generation for structured 3D worlds—an early bridge between generative modeling and spatial intelligence.

Autoregressive3DICCV
05 Vision Pro

Apple · Spatial Intelligence

Vision Pro

Led production 3D perception with voxelized RGB-D inputs, temporal fusion, detection, and scene reconstruction deployed in Apple Vision Pro.

3D PerceptionRGB-DProduct

Experience

Research, training,
and production.

xAI

xAI

Member of Technical Staff

Omni / Multimodal Foundation Models
Google logo

Google

Staff Research Engineer & Tech Lead

Vertex AI / Gemini
Apple logo

Apple

Staff Research Scientist & Tech Lead

Applied Foundation Models
Uber logo

Uber ATG

Senior Research Scientist

Autonomous Driving Research
Meta logo

Meta FAIR

Research Engineer

Reinforcement Learning & Generative Models
Baidu logo

Baidu Research

Senior Research Scientist

Institute of Deep Learning

Education

Northwestern UniversityPh.D., Electrical Engineering & Computer Science · 2014
Tsinghua UniversityM.S., Computer Science · B.S., Mathematics & Physics

Publications & technical releases

A long arc of
learning systems.

Selected model releases, conference papers, technical reports, and foundational work from 2009–2026. For citation details, see Google Scholar ↗.

  1. 012026Introducing Grok 4.5xAI Model Release
  2. 022026Grok 4.3xAI Flagship Model Release
  3. 032026Grok 4.20 System CardxAI Technical Report
  4. 042025Gemini 3: Introducing the Latest Gemini AI Model from GoogleGoogle Model Release
  5. 052025Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next-Generation Agentic CapabilitiesTechnical Report
  6. 062024Gemini 2.0: Our New AI Model for the Agentic EraGoogle Model Release
  7. 072024Gemini 1.5: Unlocking Multimodal Understanding Across Millions of Tokens of ContextTechnical Report
  8. 082023A Bitter Lesson: Is MAE Really a Good Generator?arXiv
  9. 092022GAUDI: A Neural Architect for Immersive 3D Scene GenerationNeurIPS
  10. 102021ARKitScenes: A Diverse Real-World Dataset for 3D Indoor Scene Understanding Using Mobile RGB-D DataNeurIPS
  11. 112020ShapeAdv: Generating Shape-Aware Adversarial 3D Point CloudsECCV
  12. 122019Order-Aware Generative Modeling Using the 3D-Craft DatasetICCV
  13. 132019CraftAssist: A Framework for Dialogue-Enabled Interactive AgentsarXiv
  14. 142019Why Build an Assistant in Minecraft?arXiv
  15. 152019ELF OpenGo: An Analysis and Open Reimplementation of AlphaZeroICML Oral
  16. 162017An Analysis of Feature Regularization for Low-Shot LearningTechnical Report
  17. 172016Mining Spatial and Spatio-Temporal ROIs for Video Action RecognitionECCV Workshop
  18. 182016Integrated Variational and Nearest Neighbor Field for Optical FlowECCV
  19. 192015A Deep Visual Correspondence Embedding Model for Stereo Matching CostsICCV
  20. 202013Robust Dictionary Learning by Error Source DecompositionICCV
  21. 212013Caranx: Scalable Social Image Index Using a Phylogenetic Tree of HashtagsSC
  22. 222013Large Displacement Optical Flow from Nearest Neighbor FieldsCVPR
  23. 232012Robust 3D Action Recognition with Random Occupancy PatternsECCV
  24. 242012Decomposing and Regularizing Sparse/Non-Sparse Components for Motion Field EstimationCVPR
  25. 252012Spatial Locality-Aware Sparse Coding and Dictionary LearningJMLR
  26. 262011Action Recognition with Multiscale Spatio-Temporal ContextsCVPR
  27. 272011A Compressive Sensing Reconstruction Algorithm for Trinary and Binary Sparse Signals Using Pre-MappingDCC Oral
  28. 282010Reconstruction of Sparse Binary Signals Using Compressive SensingDCC
  29. 292010Blind Motion Deblurring Using Local MinimumInverse Problems
  30. 302009Auto-Cut for Web ImagesACM Multimedia