Shan Yang

Shan Yang

Research Engineer · RL post-training infrastructure for multimodal and agentic models

Ten years building ML systems end-to-end — from a production video understanding stack at Google DeepMind to RL post-training pipelines at Amazon and Adobe. Most recently I built the full GSPO+DAPO training stack (verl on FSDP, vLLM rollouts, rule-based and LLM-judge verifiers) that took Qwen3-VL-8B-Thinking from 8.0 to 26.3 on a held-out olympiad benchmark in ~360 H200-GPU-hours across three runs.

Technical Stack

Post-training algorithms

SFT, DPO, RLHF, GRPO, GSPO, DAPO. Physics-R1 settled on GSPO+DAPO at 8B; the multi-turn follow-up runs decoupled advantages over a two-turn episode.

Training & rollout infrastructure

verl 0.6.1 on FSDP; vLLM for rollout serving; PyTorch, DeepSpeed; TensorFlow (2017–2021). Prior production work at 10B+ parameters.

Reward systems

Rule-based verifiers, LLM judges, unit-consistency checks, component-level reward logging — each term logged separately so it's visible which one is actually driving the gradient.

Data pipelines

Multi-stage audit and contamination screening — lexical, embedding, then judge — gating every training corpus; held-out benchmark construction.

Experience

Adobe Foundry — Staff Applied Scientist

2026–now

Own end-to-end post-training for multi-reference image generation on enterprise foundation models: data curation and SFT pipeline to date, now extending to RLVR/RLAIF. Sole owner of post-training for multi-reference product-detail-page image generation from tech packs, for an enterprise customer.

Amazon — Senior Applied Scientist, Tech Lead

2021–2026

Built Amazon's large-scale multimodal video search stack; applied RL post-training that reduced hallucination by 28%. Led the GenAI Live Action Studio video-generation infrastructure.

Google DeepMind — Senior Research Software Engineer

2017–2021

Built and maintained a video object detection training and inference system in TensorFlow for the research team, owning training and serving (through MediaPipe). This infrastructure underpinned AI Choreographer (ICCV 2021, with the public AIST++ dataset) and Attention Bottlenecks for Multimodal Fusion (NeurIPS 2021).

UNC-Chapel Hill — PhD, Computer Science

2012–2017

Learning physical parameters from video, advised by Prof. Ming C. Lin.

Selected Project
Physics-R1

Physics-R1

Built the end-to-end RL post-training pipeline for visual reasoning on olympiad problems: a three-stage data audit pipeline that screened 14,294 candidate records for contamination and produced a 6,432-record training corpus (PhysCorp-A) and a 500-problem held-out benchmark (PhysOlym-A); a reward system combining rule-based verifiers, LLM judges, and unit-consistency checks with component-level logging; and a GSPO+DAPO training recipe on verl/FSDP with vLLM rollouts. Three training runs, ~360 H200-GPU-hours, +18.3 on PhysOlym-A and +15.7 on PhysReason at 8B.

In Progress

Physics-R2 — multi-turn RL

Denser reward for multimodal physics reasoning. Under submission.

Open Source & Releases

Code · GitHub
physics-r1-code

Training, evaluation, and reward code for Physics-R1.

Dataset · Hugging Face
PhysCorp-A

6,432-record audited training corpus for visual physics reasoning, cut from a 14,294-record pool by a three-stage contamination audit, with full provenance and license audit.

Benchmark · Hugging Face
PhysOlym-A

500-question held-out olympiad benchmark for evaluating visual physics reasoning in VLMs, 99.8% novel-source.

Code · GitHub
AIST++ Dataset API

Loaders and tooling for the AIST++ 3D dance dataset (from AI Choreographer, ICCV 2021).

Agentic System
Lumi Research Manager

Seven specialist agents orchestrated over LangGraph across a staged research pipeline, dispatched in natural language rather than through a fixed DAG. Agents run as Claude CLI subprocesses with MCP tool access, share state through Prisma/SQLite, hold multi-round discussions logged per turn, and sync progress to Notion.

Selected Publications

Research outputs from the systems above.

TTC-Net

Beyond Test-Time Training: Learning to Reason via Hardware-Efficient Optimal Control

Peihao Wang, Shan Yang, Xijun Wang, Tesi Xiao, Xin Liu, Changlong Yu, Yu Lou, Pan Li, Zhangyang Wang, Ming Lin, René Vidal

ICML 2026

VLAP

VLAP: Efficient Video-Language Alignment via Frame Prompting and Distilling for Video Question Answering

Xijun Wang, Junbang Liang, Chun-Kai Wang, Kenan Deng, Yu Lou, Ming Lin, Shan Yang

ECCV 2024

MBT

Attention Bottlenecks for Multimodal Fusion

Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, Chen Sun

NeurIPS 2021

AI Choreographer

AI Choreographer: Music Conditioned 3D Dance Generation with AIST++

Shan Yang*, Ruilong Li*, David A. Ross, Angjoo Kanazawa

ICCV 2021

Show all publications

Earlier work, 2012–2021 — physical parameter estimation from observation. Cloth and tissue elasticity recovered from video and single images (ICCV 2017, TOG 2018, TVCG 2016, ICRA 2016, MICCAI 2016). The same question I work on now, before it was tractable in a language model: how does a model learn physics it was never told?

ICAR

ICAR: Image-based Complementary Auto Reasoning

Xijun Wang, Anqi Liang, Junbang Liang, Ming Lin, Yu Lou, Shan Yang

AAAI 2024

Cloth Material Recovery

Learning-based Cloth Material Recovery from Video

Shan Yang, Junbang Liang, Ming C. Lin

ICCV 2017

MeSa

MeSa: Masked, Geometric, and Supervised Pre-training for Monocular Depth Estimation

Muhammad Osama Khan, Junbang Liang, Chun-Kai Wang, Shan Yang, Yu Lou

NeurIPS 2023 Workshop SSLTheoryPractice

Schema Perception

Schema Perception for Robust Video Question Answering

Xijun Wang, Shan Yang

Adobe Technical Report 2025

RoSI

RoSI: Recovering 3D Shape Interiors from Few Articulation Images

Akshay Gadi Patil, Yiming Qian, Shan Yang, Brian Jackson, Eric Bennett, Hao Zhang

arXiv 2023

Optical Mouse

Optical Mouse: 3D Mouse Pose From Single-View Video

Shan Yang*, Bo Hu*, David A. Ross, Avneesh Sud, Yi Liu, Graham Ruby, Bryan Seybold

CVPR 2021 (CV4Animal Workshop)

Garment Recovery

Physics-Inspired Garment Recovery from a Single-View Image

Shan Yang, Zherong Pan, Tanya Amert, Ke Wang, Licheng Yu, Tamara Berg, Ming C. Lin

ACM TOG 2018

Referring Expressions

Modeling Context in Referring Expressions

Licheng Yu, Patric Poirson, Shan Yang, Alex Berg, Tamara Berg

ECCV 2016

Prostate Cancer Classification

Classification of Prostate Cancer Grades and T-Stages based on Tissue Elasticity Using Medical Image Analysis

Shan Yang, Vladimir Jojic, Jun Lian, Ronald Chen, Hongtu Zhu, Ming C. Lin

MICCAI 2016

Bayesian Estimation

Bayesian Estimation of Non-Rigid Mechanical Parameters Using Temporal Sequences of Deformation Samples

Shan Yang, Ming C. Lin

ICRA 2016

MaterialCloning

MaterialCloning: Acquiring Elasticity Parameters from Images for Medical Applications

Shan Yang, Ming C. Lin

IEEE TVCG 2016

Simultaneous Estimation

Simultaneous Estimation of Elasticity for Multiple Deformable Bodies

Shan Yang, Ming C. Lin

Computer Animation and Virtual Worlds 2015

Buried Suture

Real-time Simulation for Buried Suture

Shan Yang, Wenlong Lu, Lixu Gu

CARS 2012