Researcher · RL Post-Training & Agentic Reasoning
I work on RL post-training — models that reason and use tools reliably. My focus is the unglamorous half of the stack: audited training data, hard evaluation benchmarks, reward design, and RL recipes that survive real-world deployment. Physics reasoning is where I proved it out (Physics-R1); I'm now extending it to agentic tool-use and reward hacking.
Currently Staff Applied Scientist at Adobe Foundry, building custom generative foundation models for enterprise customers (SFT, DPO, RL, with data curation as the main lever). Previously Tech Lead at Amazon Video Search (multimodal search + RL post-training, shipped to customers) and GenAI Live Action Studio, and Senior Research SDE at Google Research (multi-modal modeling, AIST++). PhD from UNC-Chapel Hill with Prof. Ming C. Lin on learning physical parameters from video.
Currently exploring: tool-use RL and reward hacking in agents (Physics-R2, in preparation).
Visual physics reasoning for vision-language models: a 2,434-record audited training corpus, a 500-question novel-source olympiad benchmark (PhysOlym-A), and an RL recipe that pushes SOTA on visual physics reasoning at the 7B scale. All artifacts open.
ECCV 2024
NeurIPS 2023 Workshop SSLTheoryPractice
CVPR 2021 (CV4Animal Workshop)
CARS 2012
2,434-record audited training corpus for visual physics reasoning, with provenance and license audit.
500-question novel-source olympiad benchmark for evaluating visual physics reasoning in VLMs.
11.6K episodes of physics-grounded video annotations for world-model training.
Loaders and tooling for the AIST++ 3D dance dataset (from AI Choreographer, ICCV 2021).
AI-powered research project manager with pixel-art agent teams (Scout, Theorist, Architect, Coder).
Personal log of learning reinforcement learning — notes, experiments, and insights.