Research Engineer · RL post-training infrastructure for multimodal and agentic models
Ten years building ML systems end-to-end — from a production video understanding stack at Google DeepMind to RL post-training pipelines at Amazon and Adobe. Most recently I built the full GSPO+DAPO training stack (verl on FSDP, vLLM rollouts, rule-based and LLM-judge verifiers) that took Qwen3-VL-8B-Thinking from 8.0 to 26.3 on a held-out olympiad benchmark in ~360 H200-GPU-hours across three runs.
SFT, DPO, RLHF, GRPO, GSPO, DAPO. Physics-R1 settled on GSPO+DAPO at 8B; the multi-turn follow-up runs decoupled advantages over a two-turn episode.
verl 0.6.1 on FSDP; vLLM for rollout serving; PyTorch, DeepSpeed; TensorFlow (2017–2021). Prior production work at 10B+ parameters.
Rule-based verifiers, LLM judges, unit-consistency checks, component-level reward logging — each term logged separately so it's visible which one is actually driving the gradient.
Multi-stage audit and contamination screening — lexical, embedding, then judge — gating every training corpus; held-out benchmark construction.
Own end-to-end post-training for multi-reference image generation on enterprise foundation models: data curation and SFT pipeline to date, now extending to RLVR/RLAIF. Sole owner of post-training for multi-reference product-detail-page image generation from tech packs, for an enterprise customer.
Built Amazon's large-scale multimodal video search stack; applied RL post-training that reduced hallucination by 28%. Led the GenAI Live Action Studio video-generation infrastructure.
Built and maintained a video object detection training and inference system in TensorFlow for the research team, owning training and serving (through MediaPipe). This infrastructure underpinned AI Choreographer (ICCV 2021, with the public AIST++ dataset) and Attention Bottlenecks for Multimodal Fusion (NeurIPS 2021).
Learning physical parameters from video, advised by Prof. Ming C. Lin.
Built the end-to-end RL post-training pipeline for visual reasoning on olympiad problems: a three-stage data audit pipeline that screened 14,294 candidate records for contamination and produced a 6,432-record training corpus (PhysCorp-A) and a 500-problem held-out benchmark (PhysOlym-A); a reward system combining rule-based verifiers, LLM judges, and unit-consistency checks with component-level logging; and a GSPO+DAPO training recipe on verl/FSDP with vLLM rollouts. Three training runs, ~360 H200-GPU-hours, +18.3 on PhysOlym-A and +15.7 on PhysReason at 8B.
Denser reward for multimodal physics reasoning. Under submission.
6,432-record audited training corpus for visual physics reasoning, cut from a 14,294-record pool by a three-stage contamination audit, with full provenance and license audit.
500-question held-out olympiad benchmark for evaluating visual physics reasoning in VLMs, 99.8% novel-source.
Loaders and tooling for the AIST++ 3D dance dataset (from AI Choreographer, ICCV 2021).
Seven specialist agents orchestrated over LangGraph across a staged research pipeline, dispatched in natural language rather than through a fixed DAG. Agents run as Claude CLI subprocesses with MCP tool access, share state through Prisma/SQLite, hold multi-round discussions logged per turn, and sync progress to Notion.
Research outputs from the systems above.
ECCV 2024
Earlier work, 2012–2021 — physical parameter estimation from observation. Cloth and tissue elasticity recovered from video and single images (ICCV 2017, TOG 2018, TVCG 2016, ICRA 2016, MICCAI 2016). The same question I work on now, before it was tractable in a language model: how does a model learn physics it was never told?
NeurIPS 2023 Workshop SSLTheoryPractice
CVPR 2021 (CV4Animal Workshop)
CARS 2012