I work on generalizable robot learning: how can available experience become behavior that works in a new situation? My research connects reinforcement learning, visually grounded action, continuous execution, and human data for dexterous manipulation.
I am a master’s student in Computer Science and Technology at Tongji University, advised by Prof. Junqiao Zhao. My experience spans language-model alignment at TAL and reinforcement learning with Prof. Eduardo Veas during an Erasmus+ exchange at TU Graz.
At Spirit-AI, I worked with Junliang Guo, Junyuan Xie, and Yang Gao on gripper policy pretraining. At Liber-AI, I work with Fanqi Lin and Songming Liu on human-data-driven dexterous learning.
I will join the School of Computing and Data Science, The University of Hong Kong, as a PhD student in Fall 2027, advised by Prof. Hongyang Li.
Generalist manipulation from diverse robot experience.
Robot-free glove collection for rich hand–object interactions.
Coordinating hands, arms, torso, and locomotion.
For my PhD, I hope to address mobility in confined spaces, tasks requiring whole-body force, combined locomotion and manipulation in complex environments, and tasks requiring body contact.
Language alone can leave manipulation targets ambiguous. Point-VLA combines a visual target cue with a language action, using a fixed reference image, multimodal co-training, and spatial augmentation.
92.5% mean success across six reported tasks, versus 32.4% for text-only instructions and 40.0% for interleaved prompting (up to two retries within 30 seconds).
Role: Independently developed the method and carried out the experiments.
Action chunks should continue smoothly without a training–inference mismatch. Legato couples scheduled initialization with a consistent velocity field to learn native continuation.
In the reported pouring comparison, completion time decreased from 95.07 to 75.73 seconds while task score increased from 9.34 to 9.72.
Role: Theory and derivations, the training–inference-consistent velocity-field formulation, and experimental implementation.
Fragmented offline trajectories leave useful experience disconnected. ASTRO combines temporal target selection, a diffusion action planner with frozen dynamics, and deviation feedback for self-correction during training and rollout.
In the reported state-based offline RL evaluation, IQL improved from 36.08 to 45.52 and FQL from 55.52 to 65.71.
Role: Independently developed the method and carried out the experiments.
What improves when an LLM reasons better? Knowledge Index and Information Gain disentangle knowledge from reasoning, revealing different effects of supervised fine-tuning and reinforcement learning in the evaluated models.
Visual RL can overfit to irrelevant backgrounds. SMG separates task-relevant foreground learning using reward and inverse-dynamics signals, improving generalization under visual distractions.
Better predictive models do not always produce better policies. USB-PO jointly accounts for model shift and model bias with an adaptive model-update objective, improving sample efficiency on six MuJoCo tasks.
Aug. 2026 – Present
Core member working with Fanqi Lin and Songming Liu. I study scaling laws for pretraining both World Action Models (WAM) and Vision-Language-Action (VLA) models, focusing on learning dexterous manipulation from human videos and data.
Robot-free glove collection captures direct human interactions efficiently and covers dexterous tasks that are difficult to perform through teleoperation. The goal is to turn this richer human experience into capable robotic-hand policies. “Robot-free” describes data collection, not a claim of glove-only model training.
Company demonstrations: dexterous manipulation with glove-trained policies →
Aug. 2025 – Aug. 2026
Worked with Prof. Yang Gao and Dr. Junyuan Xie. I participated throughout the VLA pretraining pipeline, including benchmark evaluation and post-training demonstrations.
RoboChallenge · #1 / 50.33%
Spirit v1.5, Table 30, reported Jan. 11, 2026. Task-specific fine-tuning across 30 tasks and four robot platforms. Report & demonstrations
RoboArena · #1 / 1918 rating
Spirit v1.6, June 2, 2026 snapshot with 301 A/B comparisons. Company report
These are historical team-level benchmark results under different evaluation protocols, not claims of current leaderboard positions.
Mar. – Jul. 2025 · Graz, Austria
Worked with Prof. Eduardo Veas on reinforcement learning and dynamics-guided trajectory stitching, connecting fragmented offline experience into useful training trajectories.
Mar. – Jul. 2023
Contributed to the RLHF pipeline for MathGPT, including reward modeling and PPO/DPO alignment for mathematical reasoning.
Class monitor and team captain for the National Undergraduate Intelligent Vehicle Competition, with experience coordinating teammates and communicating technical work.
Invited Session Chair and Co-Chair, IROS 2026.