Hang Yu

Hang Yu

Master's Student
Tongji University
yuhang.harry@gmail.com


About Me

I work on generalizable robot learning: how can available experience become behavior that works in a new situation? My research connects reinforcement learning, visually grounded action, continuous execution, and human data for dexterous manipulation.

I am a master’s student in Computer Science and Technology at Tongji University, advised by Prof. Junqiao Zhao. My experience spans language-model alignment at TAL and reinforcement learning with Prof. Eduardo Veas during an Erasmus+ exchange at TU Graz.

At Spirit-AI, I worked with Junliang Guo, Junyuan Xie, and Yang Gao on gripper policy pretraining. At Liber-AI, I work with Fanqi Lin and Songming Liu on human-data-driven dexterous learning.

I will join the School of Computing and Data Science, The University of Hong Kong, as a PhD student in Fall 2027, advised by Prof. Hongyang Li.

Research Direction

01 · Spirit-AIGripper policy pretraining

Generalist manipulation from diverse robot experience.

02 · Liber-AIDexterous learning from human data

Robot-free glove collection for rich hand–object interactions.

03 · FutureWhole-Body Intelligence

Coordinating hands, arms, torso, and locomotion.

For my PhD, I hope to address mobility in confined spaces, tasks requiring whole-body force, combined locomotion and manipulation in complex environments, and tasks requiring body contact.

Publications

  1. Point-VLA overview figure
    Point-VLA · First author · IROS 2026
    Hang Yu, Juntu Zhao, Yufeng Liu, Kaiyu Li, Cheng Ma, Di Zhang, Yingdong Hu, Guang Chen, Junyuan Xie, Junliang Guo‡, Junqiao Zhao†, Yang Gao†

    Language alone can leave manipulation targets ambiguous. Point-VLA combines a visual target cue with a language action, using a fixed reference image, multimodal co-training, and spatial augmentation.

    92.5% mean success across six reported tasks, versus 32.4% for text-only instructions and 40.0% for interleaved prompting (up to two retries within 30 seconds).

    Role: Independently developed the method and carried out the experiments.

  2. Legato overview figure
    Legato · Core author · RSS 2026
    Yufeng Liu, Hang Yu, Juntu Zhao, Bocheng Li, Di Zhang, Mingzhu Li, Wenxuan Wu, Yingdong Hu, Junyuan Xie, Junliang Guo‡, Dequan Wang†, Yang Gao†

    Action chunks should continue smoothly without a training–inference mismatch. Legato couples scheduled initialization with a consistent velocity field to learn native continuation.

    In the reported pouring comparison, completion time decreased from 95.07 to 75.73 seconds while task score increased from 9.34 to 9.72.

    Role: Theory and derivations, the training–inference-consistent velocity-field formulation, and experimental implementation.

  3. ASTRO overview figure
    ASTRO · First author · RSS 2026 Workshop
    Hang Yu, Di Zhang, Qiwei Du, Yanping Zhao, Hai Zhang, Guang Chen, Junqiao Zhao†, Eduardo E. Veas†

    Fragmented offline trajectories leave useful experience disconnected. ASTRO combines temporal target selection, a diffusion action planner with frozen dynamics, and deviation feedback for self-correction during training and rollout.

    In the reported state-based offline RL evaluation, IQL improved from 36.08 to 45.52 and FQL from 55.52 to 65.71.

    Role: Independently developed the method and carried out the experiments.

  4. ReasoningEval overview figure
    ReasoningEval · Co-first author · Preprint (2025)
    Juncheng Wu*, Sheng Liu*, Haoqin Tu*, Hang Yu*, Xiaoke Huang, Cihang Xie, Yuyin Zhou†

    What improves when an LLM reasons better? Knowledge Index and Information Gain disentangle knowledge from reasoning, revealing different effects of supervised fine-tuning and reinforcement learning in the evaluated models.

  5. SMG overview figure
    SMG · Co-author · NeurIPS 2024
    Di Zhang, Bowen Lv, Hai Zhang, Feifan Yang, Junqiao Zhao†, Hang Yu, Chang Huang, Hongtu Zhou, Chen Ye, Changjun Jiang

    Visual RL can overfit to irrelevant backgrounds. SMG separates task-relevant foreground learning using reward and inverse-dynamics signals, improving generalization under visual distractions.

  6. USB-PO overview figure
    USB-PO · Core author · NeurIPS 2023
    Hai Zhang, Hang Yu, Junqiao Zhao†, Di Zhang, Chang Huang, Hongtu Zhou, Xiao Zhang, Chen Ye

    Better predictive models do not always produce better policies. USB-PO jointly accounts for model shift and model bias with an adaptive model-update objective, improving sample efficiency on six MuJoCo tasks.

Research Experience

Liber-AI · Pretrain Scaling Law Team

Aug. 2026 – Present

Core member working with Fanqi Lin and Songming Liu. I study scaling laws for pretraining both World Action Models (WAM) and Vision-Language-Action (VLA) models, focusing on learning dexterous manipulation from human videos and data.

Robot-free glove collection captures direct human interactions efficiently and covers dexterous tasks that are difficult to perform through teleoperation. The goal is to turn this richer human experience into capable robotic-hand policies. “Robot-free” describes data collection, not a claim of glove-only model training.

Company demonstrations: dexterous manipulation with glove-trained policies →

Spirit-AI · Research Intern

Aug. 2025 – Aug. 2026

Worked with Prof. Yang Gao and Dr. Junyuan Xie. I participated throughout the VLA pretraining pipeline, including benchmark evaluation and post-training demonstrations.

RoboChallenge · #1 / 50.33%
Spirit v1.5, Table 30, reported Jan. 11, 2026. Task-specific fine-tuning across 30 tasks and four robot platforms. Report & demonstrations

RoboArena · #1 / 1918 rating
Spirit v1.6, June 2, 2026 snapshot with 301 A/B comparisons. Company report

These are historical team-level benchmark results under different evaluation protocols, not claims of current leaderboard positions.

TU Graz · Erasmus+ Exchange Researcher

Mar. – Jul. 2025 · Graz, Austria

Worked with Prof. Eduardo Veas on reinforcement learning and dynamics-guided trajectory stitching, connecting fragmented offline experience into useful training trajectories.

TAL · MathGPT Research Intern

Mar. – Jul. 2023

Contributed to the RLHF pipeline for MathGPT, including reward modeling and PPO/DPO alignment for mathematical reasoning.

Education

Leadership & Honors

Class monitor and team captain for the National Undergraduate Intelligent Vehicle Competition, with experience coordinating teammates and communicating technical work.

Services

Invited Session Chair and Co-Chair, IROS 2026.

Conference Reviewers