Register Any Point: Scaling 3D Point Cloud Registration by Flow Matching
A single flow-matching model for pairwise and multi-view registration, generalizing across scene scales and sensor modalities.
I am a PhD candidate in the Gradient Spaces Lab at Stanford University, working with Iro Armeni and Shuran Song. I am fortunate to be supported by the Stanford Graduate Fellowship (SGF) and the Stanford Robotics Center (SRC).
My research connects 3D vision, generative models, and robotics. I develop geometric representations and generative methods for understanding scenes, estimating poses, and assembling shapes, with the goal of enabling robots to reason about and interact with the physical world.
Since June 2026, I have been a research intern at NVIDIA Cosmos Lab, working on World-Action Models (WAMs) for robot planning and control. Prior to Stanford, I obtained my M.Sc. in Computer Science from ETH Zurich and my B.Sc. in Computer Engineering from Tongji University.
Register Any Point is an ECCV 2026 Best Paper Award Candidate (top 10) and an oral presentation.
Register Any Point was accepted to ECCV 2026.
I joined NVIDIA Cosmos Lab as a research intern, working on world-action models for robot planning and control.
I co-organized the Nothing Stands Still Challenge at CVPR 2026, as part of the Computer Vision for the Built World workshop.
I received the Stanford Robotics Center Robotics Scholars Fellowship.
Energy-based Compositional Diffusion Planning was accepted to ICML 2026.
Nothing Stands Still received the ISPRS Journal Best Paper Award for 2025.
Real-3DQA was accepted to ICLR 2026.
Rectified Point Flow was accepted to NeurIPS 2025 as a Spotlight.
I co-organized the Nothing Stands Still Challenge with HILTI at ICRA 2025, as part of the 4th Workshop on Future of Construction.
We released Seed1.5-Thinking, our technical report on advancing reasoning models with reinforcement learning.
Nothing Stands Still was accepted to the ISPRS Journal of Photogrammetry and Remote Sensing.
I received the Stanford Graduate Fellowship, providing three years of support for my PhD research.
I co-organized the Nothing Stands Still Challenge with HILTI at ICRA 2024, as part of the 3rd Workshop on Future of Construction.
I study how machines can understand the 3D world and use that understanding to act. My work connects geometric learning, generative models, and robot planning.
Learning geometry and correspondences to register, reconstruct, and understand scenes across sensors, scales, and time.
Formulating point cloud registration and shape assembly as generation, capturing structure, symmetry, and uncertainty.
Connecting world-action models and vision-language-action models with robot planning, control, and assembly.
* Equal contribution
A single flow-matching model for pairwise and multi-view registration, generalizing across scene scales and sensor modalities.
Composing short trajectory fragments into coherent long-horizon robot plans with a global energy and consistent boundary corrections.
Exposing language shortcuts in 3D reasoning benchmarks and introducing Real-3DQA to test whether models actually use visual geometry.
One generative formulation for point cloud registration and multi-part assembly, learning geometric symmetries without symmetry labels.
Scaling reasoning through reinforcement learning, with advances in training data, reward modeling, algorithms, and distributed infrastructure.
A real-world benchmark for aligning 3D scenes across construction and renovation, with large structural and temporal changes.
Testing whether deterministic uncertainty estimates remain reliable under distribution shifts and in dense prediction tasks.
Leveraging crowdsourced GPS data to improve and support road extraction from aerial imagery.
Combining satellite imagery with GPS data to improve road extraction quality
Stacked U-Nets, auxiliary supervision, and topology-aware post-processing for extracting connected roads from satellite imagery.
Previously at ByteDance Seed Lab, working on LLM post-training and reasoning, and the Computer Vision Lab at ETH Zurich, working on scene understanding and multitask learning.
M.Sc. in Computer Science
ETH Zurich
B.Sc. in Computer Engineering
Tongji University
Get in touch
I’m happy to connect about 3D vision, world models, and robotics.