I am a PhD candidate in the Gradient Spaces Lab at Stanford University, working with Iro Armeni and Shuran Song. I am fortunate to be supported by the Stanford Graduate Fellowship (SGF) and the Stanford Robotics Center (SRC).

My research connects 3D vision, generative models, and robotics. I develop geometric representations and generative methods for understanding scenes, estimating poses, and assembling shapes, with the goal of enabling robots to reason about and interact with the physical world.

Since June 2026, I have been a research intern at NVIDIA Cosmos Lab, working on World-Action Models (WAMs) for robot planning and control. Prior to Stanford, I obtained my M.Sc. in Computer Science from ETH Zurich and my B.Sc. in Computer Engineering from Tongji University.

Interests

  • World Models & Embodied AI
  • 3D Vision & Geometric Learning
  • Robot Learning, Planning & Control

News & updates

Register Any Point is an ECCV 2026 Best Paper Award Candidate (top 10) and an oral presentation.

Register Any Point was accepted to ECCV 2026.

I joined NVIDIA Cosmos Lab as a research intern, working on world-action models for robot planning and control.

I co-organized the Nothing Stands Still Challenge at CVPR 2026, as part of the Computer Vision for the Built World workshop.

I received the Stanford Robotics Center Robotics Scholars Fellowship.

Energy-based Compositional Diffusion Planning was accepted to ICML 2026.

Nothing Stands Still received the ISPRS Journal Best Paper Award for 2025.

Real-3DQA was accepted to ICLR 2026.

Show moreShow less

Rectified Point Flow was accepted to NeurIPS 2025 as a Spotlight.

I co-organized the Nothing Stands Still Challenge with HILTI at ICRA 2025, as part of the 4th Workshop on Future of Construction.

We released Seed1.5-Thinking, our technical report on advancing reasoning models with reinforcement learning.

Nothing Stands Still was accepted to the ISPRS Journal of Photogrammetry and Remote Sensing.

I received the Stanford Graduate Fellowship, providing three years of support for my PhD research.

I co-organized the Nothing Stands Still Challenge with HILTI at ICRA 2024, as part of the 3rd Workshop on Future of Construction.

Research

I study how machines can understand the 3D world and use that understanding to act. My work connects geometric learning, generative models, and robot planning.

3D world understanding

Learning geometry and correspondences to register, reconstruct, and understand scenes across sensors, scales, and time.

Generative geometry

Formulating point cloud registration and shape assembly as generation, capturing structure, symmetry, and uncertainty.

World models for action

Connecting world-action models and vision-language-action models with robot planning, control, and assembly.

Publications

Google Scholar

* Equal contribution

2026

2025

2022

2019

2018

Background

Previously at ByteDance Seed Lab, working on LLM post-training and reasoning, and the Computer Vision Lab at ETH Zurich, working on scene understanding and multitask learning.

M.Sc. in Computer Science
ETH Zurich

B.Sc. in Computer Engineering
Tongji University

Open resources

Lecture notes from my time at ETH Zurich.

Guarantees for ML 2021Advanced ML 2020

Get in touch

Let’s talk research.

I’m happy to connect about 3D vision, world models, and robotics.

taosun@stanford.edu