Research Shanghai Jiao Tong University
Spatial Intelligence
Neural Rendering3D Gaussian Splatting
Spatial AgentsVision-Language Models
I am Zhihang Zhong (), an associate professor at the School of Artificial Intelligence, Shanghai Jiao Tong University, where I lead Visionary Laboratory (). Previously, I was a researcher at the Shanghai AI Laboratory.
I received my PhD in Computer Science and ME in Precision Engineering from the University of Tokyo, and my BE in Mechatronics from Chu Kochen Honors College, Zhejiang University.
Join our research. We are looking for Masters students, PhD students, interns and RAs.
Send your rsum to zhongzhihang@sjtu.edu.cn.
27
News
- 2026.09I am appointed as an Area Chair for ICLR.
- 2026.07I am appointed as a Senior Program Committee member for AAAI.
- 2026.07Two papers are accepted to ACM MM 2026.
- 2026.06Two papers (one Spotlight) are accepted to ECCV 2026!
- 2026.05We release SpaceDG, the first benchmark for Spatial Intelligence under visual degradations!
- 2026.05Two papers (one Oral: Holi-Spatial) are accepted to ICML 2026!
- 2026.03We release CourtSI, the first benchmark for Sports Spatial Intelligence!
- 2026.03We release Holi-Spatial, a data creation engine that transforms video into spatial intelligence!
- 2026.02Two papers are accepted to CVPR 2026 (one Best Paper Candidate : Proxy-GS)
- 2026.02InterpAny is accepted to TPAMI!
- 2025.12We are thrilled to release Visionary, the World Model Carrier!!
- 2025.11One paper (Oral) is accepted to AAAI 2026.
- 2025.06Three papers are accepted to ICCV 2025.
- 2025.02MaskGaussian is accepted to CVPR 2025.
- 2024.08
- 2024.07Glad to receive the 2023 Chinese Government Award for Outstanding
Self-financed Students Abroad (Group B, Global Top 50)! - 2024.07Three papers (one Oral) are accepted to ECCV 2024!
- 2024.05
- 2024.03Glad to receive the Dean's Award for Academic Achievement from the UTokyo!
- 2024.02Two papers are accepted to CVPR 2024.
- 2023.12I give a talk at OpenMMLab about temporal super-resolution.
- 2023.11Glad to release InterpAny-Clearer project!
- 2023.07Two papers are accepted to ICCV 2023.
- 2023.07One paper is accepted to ACM MM 2023.
- 2023.03I am accepted to CVPR 2023's Doctoral Consortium.
- 2023.02Two papers are accepted to CVPR 2023.
- 2022.10I give a talk at MIPI Workshop 2022.
- 2022.10One paper is accepted to IJCV.
- 2022.09I become a JSPS DC fellow!
- 2022.07Three papers (one Oral) are accepted to ECCV 2022!
- 2022.04I become a JEM intern at Microsoft.
- 2022.03One paper is accepted to CVPR 2022.
- 2021.09I become a research intern in the Visual Computing group at MSRA.
- 2021.04I become a IIW fellow of UTokyo!
- 2021.04One paper is accepted to IoTJ.
- 2021.02One paper is accepted to CVPR 2021.
- 2020.11I become a MSRA D-CORE fellow!
- 2020.09I obtain my M.E. degree from UTokyo with an outstanding thesis award!
- 2020.07One paper (Spotlight) is accepted to ECCV 2020!
- 2019.12One paper is accepted to IUI 2020.
Projects
Publications
denotes corresponding author 2026| Intern-S2-Preview: Scientific Agentic Foundation Model arXiv, 2026 arXiv / code |
| SpaceDG: Benchmarking Spatial Intelligence under Visual Degradation arXiv, 2026 project / arXiv / code |
| PhotoFlow: Agentic 3D Virtual Photography Missions arXiv, 2026 project / arXiv / code |
| Segment and Select: Vision-Language Segmentation in 3D Scenarios arXiv, 2026 arXiv |
| Holi-Spatial: Evolving Video Streams into Holistic 3D Spatial Intelligence ICML, 2026, Oral project / arXiv / code |
| Perceptual Flow Network for Visually Grounded Reasoning ICML, 2026 arXiv |
| Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports arXiv, 2026 project / arXiv / code |
| InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing Technical Report, 2026 arXiv / code |
| GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing ECCV, 2026, Spotlight project / arXiv / code |
| Aligning Anything: Hierarchical Motion Estimation for Video Frame Interpolation ECCV, 2026 |
| Proxy-GS: Unified Occlusion Priors for Training and Inference in Structured 3D Gaussian Splatting CVPR, 2026, Oral, Best Paper Candidate project / arXiv / code |
| Motion-Aware Animatable Gaussian Avatars Deblurring CVPR, 2026 arXiv / code |
| Velocity Disambiguation for Video Frame Interpolation TPAMI, 2026 paper / arXiv |
| RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket Analysis AAAI, 2026, Oral arXiv / code |
| AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models ACM MM, 2026 project / arXiv / code |
| Toward Robust and 3D-Aware RGB-NIR Imaging in the Dark ACM MM, 2026 |
| Visionary: The World Model Carrier Built on WebGPU-Powered Gaussian Splatting Platform Technical Report, 2025 project / arXiv / code / editor |
| CityGS-X: A Scalable Architecture for Efficient and Geometrically Accurate Large-Scale Scene Reconstruction ICCV, 2025 project / arXiv / code |
| Sequential Gaussian Avatars with Hierarchical Motion Context ICCV, 2025 project / arXiv / code |
| Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars ICCV, 2025 arXiv / code |
| MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic Masks CVPR, 2025 project / arXiv / code |
| DiffBody: Human Body Image Restoration with Generative Diffusion Prior ICCP, 2025 arXiv |
| Within the Dynamic Context: Inertia-aware 3D Human Modeling with Pose Sequence ECCV, 2024 project / arXiv / code |
| Clearer Frames, Anytime: Resolving Velocity Ambiguity in Video Frame Interpolation ECCV, 2024, Oral project / arXiv / code |
| KFD-NeRF: Rethinking Dynamic NeRF with Kalman Filter ECCV, 2024 arXiv / code |
| IQ-VFI: Implicit Quadratic Motion Estimation for Video Frame Interpolation CVPR, 2024 paper |
| Fooling Polarization-based Vision using Locally Controllable Polarizing Projection CVPR, 2024 paper / arXiv |
| NIR-assisted Video Enhancement via Unpaired 24-hour Data ICCV, 2023 paper / code |
| Rethinking Video Frame Interpolation from Shutter Mode Induced Degradation ICCV, 2023 paper / code |
| Event-guided Frame Interpolation and Dynamic Range Expansion of Single Rolling Shutter Image ACM MM, 2023 paper |
| Blur Interpolation Transformer for Real-World Motion from Blur CVPR, 2023 project / paper / arXiv / code / zhihu |
| Visibility Constrained Wide-band Illumination Spectrum Design for Seeing-in-the-Dark CVPR, 2023 paper / arXiv / code |
| Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance ECCV, 2022 project / paper / arXiv / code / zhihu |
| Bringing Rolling Shutter Images Alive with Dual Reversed Distortion ECCV, 2022, Oral project / paper / arXiv / code |
| Efficient Video Deblurring Guided by Motion Magnitude ECCV, 2022 paper / arXiv / code |
| Learning Adaptive Warping for Real-World Rolling Shutter Correction CVPR, 2022 paper / arXiv / code |
| Real-world Video Deblurring: A Benchmark Dataset and An Ecient Recurrent Neural Network International Journal of Computer Vision (IJCV), 2022 paper / arXiv / code |
| Towards Rolling Shutter Correction and Deblurring in Dynamic Scenes CVPR, 2021 paper / arXiv / code |
| Multistream Temporal Convolutional Network for Correct/Incorrect Patient Transfer Action Detection using Body Sensor Network IEEE Internet of Things Journal (IoTJ), 2021 paper / code |
| Efficient Spatio-Temporal Recurrent Neural Network for Video Deblurring ECCV, 2020, Spotlight paper / arXiv / code |
| Multi-attention Deep Recurrent Neural Network for Nursing Action Evaluation using Wearable Sensor IUI, 2020 paper |
Teaching
Fall 2026
Parallel Computing and Operator Programming ()
Shanghai Jiao Tong University
AI Engineering (AI)
Shanghai Innovation Institute