影响力指数
97.47/100
前 0.1%
全站排名 #73
发表论文52
平均评分5.5
年均产出17.3 篇/年

Hengshuang Zhao

Assistant Professor@The University of Hong Kong·中国香港·OpenReview
研究方向

Image/Video/3D Understanding · Classification · Segmentation · Detection · Representation Learning · Multi-model Learning · Unified Architecture Design · Generative Modeling · Visual Content Creation · Generation · Manipulation · Autonomous Driving · Embodied AI · Robot Learning · LLM Applications

8.7
21

Orient Anything V2: Unifying Orientation and Rotation Understanding

NeurIPS 2025Spotlight
7.5
42

PlayerOne: Egocentric World Simulator

NeurIPS 2025Oral
通讯
7.3
20

Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance

NeurIPS 2025Poster
7.2
11

BOOD: Boundary-based Out-Of-Distribution Data Generation

ICML 2025Poster
通讯
7.1
42

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

NeurIPS 2025Poster
6.8
24

Seg-VAR:Image Segmentation with Visual Autoregressive Modeling

NeurIPS 2025Poster
通讯
6.8
22

Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations

NeurIPS 2025Poster
通讯
6.8
25

ROSE: Remove Objects with Side Effects in Videos

NeurIPS 2025Poster
通讯
6.6
10

HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding

ICML 2025Poster
通讯
6.4
27

LiteReality: Graphic-Ready 3D Scene Reconstruction from RGB-D Scans

NeurIPS 2025Poster
6.3
22

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

ICLR 2025Poster
6.1
12

VIP: Vision Instructed Pre-training for Robotic Manipulation

ICML 2025Poster
通讯
6.0
33

MiCo: Multi-image Contrast for Reinforcement Visual Reasoning

NeurIPS 2025Poster
通讯
5.5
24

BOOD: Boundary-based Out-Of-Distribution Data Generation

ICLR 2025Rejected
通讯
5.5
13

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models

ICML 2025Poster
5.3
5

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

ICLR 2025Withdrawn
4.9
13

TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

ICML 2025Poster
4.8
5

HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models

ICLR 2025Withdrawn
4.8
30

Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images

ICLR 2025Withdrawn
通讯
4.7
21

Effective LLM Knowledge Learning Requires Rethinking Generalization

ICLR 2025Rejected
4.5
6

LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence

ICLR 2025Withdrawn
通讯
4.4
12

LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence

ICML 2025Poster
通讯
3.5
5

VIRT: Vision Instructed Transformer for Robotic Manipulation

ICLR 2025Withdrawn
通讯