影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
97.47/100
前 0.1%
全站排名 #73
发表论文52 篇
平均评分
年均产出17.3 篇/年
Hengshuang Zhao
研究方向
Image/Video/3D Understanding · Classification · Segmentation · Detection · Representation Learning · Multi-model Learning · Unified Architecture Design · Generative Modeling · Visual Content Creation · Generation · Manipulation · Autonomous Driving · Embodied AI · Robot Learning · LLM Applications
11
Anime-Ready: Controllable 3D Anime Character Generation with Body-Aligned Component-Wise Garment Modeling
ICLR 2026Poster
通讯21
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
ICLR 2026Poster
通讯15
SpatialHand: Generative Object Manipulation from 3D Prespective
ICLR 2026Poster
19
Depth Anything with Any Prior
ICLR 2026Poster
17
SigLIP-HD by Fine-to-Coarse Supervision
ICLR 2026Poster
三作17
Diffusion Fine-Tuning: Iterative Refinement for Advanced Grounding with Diffusion Large Language Models
ICLR 2026Desk Rejected
通讯34
Stratified GRPO: Handling Structural Heterogeneity in Reinforcement Learning of LLM Search Agents
ICLR 2026Rejected
16
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
ICLR 2026Poster
通讯23
Less Gaussians, Texture More: 4K Feed-Forward Textured Splatting
ICLR 2026Poster
通讯5
Seeing Beyond Points: Adaptive Gaussian Primitives for 3D Perception
ICLR 2026Withdrawn
通讯5
Visual Spatial Tuning
ICLR 2026Withdrawn
通讯6
From Noisy Traces to Stable Gradients: Bias--Variance Optimized Preference Optimization for Aligning Large Reasoning Models
ICLR 2026Rejected
5
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
ICLR 2026Withdrawn
通讯5
MVGE: Scale-invariant and Temporal-consistent Monocular Video Geometry Estimation
ICLR 2026Withdrawn
通讯5
Bowtie-flow: Efficient High-Resolution Video Generation with Prior Preservation
ICLR 2026Withdrawn
通讯21
Orient Anything V2: Unifying Orientation and Rotation Understanding
NeurIPS 2025Spotlight
42
PlayerOne: Egocentric World Simulator
NeurIPS 2025Oral
通讯20
Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
NeurIPS 2025Poster
11
BOOD: Boundary-based Out-Of-Distribution Data Generation
ICML 2025Poster
通讯42
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
NeurIPS 2025Poster
24
Seg-VAR:Image Segmentation with Visual Autoregressive Modeling
NeurIPS 2025Poster
通讯22
Concerto: Joint 2D-3D Self-Supervised Learning Emerges Spatial Representations
NeurIPS 2025Poster
通讯25
ROSE: Remove Objects with Side Effects in Videos
NeurIPS 2025Poster
通讯10
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
ICML 2025Poster
通讯27
LiteReality: Graphic-Ready 3D Scene Reconstruction from RGB-D Scans
NeurIPS 2025Poster
22
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
ICLR 2025Poster
12
VIP: Vision Instructed Pre-training for Robotic Manipulation
ICML 2025Poster
通讯33
MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
NeurIPS 2025Poster
通讯24
BOOD: Boundary-based Out-Of-Distribution Data Generation
ICLR 2025Rejected
通讯13
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
ICML 2025Poster
5
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
ICLR 2025Withdrawn
13
TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
ICML 2025Poster
5
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
ICLR 2025Withdrawn
30
Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images
ICLR 2025Withdrawn
通讯21
Effective LLM Knowledge Learning Requires Rethinking Generalization
ICLR 2025Rejected
6
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
ICLR 2025Withdrawn
通讯12
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
ICML 2025Poster
通讯5
VIRT: Vision Instructed Transformer for Robotic Manipulation
ICLR 2025Withdrawn
通讯