影响力指数
94.31/100
前 0.3%
全站排名 #197
发表论文35
平均评分5.7
年均产出11.7 篇/年

Hanwang Zhang

Full Professor@Nanyang Technological University·新加坡·OpenReview
研究方向

causal inference · scene graph generation · vision-language

8.2
16

On Path to Multimodal Generalist: General-Level and General-Bench

ICML 2025Oral
通讯
7.3
19

Enhancing CLIP Robustness via Cross-Modality Alignment

NeurIPS 2025Spotlight
通讯
7.3
19

Selftok-Zero: Reinforcement Learning for Visual Generation via Discrete and Autoregressive Visual Tokens

NeurIPS 2025Poster
通讯
7.2
13

$\mathcal{V}ista\mathcal{DPO}$: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

ICML 2025Poster
6.8
22

Co-Reinforcement Learning for Unified Multimodal Understanding and Generation

NeurIPS 2025Spotlight
6.4
30

Vinci: Deep Thinking in Text-to-Image Generation using Unified Model with Reinforcement Learning

NeurIPS 2025Poster
通讯
6.2
8

Towards Semantic Equivalence of Tokenization in Multimodal LLM

ICLR 2025Poster
6.0
28

VR-Sampling: Accelerating Flow Generative Model Training with Variance Reduction Sampling

ICLR 2025Withdrawn
三作
5.5
5

Learning to Animate Images from A Few Videos to Portray Delicate Human Actions

ICLR 2025Withdrawn
5.5
11

3D Question Answering via only 2D Vision-Language Models

ICML 2025Poster
5.0
15

Geo-3DGS: Multi-view Geometry Consistency for 3D Gaussian Splatting and Surface Reconstruction

ICLR 2025Rejected
通讯
4.9
9

Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache Sharing

ICML 2025Poster
三作
4.8
5

Object Fusion via Diffusion Time-step for Customized Image Editing with Single Example

ICLR 2025Withdrawn
4.5
5

Towards Debiased Source-Free Domain Adaptation

ICLR 2025Withdrawn
三作
4.0
5

A Closer Look at Time Steps is Worthy of Triple Speed-Up for Diffusion Model Training

ICLR 2025Withdrawn