影响力指数
88.23/100
前 0.7%
全站排名 #450
发表论文33
平均评分5.3
年均产出11.0 篇/年

Zehan Wang

PhD student@Zhejiang University·中国·OpenReview
研究方向

multi-modal learning

8.7
21

Orient Anything V2: Unifying Orientation and Rotation Understanding

NeurIPS 2025Spotlight
一作
6.6
22

VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?

ICLR 2025Poster
6.5
29

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

ICLR 2025Poster
6.3
22

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

ICLR 2025Poster
一作
6.0
25

OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup

ICLR 2025Poster
三作
5.8
32

Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision

ICLR 2025Poster
5.8
27

Improving Long-Text Alignment for Text-to-Image Diffusion Models

ICLR 2025Poster
5.5
13

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models

ICML 2025Poster
一作
5.2
6

ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control

ICLR 2025Withdrawn
5.0
23

T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback

ICLR 2025Withdrawn
一作
5.0
14

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

ICLR 2025Withdrawn
4.8
5

AVSET-10M: An Open Large-Scale Audio-Visual Dataset with High Correspondence

ICLR 2025Withdrawn
4.4
6

Noise-Robust Audio-Visual Speech-Driven Body Language Synthesis

ICLR 2025Withdrawn
三作
4.3
7

Advancing Multimodal Unified Discrete Representations

ICLR 2025Withdrawn
4.0
5

Dynamic Switching Teacher: How to Generalize Temporal Action Detection Models

ICLR 2025Withdrawn
2.3
4

MindLoc: A Secure Brain-Based System for Object Localization

ICLR 2025Withdrawn