影响力指数
96.95/100
前 0.2%
全站排名 #97
发表论文62
平均评分5.1
年均产出20.7 篇/年

Zhou Zhao

Full Professor@Zhejiang University·中国·OpenReview
研究方向

Machine Learning · Computer Vision · deep learning

8.7
21

Orient Anything V2: Unifying Orientation and Rotation Understanding

NeurIPS 2025Spotlight
通讯
7.3
17

SPMDM: Enhancing Masked Diffusion Models through Simplifing Sampling Path

NeurIPS 2025Poster
通讯
6.8
31

ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and Editing

NeurIPS 2025Poster
6.6
17

OmniAudio: Generating Spatial Audio from 360-Degree Video

ICML 2025Poster
6.6
22

VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?

ICLR 2025Poster
通讯
6.5
33

EcoFace: Audio-Visual Emotional Co-Disentanglement Speech-Driven 3D Talking Face Generation

ICLR 2025Poster
6.5
29

WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

ICLR 2025Poster
通讯
6.4
30

Vinci: Deep Thinking in Text-to-Image Generation using Unified Model with Reinforcement Learning

NeurIPS 2025Poster
6.3
22

OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces

ICLR 2025Poster
通讯
6.0
25

OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup

ICLR 2025Poster
通讯
5.8
46

Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

ICLR 2025Withdrawn
通讯
5.7
32

MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization

ICLR 2025Rejected
通讯
5.5
12

Dataflow-Guided Neuro-Symbolic Language Models for Type Inference

ICML 2025Poster
5.5
13

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models

ICML 2025Poster
通讯
5.2
6

ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control

ICLR 2025Withdrawn
通讯
5.0
23

T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback

ICLR 2025Withdrawn
通讯
5.0
14

OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios

ICLR 2025Withdrawn
通讯
4.9
17

IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models

ICML 2025Poster
通讯
4.8
7

CodeSync: Synchronizing Large Language Models with Dynamic Code Evolution at Scale

ICML 2025Poster
4.8
5

AVSET-10M: An Open Large-Scale Audio-Visual Dataset with High Correspondence

ICLR 2025Withdrawn
通讯
4.6
42

Analyzing and Mitigating Inconsistency in Discrete Audio Tokens for Neural Codec Language Models

ICLR 2025Rejected
4.4
6

Noise-Robust Audio-Visual Speech-Driven Body Language Synthesis

ICLR 2025Withdrawn
通讯
4.3
7

Advancing Multimodal Unified Discrete Representations

ICLR 2025Withdrawn
通讯
4.3
5

MultiBand: Multi-Task Song Generation with Personalized Prompt-Based Control

ICLR 2025Withdrawn
通讯
4.2
6

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style Representation

ICLR 2025Withdrawn
4.0
5

MEDIC: Zero-shot Music Editing with Disentangled Inversion Control

ICLR 2025Withdrawn
通讯
3.0
5

Fox-TTS: Scalable Flow Transformers for Expressive Zero-Shot Text to Speech

ICLR 2025Withdrawn
通讯
2.3
4

MindLoc: A Secure Brain-Based System for Object Localization

ICLR 2025Withdrawn
通讯