影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
96.95/100
前 0.2%
全站排名 #97
发表论文62 篇
平均评分
年均产出20.7 篇/年
Zhou Zhao
研究方向
Machine Learning · Computer Vision · deep learning
15
SpatialHand: Generative Object Manipulation from 3D Prespective
ICLR 2026Poster
通讯18
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
ICLR 2026Poster
15
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
ICLR 2026Poster
19
Depth Anything with Any Prior
ICLR 2026Poster
通讯37
Figma2Code: Automating Multimodal Design to Code in the Wild
ICLR 2026Poster
16
Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models
ICLR 2026Rejected
通讯17
AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow Matching
ICLR 2026Poster
通讯16
CIAR: Interval-based Collaborative Decoding for Image Generation Acceleration
ICLR 2026Poster
二作21
WavReward: Spoken Dialogue Models With Generalist Reward Evaluators
ICLR 2026Rejected
通讯24
ETC: training-free diffusion models acceleration with Error-aware Trend Consistency
ICLR 2026Rejected
4
PSR: Subject-Consistency Rewards for Multi-Subject Personalized Generation
ICLR 2026Withdrawn
5
DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
ICLR 2026Withdrawn
通讯6
Keep Refining Your Discrete Diffusion Model: A Mixture of Absorbing and Uniform Processes
ICLR 2026Rejected
三作6
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
ICLR 2026Rejected
通讯6
ACDC: Adaptive Cloud-Device Collaboration for Efficient and Accurate Semantic Segmentation
ICLR 2026Rejected
二作21
Orient Anything V2: Unifying Orientation and Rotation Understanding
NeurIPS 2025Spotlight
通讯17
SPMDM: Enhancing Masked Diffusion Models through Simplifing Sampling Path
NeurIPS 2025Poster
通讯31
ThinkSound: Chain-of-Thought Reasoning in Multimodal LLMs for Audio Generation and Editing
NeurIPS 2025Poster
17
OmniAudio: Generating Spatial Audio from 360-Degree Video
ICML 2025Poster
22
VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?
ICLR 2025Poster
通讯33
EcoFace: Audio-Visual Emotional Co-Disentanglement Speech-Driven 3D Talking Face Generation
ICLR 2025Poster
29
WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
ICLR 2025Poster
通讯30
Vinci: Deep Thinking in Text-to-Image Generation using Unified Model with Reinforcement Learning
NeurIPS 2025Poster
22
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
ICLR 2025Poster
通讯25
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
ICLR 2025Poster
通讯46
Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis
ICLR 2025Withdrawn
通讯32
MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
ICLR 2025Rejected
通讯12
Dataflow-Guided Neuro-Symbolic Language Models for Type Inference
ICML 2025Poster
13
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
ICML 2025Poster
通讯6
ControlSpeech: Towards Simultaneous Zero-shot Speaker Cloning and Zero-shot Language Style Control
ICLR 2025Withdrawn
通讯23
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
ICLR 2025Withdrawn
通讯14
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
ICLR 2025Withdrawn
通讯17
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
ICML 2025Poster
通讯7
CodeSync: Synchronizing Large Language Models with Dynamic Code Evolution at Scale
ICML 2025Poster
5
AVSET-10M: An Open Large-Scale Audio-Visual Dataset with High Correspondence
ICLR 2025Withdrawn
通讯42
Analyzing and Mitigating Inconsistency in Discrete Audio Tokens for Neural Codec Language Models
ICLR 2025Rejected
6
Noise-Robust Audio-Visual Speech-Driven Body Language Synthesis
ICLR 2025Withdrawn
通讯7
Advancing Multimodal Unified Discrete Representations
ICLR 2025Withdrawn
通讯5
MultiBand: Multi-Task Song Generation with Personalized Prompt-Based Control
ICLR 2025Withdrawn
通讯6
Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style Representation
ICLR 2025Withdrawn
5
MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
ICLR 2025Withdrawn
通讯5
Fox-TTS: Scalable Flow Transformers for Expressive Zero-Shot Text to Speech
ICLR 2025Withdrawn
通讯4
MindLoc: A Secure Brain-Based System for Object Localization
ICLR 2025Withdrawn
通讯