影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
86/100
前 0.9%
全站排名 #549
发表论文25 篇
平均评分
年均产出8.3 篇/年
Chao Zhang
研究方向
Audio-visual Large Language Model · Neuroscience · Language processing · Speech recognition
14
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
ICLR 2026Oral
通讯11
YuE: Scaling Open Foundation Models for Long-Form Music Generation
ICLR 2026Poster
21
End-to-end Listen, Look, Speak and Act
ICLR 2026Poster
通讯39
SPEAR: A Unified SSL Framework for Learning Speech and Audio Representations
ICLR 2026Rejected
24
UniFlow-Audio: Unified Flow Matching for Audio Generation from Omni-Modalities
ICLR 2026Rejected
通讯15
SciTS: Scientific Time Series Understanding and Generation with LLMs
ICLR 2026Poster
通讯34
sleep2vec: Unified Cross-Modal Alignment for Heterogeneous Nocturnal Biosignals
ICLR 2026Poster
22
video-SALMONN S: Streaming Audio-Visual LLMs Beyond Length Limits via Memory
ICLR 2026Rejected
通讯19
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
ICLR 2026Rejected
通讯25
Bayesian Speech Synthesisers Can Learn from Multiple Teachers
ICLR 2026Rejected
通讯21
SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation
NeurIPS 2025Poster
通讯19
BrainOmni: A Brain Foundation Model for Unified EEG and MEG Signals
NeurIPS 2025Poster
通讯14
Audio Large Language Models Can Be Descriptive Speech Quality Evaluators
ICLR 2025Poster
35
Bayesian WeakS-to-Strong from Text Classification to Generation
ICLR 2025Poster
通讯19
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
ICLR 2025Rejected
通讯11
video-SALMONN-o1: Reasoning-enhanced Audio-visual Large Language Model
ICML 2025Poster
通讯8
Improving LLM Video Understanding with 16 Frames Per Second
ICML 2025Poster
通讯40
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
ICLR 2025Rejected
通讯5
SALMONN-omni: A Speech Understanding and Generation LLM in a Codec-free Full-duplex Framework
ICLR 2025Withdrawn
通讯