影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
76.05/100
前 1.8%
全站排名 #1,140
发表论文31 篇
平均评分
年均产出10.3 篇/年
16
Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models
ICLR 2026Rejected
18
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
ICLR 2026Poster
15
MARS-Sep: Multimodal-Aligned Reinforced Sound Separation
ICLR 2026Poster
通讯19
PELICAN: Personalized Education via LLM-powered Cognitive Diagnosis and Adaptive Tutoring
ICLR 2026Rejected
17
AlignSep: Temporally-Aligned Video-Queried Sound Separation with Flow Matching
ICLR 2026Poster
14
LLM-I: LLMs are Naturally Interleaved Multimodal Creators
ICLR 2026Rejected
通讯5
FarsightAlign: Early-Stage Test-Time Scaling for Prompt-Aligned Text-to-Image Generation
ICLR 2026Withdrawn
通讯6
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
ICLR 2026Rejected
4
Parameter-Efficient Attention Transfer for Multi-Modal Test-Time Adaptation
ICLR 2026Withdrawn
通讯5
Character Beyond Speech: Leveraging Role-Playing Evaluation in Large Audio Language Models via Reinforcement Learning
ICLR 2026Withdrawn
通讯5
Spot the Key, Recover the Rest: Dual-Path&View Representation Learning for Text-Video Retrieval
ICLR 2026Withdrawn
通讯22
VoxDialogue: Can Spoken Dialogue Systems Understand Information Beyond Words?
ICLR 2025Poster
22
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
ICLR 2025Poster
25
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
ICLR 2025Poster
32
Diff-Prompt: Diffusion-Driven Prompt Generator with Mask Supervision
ICLR 2025Poster
通讯20
Smoothing the Shift: Towards Stable Test-Time Adaptation under Complex Multimodal Noises
ICLR 2025Poster
二作23
T2A-Feedback: Improving Basic Capabilities of Text-to-Audio Generation via Fine-grained AI Feedback
ICLR 2025Withdrawn
14
OmniChat: Enhancing Spoken Dialogue Systems with Scalable Synthetic Data for Diverse Scenarios
ICLR 2025Withdrawn
17
IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models
ICML 2025Poster
二作5
AVSET-10M: An Open Large-Scale Audio-Visual Dataset with High Correspondence
ICLR 2025Withdrawn
6
Noise-Robust Audio-Visual Speech-Driven Body Language Synthesis
ICLR 2025Withdrawn
5
Dynamic Switching Teacher: How to Generalize Temporal Action Detection Models
ICLR 2025Withdrawn
通讯4
MindLoc: A Secure Brain-Based System for Object Localization
ICLR 2025Withdrawn