影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
57.77/100
前 5.4%
全站排名 #3,476
发表论文17 篇
平均评分
年均产出5.7 篇/年
Yunhang Shen
研究方向
Large Multimodal Models · object detection · image segmentation · weakly supervised learning · semi-supervised learning
20
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
ICLR 2026Poster
19
VITA-E: A Dual-Model Framework for Real-Time, Interruptible, and Concurrent Human-Robot Interaction
ICLR 2026Rejected
19
FlexibleLLM: Making Low-Bit Quantization for Large Language Models More Flexible and Efficient
ICLR 2026Rejected
18
DeepOmni: Towards Seamless and Smart Speech Interaction with Adaptive Modality-Specific MoE
ICLR 2026Rejected
三作14
LLaVA-RadZ: Can Multimodal Large Language Models Effectively Tackle Zero-shot Radiology Recognition?
ICLR 2026Withdrawn
7
Breaking the Bias: Quantifying the Attention of Industrial Anomaly Detection
ICLR 2026Withdrawn
5
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
ICLR 2026Withdrawn
5
Pseudo-Label Supervision in Unsupervised Industrial Anomaly Detection
ICLR 2026Withdrawn
24
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
NeurIPS 2025Spotlight
26
Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
ICLR 2025Poster
18
VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model
NeurIPS 2025Poster
二作11
FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
ICML 2025Poster
三作30
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
NeurIPS 2025Poster
42
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification
ICLR 2025Poster
三作15
DS-VLM: Diffusion Supervision Vision Language Model
ICML 2025Poster
二作4
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
ICML 2025Poster