影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
70.89/100
前 2.5%
全站排名 #1,592
发表论文23 篇
平均评分
年均产出11.5 篇/年
Chaoyou Fu
研究方向
Multimodality · Face Recognition · Image Synthesis
16
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
ICLR 2026Poster
16
Thyme: Think Beyond Images
ICLR 2026Poster
19
VITA-E: A Dual-Model Framework for Real-Time, Interruptible, and Concurrent Human-Robot Interaction
ICLR 2026Rejected
二作20
BaseReward: A Strong Baseline for Multimodal Reward Model
ICLR 2026Poster
20
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
ICLR 2026Poster
43
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
ICLR 2026Poster
三作16
MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs
ICLR 2026Withdrawn
13
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
ICLR 2026Rejected
20
CUARewardBench: Benchmark for Evaluating Reward Models on Computer-using Agent Trajectories
ICLR 2026Rejected
5
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
ICLR 2026Withdrawn
二作5
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
ICLR 2026Withdrawn
5
MME-CC: A Challenging Multi-Modal Evaluation Benchmark of Cognitive Capacity
ICLR 2026Withdrawn
6
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
ICLR 2026Rejected
24
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
NeurIPS 2025Spotlight
一作30
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
NeurIPS 2025Poster
34
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
ICLR 2025Poster
26
Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
ICLR 2025Poster
18
VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model
NeurIPS 2025Poster
三作30
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
NeurIPS 2025Poster
11
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
ICML 2025Poster
13
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
ICML 2025Poster
5
MME-FINANCE: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
ICLR 2025Withdrawn
4
Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
ICML 2025Poster
三作