影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
72.59/100
前 2.2%
全站排名 #1,438
发表论文22 篇
平均评分
年均产出7.3 篇/年
YiFan Zhang
研究方向
Machine Learning · Computer Vision · Multimodal Large Language Models
26
AudioTrust: Benchmarking The Multifaceted Trustworthiness of Audio Large Language Models
ICLR 2026Poster
16
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
ICLR 2026Poster
一作16
Thyme: Think Beyond Images
ICLR 2026Poster
一作20
BaseReward: A Strong Baseline for Multimodal Reward Model
ICLR 2026Poster
一作19
VITA-E: A Dual-Model Framework for Real-Time, Interruptible, and Concurrent Human-Robot Interaction
ICLR 2026Rejected
43
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models
ICLR 2026Poster
二作29
OmniEarth-Bench: Probing Cognitive Abilities of MLLMs for Earth's Multi-sphere Observation Data
ICLR 2026Withdrawn
31
MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
ICLR 2026Poster
5
Why Distillation can Outperform Zero-RL: The Role of Flexible Reasoning
ICLR 2026Withdrawn
13
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
ICLR 2026Rejected
通讯5
InstructEngine: Instruction-driven Text-to-Image Alignment
ICLR 2026Withdrawn
三作5
VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
ICLR 2026Withdrawn
5
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
ICLR 2026Withdrawn
14
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
ICLR 2026Rejected
7
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
ICML 2025Poster
24
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
NeurIPS 2025Spotlight
34
MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
ICLR 2025Poster
一作11
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
ICML 2025Poster
一作6
ErrorRadar: Benchmarking Complex Mathematical Reasoning of Multimodal Large Language Models Via Error Detection
ICLR 2025Rejected
28
TimeRAF: Retrieval-Augmented Foundation model for Zero-shot Time Series Forecasting
ICLR 2025Withdrawn
三作