影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
81.3/100
前 1.3%
全站排名 #810
发表论文24 篇
平均评分
年均产出8.0 篇/年
Zhe Gan
研究方向
deep learning · vision and language · deep generative models
13
Scaling Synthetic Task Generation for Agents via Exploration
ICLR 2026Poster
10
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
ICLR 2026Poster
15
UltraCUA: Scaling Computer Use Agent through GUI and Programmatic Control
ICLR 2026Desk Rejected
通讯16
Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
ICLR 2026Rejected
通讯7
Where Did the Reasoning Go Wrong? A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
ICLR 2026Withdrawn
通讯12
GIE-Bench: Towards Grounded Evaluation for Text-Guided Image Editing
ICLR 2026Rejected
通讯5
DeepMMSearch-R1: Empowering Multimodal LLMs in Multi-Modal Web Search
ICLR 2026Withdrawn
通讯11
Contrastive Localized Language-Image Pre-Training
ICML 2025Poster
通讯23
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
ICLR 2025Poster
三作10
Understanding Alignment in Multimodal LLMs: A Comprehensive Study
ICLR 2025Rejected
17
Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms
ICLR 2025Poster
通讯17
SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
COLM 2025Poster
16
MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
ICLR 2025Poster
通讯16
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
ICLR 2025Poster
7
MMEgo: Towards Building Egocentric Multimodal LLMs for Video QA
ICLR 2025Poster
21
SlowFast-LLaVA: A strong training-free baseline for video large language models
ICLR 2025Rejected
三作19
Contrastive Localized Language-Image Pre-Training
ICLR 2025Rejected
通讯15
Improve Vision Language Model Chain-of-thought Reasoning
ICLR 2025Withdrawn
14
Pixelated Instructions: Can Multimodal Large Language Models Follow Printed Instructions in Images?
ICLR 2025Rejected
三作