影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
56.79/100
前 5.8%
全站排名 #3,726
发表论文29 篇
平均评分
年均产出9.7 篇/年
Ruiyi Zhang
研究方向
Vision-Language · Natural Language Processing · Reinforcement Learning · Machine Learning
11
Bayesian Data Reweighting Improves Retrieval in Knowledge-Based VQA
ICLR 2026Rejected
二作18
MusiXQA: Advancing Visual Music Understanding in Multimodal LLMs
ICLR 2026Rejected
13
Towards Visual Text Grounding of Multimodal Large Language Model
ICLR 2026Rejected
二作11
Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations
ICLR 2026Poster
14
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
ICLR 2026Rejected
通讯14
CURV: Enhancing Chart Understanding through Visual Grounded Reasoning
ICLR 2026Withdrawn
三作16
GUI‑AIMA: Aligning Intrinsic Multi-Modal Attention with a Context Anchor for GUI Grounding
ICLR 2026Rejected
通讯14
CaughtCheating: Is Your MLLM a Good Cheating Detective? Exploring the Boundary of Visual Perception and Reasoning
ICLR 2026Withdrawn
6
MDocAgent: A Multi-Modal Multi-Agent Framework for Document Question Answering
ICLR 2026Rejected
三作13
Reasoning-Based Personalized Generation for Users with Sparse Data
ICLR 2026Rejected
15
A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations
ICLR 2026Withdrawn
21
DynaSaur: Large Language Agents Beyond Predefined Actions
COLM 2025Poster
15
SV-RAG: LoRA-Contextualizing Adaptation of MLLMs for Long Document Understanding
ICLR 2025Poster
二作14
VipAct: Visual-Perception Enhancement via Specialized VLM Agent Collaboration and Tool-use
ICLR 2025Rejected
6
VaQuitA: Enhancing Alignment in LLM-Assisted Zero-Shot Video Understanding
ICLR 2025Withdrawn
二作15
ADOPD-Instruct: A Large-Scale Multimodal Dataset for Document Editing
ICLR 2025Rejected
17
OIDA-QA: A Multimodal Benchmark for Analyzing the Opioid Industry Document Archive
ICLR 2025Rejected
7
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
ICLR 2025Rejected
5
LLaVA-Read: Enhancing Reading Ability of Multimodal Large Language Models
ICLR 2025Rejected
一作6
Enhancing Diffusion Posterior Sampling for Inverse Problems by Integrating Crafted Measurements
ICLR 2025Withdrawn