影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
80.83/100
前 1.3%
全站排名 #841
发表论文24 篇
平均评分
年均产出8.0 篇/年
Qing Li
研究方向
LLM Agents · GUI Agents · Vision-Language Models · 3D-LLM · Computer-Use Agents · Video Understanding & Grounding
23
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
ICLR 2026Poster
通讯24
When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations
ICLR 2026Poster
通讯15
Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
ICLR 2026Rejected
通讯22
MVR: Multi-view Video Reward Shaping for Reinforcement Learning
ICLR 2026Poster
通讯18
STVG-R1: Incentivizing Instance-Level Reasoning and Grounding in Videos via Reinforcement Learning
ICLR 2026Poster
通讯5
AdaReP: Plug-and-Play Acceleration for World Model Predictive Control using Adaptive Re-Planning
ICLR 2026Withdrawn
通讯19
GUI Knowledge Bench: Revealing the Knowledge Gap Behind VLM Failures in GUI Tasks
ICLR 2026Withdrawn
通讯15
On the Effect of Positional Encoding for In-context Learning in Transformers
ICLR 2026Rejected
三作5
Towards the Decisive Factor of Symbolic Generalization of DNNs
ICLR 2026Rejected
28
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
ICLR 2026Rejected
20
KORE: Enhancing Knowledge Injection for Large Multimodal Models via Knowledge-Oriented Augmentations and Constraints
ICLR 2026Rejected
通讯6
Towards the Mysteries of Convergent Interaction Representations through DNNs
ICLR 2026Rejected
二作5
COIN: Chain Of INteraction Benchmark: When Reasoning meets Embodied interaction
ICLR 2026Withdrawn
通讯24
Multi-modal Agent Tuning: Building a VLM-Driven Agent for Efficient Tool Usage
ICLR 2025Spotlight
通讯20
NEP: Autoregressive Image Editing via Next Editing Token Prediction
NeurIPS 2025Poster
通讯24
Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
NeurIPS 2025Poster
通讯30
MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge
ICLR 2025Poster
通讯17
Falcon: Fast Visuomotor Policies via Partial Denoising
ICML 2025Poster
5
LongViTU: Instruction Tuning for Long-Form Video Understanding
ICLR 2025Withdrawn
29
Task-oriented Sequential Grounding in 3D Scenes
ICLR 2025Rejected
通讯