影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
93.25/100
前 0.4%
全站排名 #239
发表论文40 篇
平均评分
年均产出13.3 篇/年
Linjie Li
研究方向
Multimodal Understanding and Generation
21
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
ICLR 2026Poster
一作25
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
ICLR 2026Poster
12
OR-PRM: A Process Reward Model for Algorithmic Problem in Operations Research
ICLR 2026Poster
25
EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
ICLR 2026Poster
15
3D-CoS: A New 3D Reconstruction Paradigm Based on VLM Code Synthesis
ICLR 2026Rejected
三作26
STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
ICLR 2026Poster
三作36
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
ICLR 2026Poster
20
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
ICLR 2026Rejected
10
TextAtlas5M: A Large-Scale Dataset for Long and Structured Text Image Generation
ICLR 2026Rejected
22
InfoAgent: Advancing Autonomous Information‑Seeking Agents
ICLR 2026Rejected
22
V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
ICLR 2026Withdrawn
二作12
Unary Feedback as Observation: Incentivizing Self-Reflection in Large Language Models via Multi-Turn RL
ICLR 2026Rejected
三作12
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
ICLR 2026Rejected
通讯24
Shanks: Simultaneous Hearing and Thinking for Spoken Language Models
ICLR 2026Withdrawn
三作6
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
ICLR 2026Rejected
三作5
Where do Reasoning Models Make a Difference? Follow the Reasoning Leader for Efficient Decoding
ICLR 2026Withdrawn
4
Computer-Use Agents as Judges for Automatic GUI Design
ICLR 2026Withdrawn
三作15
VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
ICLR 2026Rejected
6
The Agent's Marathon: Probing the Limits of Endurance in Long-Horizon Tasks
ICLR 2026Rejected
11
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
ICML 2025Oral
15
MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
ICLR 2025Oral
24
SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
ICLR 2025Spotlight
21
VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
NeurIPS 2025Poster
22
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
NeurIPS 2025Spotlight
24
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
NeurIPS 2025Poster
20
EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
ICLR 2025Poster
26
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
NeurIPS 2025Poster
三作6
GenXD: Generating Any 3D and 4D Scenes
ICLR 2025Poster
18
CertainlyUncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
ICLR 2025Poster
二作39
MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
ICLR 2025Poster
14
OmniContrast: Vision-Language-Interleaved Contrast from Pixels All at once
ICLR 2025Rejected
三作11
Test-Time Preference Optimization: On-the-Fly Alignment via Iterative Textual Feedback
ICML 2025Poster