影响力指数
80.84/100
前 1.3%
全站排名 #838
发表论文34
平均评分4.3
年均产出11.3 篇/年

Yujun Cai

Assistant Professor@The University of Queensland·澳大利亚·OpenReview
研究方向

Vision Language Understanding · Action Recognition · 3D · Pose Estimation

6.0
24

ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations

ICLR 2026Poster
三作
6.0
22

AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving

ICLR 2026Poster
6.0
19

Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools

ICLR 2026Poster
5.5
13

From Tokens to Nodes: Semantic-Guided Motion Control for Dynamic 3D Gaussian Splatting

ICLR 2026Poster
三作
5.0
16

Unveiling the Potential of Diffusion Large Language Model in Controllable Generation

ICLR 2026Poster
二作
5.0
35

ContextNav: Towards Agentic Multimodal In-Context Learning

ICLR 2026Poster
通讯
4.7
14

T2G-Reasoner: Deep Reasoning for Text-to-Gloss Translation

ICLR 2026Rejected
4.5
19

Unveiling Impact of Frequency Components on Membership Inference Attacks for Diffusion Models

ICLR 2026Rejected
二作
4.5
29

Efficiently Disentangling CLIP for Multi-Object Perception

ICLR 2026Rejected
二作
4.5
16

STDR: Spatio-Temporal Decoupling for Real-Time Dynamic Scene Rendering

ICLR 2026Rejected
三作
4.5
27

WavefrontDiffusion: Dynamic Decoding Schedule for Improved Reasoning

ICLR 2026Poster
4.0
5

Visual CoT Makes VLMs Smarter but More Fragile

ICLR 2026Withdrawn
三作
4.0
14

Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory

ICLR 2026Poster
三作
4.0
25

Mitigating Coordinate Prediction Bias from Positional Encoding Failures

ICLR 2026Withdrawn
三作
4.0
23

FrameMind: Frame-Interleaved Chain-of-Thought for Video Reasoning via Reinforcement Learning

ICLR 2026Rejected
通讯
4.0
4

$A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement

ICLR 2026Withdrawn
通讯
4.0
11

IUT-Plug: A Plug-in tool for Interleaved Image-Text Generation

ICLR 2026Desk Rejected
4.0
5

Structuring Reasoning for Complex Rules Beyond Flat Representations

ICLR 2026Withdrawn
3.5
5

Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning for Large Audio-Language Models

ICLR 2026Withdrawn
二作
3.5
5

Beyond the Shot: Rethinking Cinematography Understanding with Foundational Skill Evaluation

ICLR 2026Withdrawn
二作
3.5
22

Not Errors but Guardians: Understanding Sink Tokens in Multimodal LLMs

ICLR 2026Rejected
二作
3.5
26

Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding

ICLR 2026Rejected
二作
3.3
6

When Symbols Speak: Understanding Logo Triggered Texts in Vision-Language Models

ICLR 2026Withdrawn
二作
3.3
4

Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion LLMs

ICLR 2026Withdrawn
三作
3.0
5

Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation

ICLR 2026Withdrawn
三作
2.5
7

Do "New Snow Tablets" Contain Snow? Large Language Models Over-Rely on Names to Identify Ingredients of Chinese Drugs

ICLR 2026Withdrawn
二作
2.5
6

Structured Attention Matters to Multimodal LLMs in Document Understanding

ICLR 2026Rejected
三作
2.5
5

Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding

ICLR 2026Withdrawn
通讯
1.5
5

Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs

ICLR 2026Withdrawn
二作