影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
80.84/100
前 1.3%
全站排名 #838
发表论文34 篇
平均评分
年均产出11.3 篇/年
Yujun Cai
研究方向
Vision Language Understanding · Action Recognition · 3D · Pose Estimation
24
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
ICLR 2026Poster
三作22
AutoDrive-R²: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
ICLR 2026Poster
19
Video-STAR: Reinforcing Open-Vocabulary Action Recognition with Tools
ICLR 2026Poster
13
From Tokens to Nodes: Semantic-Guided Motion Control for Dynamic 3D Gaussian Splatting
ICLR 2026Poster
三作16
Unveiling the Potential of Diffusion Large Language Model in Controllable Generation
ICLR 2026Poster
二作35
ContextNav: Towards Agentic Multimodal In-Context Learning
ICLR 2026Poster
通讯14
T2G-Reasoner: Deep Reasoning for Text-to-Gloss Translation
ICLR 2026Rejected
19
Unveiling Impact of Frequency Components on Membership Inference Attacks for Diffusion Models
ICLR 2026Rejected
二作29
Efficiently Disentangling CLIP for Multi-Object Perception
ICLR 2026Rejected
二作16
STDR: Spatio-Temporal Decoupling for Real-Time Dynamic Scene Rendering
ICLR 2026Rejected
三作27
WavefrontDiffusion: Dynamic Decoding Schedule for Improved Reasoning
ICLR 2026Poster
5
Visual CoT Makes VLMs Smarter but More Fragile
ICLR 2026Withdrawn
三作14
Are LLMs Really Not Knowledgeable? Mining the Submerged Knowledge in LLMs' Memory
ICLR 2026Poster
三作25
Mitigating Coordinate Prediction Bias from Positional Encoding Failures
ICLR 2026Withdrawn
三作23
FrameMind: Frame-Interleaved Chain-of-Thought for Video Reasoning via Reinforcement Learning
ICLR 2026Rejected
通讯4
$A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
ICLR 2026Withdrawn
通讯11
IUT-Plug: A Plug-in tool for Interleaved Image-Text Generation
ICLR 2026Desk Rejected
5
Structuring Reasoning for Complex Rules Beyond Flat Representations
ICLR 2026Withdrawn
5
Thinking with Sound: Audio Chain-of-Thought Enables Multimodal Reasoning for Large Audio-Language Models
ICLR 2026Withdrawn
二作5
Beyond the Shot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
ICLR 2026Withdrawn
二作22
Not Errors but Guardians: Understanding Sink Tokens in Multimodal LLMs
ICLR 2026Rejected
二作26
Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
ICLR 2026Rejected
二作6
When Symbols Speak: Understanding Logo Triggered Texts in Vision-Language Models
ICLR 2026Withdrawn
二作4
Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion LLMs
ICLR 2026Withdrawn
三作5
Detecting and Mitigating Insertion Hallucination in Video-to-Audio Generation
ICLR 2026Withdrawn
三作7
Do "New Snow Tablets" Contain Snow? Large Language Models Over-Rely on Names to Identify Ingredients of Chinese Drugs
ICLR 2026Withdrawn
二作6
Structured Attention Matters to Multimodal LLMs in Document Understanding
ICLR 2026Rejected
三作5
Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
ICLR 2026Withdrawn
通讯5
Symbolic or Numerical? Understanding Physics Problem Solving in Reasoning LLMs
ICLR 2026Withdrawn
二作21
How does Watermarking Affect Visual Language Models in Document Understanding?
COLM 2025Poster
17
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
COLM 2025Poster
三作32
HAIF-GS: Hierarchical and Induced Flow-Guided Gaussian Splatting for Dynamic Scene
NeurIPS 2025Poster
三作21
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
COLM 2025Poster
通讯