影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
97.41/100
前 0.1%
全站排名 #76
发表论文63 篇
平均评分
年均产出21.0 篇/年
Mike Zheng Shou
研究方向
Multimodal · Video Generation · Video Understanding
24
VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning
ICLR 2026Poster
通讯39
Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
ICLR 2026Poster
通讯25
TPDiff: Temporal Pyramid Video Diffusion Model
ICLR 2026Poster
二作23
Paper2Video: Automatic Video Generation from Scientific Papers
ICLR 2026Rejected
三作13
D-AR: Diffusion via Autoregressive Models
ICLR 2026Poster
二作22
DD-Ranking: Rethinking the Evaluation of Dataset Distillation
ICLR 2026Rejected
18
Personalized Vision via Visual In-Context Learning
ICLR 2026Rejected
通讯15
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
ICLR 2026Rejected
三作12
A Gain for Reconstruction, A Pain for Generation: Exploiting Representation in Visual Tokenization
ICLR 2026Rejected
通讯16
Ego-centric Predictive Model Conditioned on Hand Trajectories
ICLR 2026Withdrawn
二作16
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
ICLR 2026Rejected
三作16
Automated Movie Generation via Multi-Agent CoT Planning
ICLR 2026Rejected
三作12
Rethinking Defense for Computer-Use Agents: Context Deception Attacks are Simple to Defend
ICLR 2026Rejected
三作5
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
ICLR 2026Withdrawn
三作22
Code2Video: A Code-centric Paradigm for Educational Video Generation
ICLR 2026Rejected
三作5
Mitty: Diffusion-based Human-To-Robot Video Generation
ICLR 2026Withdrawn
通讯5
DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection
ICLR 2026Withdrawn
通讯11
WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point
ICLR 2026Withdrawn
通讯4
Computer-Use Agents as Judges for Automatic GUI Design
ICLR 2026Withdrawn
通讯5
Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation
ICLR 2026Withdrawn
通讯7
Multi-Human Interactive Talking Dataset
ICLR 2026Rejected
三作14
macOSWorld: A Multilingual Interactive Benchmark for GUI Agents
NeurIPS 2025Poster
三作25
Show-o2: Improved Native Unified Multimodal Models
NeurIPS 2025Poster
三作11
Impossible Videos
ICML 2025Poster
三作7
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
ICLR 2025Poster
通讯31
PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
NeurIPS 2025Poster
三作26
Sparse Image Synthesis via Joint Latent and RoI Flow
NeurIPS 2025Poster
三作18
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
NeurIPS 2025Poster
三作33
Image Watermarks are Removable using Controllable Regeneration from Clean Noise
ICLR 2025Poster
23
DOTA: Distributional Test-time Adaptation of Vision-Language Models
NeurIPS 2025Poster
7
MP-Mat: A 3D-and-Instance-Aware Human Matting and Editing Framework with Multiplane Representation
ICLR 2025Poster
通讯28
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
NeurIPS 2025Poster
通讯21
Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach
ICLR 2025Poster
通讯8
WMAdapter: Adding WaterMark Control to Latent Diffusion Models
ICML 2025Poster
通讯35
Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
NeurIPS 2025Poster
通讯6
Grounding Multimodal Large Language Model in GUI World
ICLR 2025Poster
三作47
DOTA: Distributional Test-Time Adaptation of Vision-Language Models
ICLR 2025Rejected
20
Personalized Vision via Visual In-Context Learning
NeurIPS 2025Rejected
通讯14
OmniContrast: Vision-Language-Interleaved Contrast from Pixels All at once
ICLR 2025Rejected
通讯33
WMAdapter: Adding WaterMark Control to Latent Diffusion Models
ICLR 2025Rejected
通讯14
VEditBench: Holistic Benchmark for Text-Guided Video Editing
ICLR 2025Rejected
通讯5
Improving Autoregressive Image Generation by Mitigating Gradient Bias in Softmax
ICLR 2025Withdrawn
二作5
Unsupervised Prior Learning: Discovering Categorical Pose Priors from Videos
ICLR 2025Withdrawn
三作5
Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation
ICLR 2025Withdrawn
通讯4
X-PlugVid: Versatile Adaptation of Image Plugins for Controllable Video Generation
ICLR 2025Withdrawn
通讯