影响力指数
97.41/100
前 0.1%
全站排名 #76
发表论文63
平均评分5.3
年均产出21.0 篇/年

Mike Zheng Shou

Assistant Professor@National University of Singapore·新加坡·OpenReview
研究方向

Multimodal · Video Generation · Video Understanding

6.5
24

VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning

ICLR 2026Poster
通讯
6.0
39

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing

ICLR 2026Poster
通讯
5.3
25

TPDiff: Temporal Pyramid Video Diffusion Model

ICLR 2026Poster
二作
5.3
23

Paper2Video: Automatic Video Generation from Scientific Papers

ICLR 2026Rejected
三作
5.0
13

D-AR: Diffusion via Autoregressive Models

ICLR 2026Poster
二作
5.0
22

DD-Ranking: Rethinking the Evaluation of Dataset Distillation

ICLR 2026Rejected
4.8
18

Personalized Vision via Visual In-Context Learning

ICLR 2026Rejected
通讯
4.7
15

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

ICLR 2026Rejected
三作
4.7
12

A Gain for Reconstruction, A Pain for Generation: Exploiting Representation in Visual Tokenization

ICLR 2026Rejected
通讯
4.5
16

Ego-centric Predictive Model Conditioned on Hand Trajectories

ICLR 2026Withdrawn
二作
4.5
16

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

ICLR 2026Rejected
三作
4.5
16

Automated Movie Generation via Multi-Agent CoT Planning

ICLR 2026Rejected
三作
4.0
12

Rethinking Defense for Computer-Use Agents: Context Deception Attacks are Simple to Defend

ICLR 2026Rejected
三作
4.0
5

MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation

ICLR 2026Withdrawn
三作
4.0
22

Code2Video: A Code-centric Paradigm for Educational Video Generation

ICLR 2026Rejected
三作
4.0
5

Mitty: Diffusion-based Human-To-Robot Video Generation

ICLR 2026Withdrawn
通讯
3.5
5

DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection

ICLR 2026Withdrawn
通讯
3.5
11

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point

ICLR 2026Withdrawn
通讯
3.3
4

Computer-Use Agents as Judges for Automatic GUI Design

ICLR 2026Withdrawn
通讯
3.0
5

Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation

ICLR 2026Withdrawn
通讯
2.7
7

Multi-Human Interactive Talking Dataset

ICLR 2026Rejected
三作
7.3
14

macOSWorld: A Multilingual Interactive Benchmark for GUI Agents

NeurIPS 2025Poster
三作
6.8
25

Show-o2: Improved Native Unified Multimodal Models

NeurIPS 2025Poster
三作
6.6
11

Impossible Videos

ICML 2025Poster
三作
6.5
7

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

ICLR 2025Poster
通讯
6.4
31

PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer

NeurIPS 2025Poster
三作
6.4
26

Sparse Image Synthesis via Joint Latent and RoI Flow

NeurIPS 2025Poster
三作
6.4
18

OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data

NeurIPS 2025Poster
三作
6.4
33

Image Watermarks are Removable using Controllable Regeneration from Clean Noise

ICLR 2025Poster
6.4
23

DOTA: Distributional Test-time Adaptation of Vision-Language Models

NeurIPS 2025Poster
6.4
7

MP-Mat: A 3D-and-Instance-Aware Human Matting and Editing Framework with Multiplane Representation

ICLR 2025Poster
通讯
6.4
28

CoFFT: Chain of Foresight-Focus Thought for Visual Language Models

NeurIPS 2025Poster
通讯
6.3
21

Bridging Information Asymmetry in Text-video Retrieval: A Data-centric Approach

ICLR 2025Poster
通讯
6.3
8

WMAdapter: Adding WaterMark Control to Latent Diffusion Models

ICML 2025Poster
通讯
6.0
35

Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models

NeurIPS 2025Poster
通讯
6.0
6

Grounding Multimodal Large Language Model in GUI World

ICLR 2025Poster
三作
6.0
47

DOTA: Distributional Test-Time Adaptation of Vision-Language Models

ICLR 2025Rejected
5.5
20

Personalized Vision via Visual In-Context Learning

NeurIPS 2025Rejected
通讯
5.5
14

OmniContrast: Vision-Language-Interleaved Contrast from Pixels All at once

ICLR 2025Rejected
通讯
5.2
33

WMAdapter: Adding WaterMark Control to Latent Diffusion Models

ICLR 2025Rejected
通讯
5.2
14

VEditBench: Holistic Benchmark for Text-Guided Video Editing

ICLR 2025Rejected
通讯
4.8
5

Improving Autoregressive Image Generation by Mitigating Gradient Bias in Softmax

ICLR 2025Withdrawn
二作
4.5
5

Unsupervised Prior Learning: Discovering Categorical Pose Priors from Videos

ICLR 2025Withdrawn
三作
4.5
5

Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation

ICLR 2025Withdrawn
通讯
4.3
4

X-PlugVid: Versatile Adaptation of Image Plugins for Controllable Video Generation

ICLR 2025Withdrawn
通讯