影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
89/100
前 0.7%
全站排名 #420
发表论文56 篇
平均评分
年均产出18.7 篇/年
30
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
ICLR 2026Poster
14
Latent Diffusion Model without Variational Autoencoder
ICLR 2026Poster
21
UniVideo: Unified Understanding, Generation, and Editing for Videos
ICLR 2026Poster
15
Unified In-Context Video Editing
ICLR 2026Poster
20
Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
ICLR 2026Poster
16
Mitigating Noise Shift in Denoising Generative Models with Noise Awareness Guidance
ICLR 2026Poster
12
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
ICLR 2026Poster
18
Improving Autoregressive Video Modeling with History Understanding
ICLR 2026Poster
24
SimpleGVR: A Simple Baseline for Latent-Cascaded Generative Video Super-Resolution
ICLR 2026Poster
33
VideoSearch Reasoner: Boosting Multimodal Reward Models through Think with Image Reasoning
ICLR 2026Rejected
20
From Inpainting to Editing: A Self-Bootstrapping Paradigm for Context-Rich Visual Dubbing
ICLR 2026Rejected
22
AdaViewPlanner: Adapting Video Diffusion Models for Viewpoint Planning in 4D Scenes
ICLR 2026Poster
13
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
ICLR 2026Rejected
19
Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
ICLR 2026Poster
15
DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers
ICLR 2026Rejected
14
Astra: General Interactive World Model with Autoregressive Denoising
ICLR 2026Poster
11
RelightMaster: Precise Video Relighting with Multi-plane Light Images
ICLR 2026Rejected
19
A Guide to Training Consistency Models
ICLR 2026Rejected
三作4
A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
ICLR 2026Withdrawn
30
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
ICLR 2026Poster
26
Scaling Image and Video Generation via Test-Time Evolutionary Search
ICLR 2026Rejected
20
TivTok: Broadcasting Time-Invariant Tokens for Scalable Video Tokenization
ICLR 2026Rejected
20
FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation
ICLR 2026Poster
5
OmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D Scenes
ICLR 2026Withdrawn
5
Interpreting Any Condition to Caption for Controllable Video Generation
ICLR 2026Withdrawn
27
UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
ICLR 2026Rejected
5
CamPilot: Improving Camera Control in Video Diffusion Model with Efficient Camera Reward Feedback
ICLR 2026Withdrawn
5
Efficient Training-Free High-Resolution Synthesis with Energy Rectification in Diffusion Models
ICLR 2026Withdrawn
13
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
ICLR 2026Rejected
5
PhysMaster: Mastering Physical Representation for Video Generation via Reinforcement Learning
ICLR 2026Withdrawn
16
SPF-Portrait: Towards Pure Text-to-Portrait Customization with Semantic Pollution-Free Fine-Tuning
ICLR 2026Rejected
4
Score Augmentation for Diffusion Models
ICLR 2026Withdrawn
5
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
ICLR 2026Withdrawn
5
RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
ICLR 2026Withdrawn
10
Terra: Explorable Native 3D World Model with Point Latents
ICLR 2026Withdrawn
5
VFXMaster: Unlocking Dynamic Visual Effect Generation via In-Context Learning
ICLR 2026Withdrawn
4
EmoDialogCN: A Multimodal Mandarin Dyadic Dialogue Dataset of Emotions
ICLR 2026Withdrawn
通讯5
Bowtie-flow: Efficient High-Resolution Video Generation with Prior Preservation
ICLR 2026Withdrawn
12
Fast to Train, Fast to Sample: Stable Velocity for Flow Matching
ICLR 2026Rejected
19
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
NeurIPS 2025Spotlight
13
MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
ICML 2025Spotlight
33
Flow-GRPO: Training Flow Matching Models via Online RL
NeurIPS 2025Poster
31
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
NeurIPS 2025Poster
18
VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information Assumption
NeurIPS 2025Poster
28
Training-Free Efficient Video Generation via Dynamic Token Carving
NeurIPS 2025Poster
27
3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
ICLR 2025Poster
25
Stable Segment Anything Model
ICLR 2025Poster
22
Improving Video Generation with Human Feedback
NeurIPS 2025Poster
33
Motion Inversion for Video Customization
ICLR 2025Rejected
15
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
ICLR 2025Poster
32
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
ICLR 2025Poster
24
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
ICLR 2025Rejected
5
Kinda-45M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
ICLR 2025Withdrawn
5
Explicit-Constrained Single Agent for Enhanced Task-Solving in LLMs
ICLR 2025Withdrawn