影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
86.44/100
前 0.8%
全站排名 #530
发表论文38 篇
平均评分
年均产出12.7 篇/年
Di ZHANG
研究方向
Machine Learning · Recommended System · LLM & MLLM · Image & Video Generation Model
16
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
ICLR 2026Poster
15
Unified In-Context Video Editing
ICLR 2026Poster
12
VMoBA: Mixture-of-Block Attention for Video Diffusion Models
ICLR 2026Poster
19
Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
ICLR 2026Poster
15
DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers
ICLR 2026Rejected
11
Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter
ICLR 2026Poster
7
Towards Subject-Consistent and Text-Aligned Personalized Image Generation via Precise Attribute Learning
ICLR 2026Rejected
20
FilMaster: Bridging Cinematic Principles and Generative AI for Automated Film Generation
ICLR 2026Poster
5
Interpreting Any Condition to Caption for Controllable Video Generation
ICLR 2026Withdrawn
22
Physical Dynamics as Next Geometric Graph Prediction
ICLR 2026Rejected
5
Efficient Training-Free High-Resolution Synthesis with Energy Rectification in Diffusion Models
ICLR 2026Withdrawn
13
PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
ICLR 2026Rejected
通讯19
OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers
NeurIPS 2025Spotlight
13
MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
ICML 2025Spotlight
33
Flow-GRPO: Training Flow Matching Models via Online RL
NeurIPS 2025Poster
31
VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
NeurIPS 2025Poster
25
Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models
NeurIPS 2025Poster
26
Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
NeurIPS 2025Poster
27
3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation
ICLR 2025Poster
25
Stable Segment Anything Model
ICLR 2025Poster
28
Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model
ICLR 2025Poster
22
Improving Video Generation with Human Feedback
NeurIPS 2025Poster
11
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
ICML 2025Poster
33
Motion Inversion for Video Customization
ICLR 2025Rejected
15
Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control
ICLR 2025Poster
27
TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types
ICLR 2025Poster
通讯32
SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints
ICLR 2025Poster
通讯24
SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance
ICLR 2025Rejected
29
Geometric Spatiotemporal Transformer to Simulate Long-Term Physical Dynamics
ICLR 2025Rejected
5
Kinda-45M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
ICLR 2025Withdrawn
通讯5
Explicit-Constrained Single Agent for Enhanced Task-Solving in LLMs
ICLR 2025Withdrawn
通讯12
DMQR-RAG: Diverse Multi-Query Rewriting in Retrieval-Augmented Generation
ICLR 2025Withdrawn
5
Recipes for Unbiased Reward Modeling Learning: An Empirically Study
ICLR 2025Withdrawn
12
Generate explorative goals with large language model guidance
ICLR 2025Withdrawn
-1
EVLM: An Efficient Vision-Language Model for Visual Understanding
ICLR 2025Desk Rejected
通讯