影响力指数
86.44/100
前 0.8%
全站排名 #530
发表论文38
平均评分5.6
年均产出12.7 篇/年

Di ZHANG

VP@Kuaishou Technology·中国·OpenReview
研究方向

Machine Learning · Recommended System · LLM & MLLM · Image & Video Generation Model

9.1
19

OmniSync: Towards Universal Lip Synchronization via Diffusion Transformers

NeurIPS 2025Spotlight
7.8
13

MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding

ICML 2025Spotlight
7.5
33

Flow-GRPO: Training Flow Matching Models via Online RL

NeurIPS 2025Poster
7.3
31

VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models

NeurIPS 2025Poster
6.8
25

Decoupling Contrastive Decoding: Robust Hallucination Mitigation in Multimodal Large Language Models

NeurIPS 2025Poster
6.8
26

Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

NeurIPS 2025Poster
6.8
27

3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation

ICLR 2025Poster
6.5
25

Stable Segment Anything Model

ICLR 2025Poster
6.4
28

Solving Token Gradient Conflict in Mixture-of-Experts for Large Vision-Language Model

ICLR 2025Poster
6.4
22

Improving Video Generation with Human Feedback

NeurIPS 2025Poster
6.1
11

MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

ICML 2025Poster
6.0
33

Motion Inversion for Video Customization

ICLR 2025Rejected
6.0
15

Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained Control

ICLR 2025Poster
6.0
27

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

ICLR 2025Poster
通讯
5.9
32

SynCamMaster: Synchronizing Multi-Camera Video Generation from Diverse Viewpoints

ICLR 2025Poster
通讯
5.5
24

SG-Adapter: Enhancing Text-to-Image Generation with Scene Graph Guidance

ICLR 2025Rejected
5.0
29

Geometric Spatiotemporal Transformer to Simulate Long-Term Physical Dynamics

ICLR 2025Rejected
4.5
5

Kinda-45M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content

ICLR 2025Withdrawn
通讯
4.3
5

Explicit-Constrained Single Agent for Enhanced Task-Solving in LLMs

ICLR 2025Withdrawn
通讯
3.5
12

DMQR-RAG: Diverse Multi-Query Rewriting in Retrieval-Augmented Generation

ICLR 2025Withdrawn
3.5
5

Recipes for Unbiased Reward Modeling Learning: An Empirically Study

ICLR 2025Withdrawn
2.0
12

Generate explorative goals with large language model guidance

ICLR 2025Withdrawn
-1

EVLM: An Efficient Vision-Language Model for Visual Understanding

ICLR 2025Desk Rejected
通讯