影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
90.01/100
前 0.6%
全站排名 #373
发表论文41 篇
平均评分
年均产出13.7 篇/年
Xiaodan Liang
研究方向
Embodied Vision · Cross-modal Understanding and generation · Image/Video Generation and Editing
14
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
ICLR 2026Poster
通讯14
iTryOn: Mastering Interactive Video Virtual Try-On with Spatial-Semantic Guidance
ICLR 2026Rejected
通讯11
FastFit: Accelerating Multi-Reference Virtual Try-On via Cacheable Diffusion Models
ICLR 2026Rejected
通讯37
Does Your 3D Encoder Really Work? A simple yet effective pathway to real 3D scene understanding
ICLR 2026Rejected
通讯16
C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
ICLR 2026Rejected
通讯7
VistaGUI: Towards More Robust and Intelligent GUI Automation
ICLR 2026Rejected
11
SimuPhy: Towards Physical Understanding, Reasoning, and Evaluation via Code Generation
ICLR 2026Rejected
通讯13
Depth-Breadth Synergy in RLVR: Unlocking LLM Reasoning Gains with Adaptive Exploration
ICLR 2026Withdrawn
5
Critique to Verify: Accurate and Honest Test-Time Scaling with RL-Trained Verifiers
ICLR 2026Withdrawn
6
MakeupAnyone: Self-Supervised Identity-Preserving MakeUp Transfer with Region-Aware Multi-Scale Alignment
ICLR 2026Rejected
4
TreeRPO: Tree Relative Policy Optimization
ICLR 2026Withdrawn
-1
CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics
ICLR 2026Desk Rejected
24
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning
NeurIPS 2025Poster
17
MMTryon: Multi-Modal Multi-Reference Control for High-Quality Fashion Generation
ICLR 2025Rejected
通讯31
WISA: World simulator assistant for physics-aware text-to-video generation
NeurIPS 2025Spotlight
通讯25
OptiBench Meets ReSocratic: Measure and Improve LLMs for Optimization Modeling
ICLR 2025Poster
10
GDrag:Towards General-Purpose Interactive Editing with Anti-ambiguity Point Diffusion
ICLR 2025Poster
通讯26
PT-T2I/V: An Efficient Proxy-Tokenized Diffusion Transformer for Text-to-Image/Video-Task
ICLR 2025Poster
通讯12
CatVTON: Concatenation Is All You Need for Virtual Try-On with Diffusion Models
ICLR 2025Poster
通讯29
UniGS: Unified Language-Image-3D Pretraining with Gaussian Splatting
ICLR 2025Poster
通讯23
Sitcom-Crafter: A Plot-Driven Human Motion Generation System in 3D Scenes
ICLR 2025Poster
通讯13
S2-Track: A Simple yet Strong Approach for End-to-End 3D Multi-Object Tracking
ICML 2025Poster
通讯5
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
ICLR 2025Withdrawn
30
UncertaintyRAG: Span Uncertainty Enhanced Long-Context Modeling for Retrieval-Augmented Generation
ICLR 2025Rejected
5
Continual LLaVA: Continual Instruction Tuning in Large Vision-Language Models
ICLR 2025Withdrawn
通讯5
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
ICLR 2025Withdrawn
通讯5
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
ICLR 2025Rejected
通讯5
Memory-Driven Multimodal Chain of Thought for Embodied Long-Horizon Task Planning
ICLR 2025Withdrawn
通讯6
ActionFiller: Fill-In-The-Blank Prompting for OS Agent
ICLR 2025Rejected
通讯