影响力指数
97.77/100
前 0.1%
全站排名 #64
发表论文56
平均评分5.6
年均产出18.7 篇/年

Ming-Hsuan Yang

Senior Staff Research Scientist@Google DeepMind·美国·OpenReview
研究方向

object tracking · object segmentation · scene parsing · low-level vision · image synthesis · vision and language · vision and learning · image generation · multimodal large model

8.0
21

No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images

ICLR 2025Oral
7.8
28

InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model

NeurIPS 2025Poster
通讯
7.5
26

RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything

ICLR 2025Oral
通讯
7.3
29

Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video

NeurIPS 2025Poster
通讯
7.3
21

RAPID Hand: Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platfrom for Embodied Intelligence

NeurIPS 2025Poster
7.3
21

DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos

NeurIPS 2025Poster
7.2
29

MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

ICLR 2025Spotlight
通讯
6.6
31

A Simple Approach to Unifying Diffusion-based Conditional Generation

ICLR 2025Poster
通讯
6.5
24

Learning Spatial-Semantic Features for Robust Video Object Segmentation

ICLR 2025Poster
通讯
6.5
37

HQGS: High-Quality Novel View Synthesis with Gaussian Splatting in Degraded Scenes

ICLR 2025Poster
6.4
44

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation

NeurIPS 2025Poster
通讯
6.4
20

EA3D: Online Open-World 3D Object Extraction from Streaming Videos

NeurIPS 2025Poster
通讯
6.4
33

HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis

NeurIPS 2025Poster
通讯
6.4
23

4KAgent: Agentic Any Image to 4K Super-Resolution

NeurIPS 2025Poster
6.3
25

Ranking-aware adapter for text-driven image ordering with CLIP

ICLR 2025Poster
三作
6.3
26

RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection

ICLR 2025Poster
通讯
6.0
19

OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities

ICLR 2025Poster
5.8
21

Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models

COLM 2025Poster
5.8
33

Kitten: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities

ICLR 2025Withdrawn
通讯
5.5
38

Layout-your-3D: Controllable and Precise 3D Generation with 2D Blueprint

ICLR 2025Poster
通讯
5.5
24

Customized Procedure Planning in Instructional Videos

ICLR 2025Rejected
5.5
13

Three-Dimensional Trajectory Prediction with 3DMoTraj Dataset

ICML 2025Poster
5.3
42

Gaga: Group Any Gaussians via 3D-aware Memory Bank

ICLR 2025Withdrawn
通讯
5.3
30

Hierarchical Information Flow for Generalized Efficient Image Restoration

ICLR 2025Withdrawn
5.0
24

Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models

ICLR 2025Rejected
三作
5.0
35

PredFormer: Transformers Are Effective Spatial-Temporal Predictive Learners

ICLR 2025Withdrawn
通讯
5.0
25

PrML: Progressive Multi-Task Learning for Monocular 3D Human Pose Estimation

ICLR 2025Rejected
通讯
4.8
11

VideoAlchemy: Open-set Personalization in Video Generation

ICLR 2025Withdrawn
4.3
27

HALO: Human-Aligned End-to-end Image Retargeting with Layered Transformations

ICLR 2025Rejected
4.3
5

RelationBooth: Towards Relation-Aware Customized Object Generation

ICLR 2025Withdrawn
通讯
4.0
5

Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration

ICLR 2025Withdrawn
通讯