影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
97.77/100
前 0.1%
全站排名 #64
发表论文56 篇
平均评分
年均产出18.7 篇/年
Ming-Hsuan Yang
研究方向
object tracking · object segmentation · scene parsing · low-level vision · image synthesis · vision and language · vision and learning · image generation · multimodal large model
17
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
ICLR 2026Poster
28
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video
ICLR 2026Poster
23
Streaming Autoregressive Video Generation via Diagonal Distillation
ICLR 2026Poster
17
Pursuing Minimal Sufficiency in Spatial Reasoning
ICLR 2026Poster
通讯40
SoCo: Progressive Spectrum Optimization for Large Language Model Compression
ICLR 2026Rejected
通讯17
Multi-Object System Identification from Videos
ICLR 2026Poster
27
OmniLens++: Blind Lens Aberration Correction via Large LensLib Pre-Training and Latent PSF Representation
ICLR 2026Rejected
29
Efficiently Disentangling CLIP for Multi-Object Perception
ICLR 2026Rejected
21
SceneAdapt: Scene-aware Adaptation of Human Motion Diffusion
ICLR 2026Rejected
18
Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization
ICLR 2026Poster
20
SUBench: Benchmarking Spatial Understanding in Vision-Language Models
ICLR 2026Rejected
26
Blockwise SFT for Diffusion Language Models: Reconciling Bidirectional Attention and Autoregressive Decoding
ICLR 2026Rejected
三作5
Scaling Laws for Deepfake Detection
ICLR 2026Withdrawn
通讯5
Beyond the Shot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
ICLR 2026Withdrawn
6
Structured Attention Matters to Multimodal LLMs in Document Understanding
ICLR 2026Rejected
21
No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
ICLR 2025Oral
28
InstaInpaint: Instant 3D-Scene Inpainting with Masked Large Reconstruction Model
NeurIPS 2025Poster
通讯26
RMP-SAM: Towards Real-Time Multi-Purpose Segment Anything
ICLR 2025Oral
通讯29
Restage4D: Reanimating Deformable 3D Reconstruction from a Single Video
NeurIPS 2025Poster
通讯21
RAPID Hand: Robust, Affordable, Perception-Integrated, Dexterous Manipulation Platfrom for Embodied Intelligence
NeurIPS 2025Poster
21
DGS-LRM: Real-Time Deformable 3D Gaussian Reconstruction From Monocular Videos
NeurIPS 2025Poster
29
MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
ICLR 2025Spotlight
通讯31
A Simple Approach to Unifying Diffusion-based Conditional Generation
ICLR 2025Poster
通讯24
Learning Spatial-Semantic Features for Robust Video Object Segmentation
ICLR 2025Poster
通讯37
HQGS: High-Quality Novel View Synthesis with Gaussian Splatting in Degraded Scenes
ICLR 2025Poster
44
IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation
NeurIPS 2025Poster
通讯20
EA3D: Online Open-World 3D Object Extraction from Streaming Videos
NeurIPS 2025Poster
通讯33
HoliGS: Holistic Gaussian Splatting for Embodied View Synthesis
NeurIPS 2025Poster
通讯23
4KAgent: Agentic Any Image to 4K Super-Resolution
NeurIPS 2025Poster
25
Ranking-aware adapter for text-driven image ordering with CLIP
ICLR 2025Poster
三作26
RobuRCDet: Enhancing Robustness of Radar-Camera Fusion in Bird's Eye View for 3D Object Detection
ICLR 2025Poster
通讯19
OmnixR: Evaluating Omni-modality Language Models on Reasoning across Modalities
ICLR 2025Poster
21
Text Speaks Louder than Vision: ASCII Art Reveals Textual Biases in Vision-Language Models
COLM 2025Poster
33
Kitten: A Knowledge-Intensive Evaluation of Image Generation on Visual Entities
ICLR 2025Withdrawn
通讯38
Layout-your-3D: Controllable and Precise 3D Generation with 2D Blueprint
ICLR 2025Poster
通讯24
Customized Procedure Planning in Instructional Videos
ICLR 2025Rejected
13
Three-Dimensional Trajectory Prediction with 3DMoTraj Dataset
ICML 2025Poster
42
Gaga: Group Any Gaussians via 3D-aware Memory Bank
ICLR 2025Withdrawn
通讯30
Hierarchical Information Flow for Generalized Efficient Image Restoration
ICLR 2025Withdrawn
24
Tex4D: Zero-shot 4D Scene Texturing with Video Diffusion Models
ICLR 2025Rejected
三作35
PredFormer: Transformers Are Effective Spatial-Temporal Predictive Learners
ICLR 2025Withdrawn
通讯25
PrML: Progressive Multi-Task Learning for Monocular 3D Human Pose Estimation
ICLR 2025Rejected
通讯11
VideoAlchemy: Open-set Personalization in Video Generation
ICLR 2025Withdrawn
27
HALO: Human-Aligned End-to-end Image Retargeting with Layered Transformations
ICLR 2025Rejected
5
RelationBooth: Towards Relation-Aware Customized Object Generation
ICLR 2025Withdrawn
通讯5
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
ICLR 2025Withdrawn
通讯