影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
87.52/100
前 0.8%
全站排名 #489
发表论文40 篇
平均评分
年均产出13.3 篇/年
Salman Khan
研究方向
Multi-modal Learning · Open-world and open-set learning · Continual Learning · Deep learning · Object Detection · Image Recognition
11
TerraFM: A Scalable Foundation Model for Unified Multisensor Earth Observation
ICLR 2026Poster
通讯20
PersonaX: Multimodal Datasets with LLM-Inferred Behavior Traits
ICLR 2026Poster
20
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
ICLR 2026Poster
通讯28
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Video
ICLR 2026Poster
19
Dr.LLM: Dynamic Layer Routing in LLMs
ICLR 2026Poster
三作19
EvoIR: Towards All-in-One Image Restoration via Evolutionary Frequency Modulation
ICLR 2026Rejected
通讯15
StageVAR: Stage-Aware Acceleration for Visual Autoregressive Models
ICLR 2026Rejected
三作18
MediX-R1: Open Ended Medical Reinforcement Learning
ICLR 2026Rejected
22
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
ICLR 2026Poster
5
TOWARDS CALIBRATING PROMPT TUNING OF VISION- LANGUAGE MODELS
ICLR 2026Withdrawn
5
RainDiff: End to End Precipitation Nowcasting Via Token-wise Attention Diffusion
ICLR 2026Withdrawn
通讯4
ThinkGeo: Evaluating Tool-Augmented Agents for Remote Sensing Tasks
ICLR 2026Withdrawn
通讯5
Towards Multimodal Understanding, Reasoning, and Tool Usage across Vision, Speech, and Audio in Long Videos
ICLR 2026Withdrawn
7
GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
ICLR 2026Withdrawn
通讯7
CASS: Nvidia to AMD Transpilation with Data, Models, and Benchmark
ICLR 2026Withdrawn
8
VideoMolmo: Spatio-Temporal Grounding meets Pointing
ICLR 2026Withdrawn
通讯5
MedGazeShift : Transferable Multimodal Adversarial Attacks for Diagnostic Misdirection in Vision-Language Models
ICLR 2026Withdrawn
5
Planner and Executor: Collaboration between Discrete Diffusion And Auto-regressive Models in Reasoning
ICLR 2026Withdrawn
-1
MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
ICLR 2026Withdrawn
通讯16
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
NeurIPS 2025Spotlight
16
Open-YOLO 3D: Towards Fast and Accurate Open-Vocabulary 3D Instance Segmentation
ICLR 2025Oral
24
DEFT: Decompositional Efficient Fine-Tuning for Text-to-Image Models
NeurIPS 2025Poster
18
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
NeurIPS 2025Poster
22
AdaIR: Adaptive All-in-One Image Restoration via Frequency Mining and Modulation
ICLR 2025Poster
三作18
On the Importance of Language-driven Representation Learning for Heterogeneous Federated Learning
ICLR 2025Poster
13
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
ICML 2025Poster
通讯10
GenZSL: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
ICML 2025Poster
三作5
Dynamic Pre-training: Towards Efficient and Scalable All-in-One Image Restoration
ICLR 2025Withdrawn
5
CPT: Consistent Proxy Tuning for Black-box Optimization
ICLR 2025Withdrawn
5
How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs
ICLR 2025Withdrawn
通讯5
Induction Rather Than Imagination: Generative Zero-Shot Learning Via Inductive Variational Autoencoder
ICLR 2025Withdrawn
三作6
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
ICLR 2025Withdrawn
三作4
GroupMamba: Parameter-Efficient and Accurate Group Visual State Space Model
ICLR 2025Withdrawn
三作