影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
41.93/100
前 13.1%
全站排名 #8,425
发表论文5 篇
平均评分
年均产出2.5 篇/年
Abhay Sheshadri
研究方向
Explainable AI · Interpretability Tools · Mechanistic Interpretability · Deep Learning · Generative Modeling
12
Why Do Some Language Models Fake Alignment While Others Don't?
NeurIPS 2025Spotlight
一作9
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
ICML 2025Spotlight
三作26
Mechanistic Unlearning: Robust Knowledge Unlearning and Editing via Mechanistic Localization
ICLR 2025Rejected
三作30
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
ICLR 2025Rejected
一作