影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
71.24/100
前 2.4%
全站排名 #1,557
发表论文11 篇
平均评分
年均产出3.7 篇/年
Sebastian Lapuschkin
研究方向
Explainable Artificial Intelligence · Interpretability · Machine Learning · Artificial Intelligence · Computer Vision
22
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
ICLR 2026Poster
12
Attribution-Guided Decoding
ICLR 2026Poster
18
Atlas-Alignment: Making Interpretability Transferable Across Language Models
ICLR 2026Rejected
三作15
ASIDE: Architectural Separation of Instructions and Data in Language Models
ICLR 2026Poster
16
Circuit Insights: Towards Interpretability Beyond Activations
ICLR 2026Poster
通讯14
Dyslexify: A Mechanistic Defense Against Typographic Attacks in CLIP
ICLR 2026Poster
22
Navigating Neural Space: Revisiting Concept Activation Vectors to Overcome Directional Divergence
ICLR 2025Poster
通讯23
Manipulating Feature Visualizations with Gradient Slingshots
NeurIPS 2025Poster
24
The Atlas of In-Context Learning: How Attention Heads Shape In-Context Retrieval Augmentation
NeurIPS 2025Poster
通讯