影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
63.14/100
前 4%
全站排名 #2,548
发表论文12 篇
平均评分
年均产出4.0 篇/年
Adrià Garriga-Alonso
研究方向
mechanistic interpretability · reinforcement learning · sokoban · planning · interpretability · hypothesis testing · estimators · deep learning theory · gaussian processes · bayesian neural networks · markov chain monte carlo · bayesian
11
Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
ICLR 2026Poster
通讯13
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
ICLR 2026Rejected
二作18
Post-Hoc Reasoning in Chain-of-Thought: Evidence from Pre-CoT Probes and Activation Steering
ICLR 2026Rejected
三作13
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
ICLR 2026Rejected
三作21
Among Us: A Sandbox for Measuring and Detecting Agentic Deception
NeurIPS 2025Spotlight
二作31
Interpreting Emergent Planning in Model-Free Reinforcement Learning
ICLR 2025Oral
24
Feature Hedging: Correlated Features Break Narrow Sparse Autoencoders
NeurIPS 2025Rejected
三作20
Interpreting learned search: finding a transition model and value function in an RNN that plays Sokoban
NeurIPS 2025Rejected
通讯16
Planning in a recurrent neural network that plays Sokoban
ICLR 2025Rejected
通讯