影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
60.76/100
前 4.6%
全站排名 #2,992
发表论文9 篇
平均评分
年均产出3.0 篇/年
Aaron Mueller
研究方向
Interpretability · Syntax · Computational psycholinguistics · Evaluation
13
Priors in time: Missing inductive biases for language model interpretability
ICLR 2026Poster
通讯10
Discovering Latent Biases in Language Models with Steering Vectors
ICLR 2026Withdrawn
二作15
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
ICLR 2026Withdrawn
一作13
In-Context Learning Without Copying
ICLR 2026Rejected
18
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
ICLR 2025Oral
通讯22
Arithmetic Without Algorithms: Language Models Solve Math with a Bag of Heuristics
ICLR 2025Poster
三作27
NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals
ICLR 2025Poster
12
MIB: A Mechanistic Interpretability Benchmark
ICML 2025Poster
一作