影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
40.19/100
前 14.3%
全站排名 #9,212
发表论文8 篇
平均评分
年均产出2.7 篇/年
Alex Troy Mallen
研究方向
AI safety · alignment · and control · NLP · time series · probabilistic forecasts · neural coding · neural computation · BCI
12
Why Do Some Language Models Fake Alignment While Others Don't?
NeurIPS 2025Spotlight
14
Automatically Interpreting Millions of Features in Large Language Models
ICLR 2025Rejected
二作11
Automatically Interpreting Millions of Features in Large Language Models
ICML 2025Poster
二作5
Balancing Label Quantity and Quality for Scalable Elicitation
ICLR 2025Rejected
一作