影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
51.39/100
前 7.9%
全站排名 #5,074
发表论文9 篇
平均评分
年均产出3.0 篇/年
Nora Belrose
研究方向
language models · interpretability · adversarial robustness
18
Can We Partially Rewrite Transformers in Natural Language?
NeurIPS 2025Rejected
二作11
Automatically Interpreting Millions of Features in Large Language Models
ICML 2025Poster
通讯14
Automatically Interpreting Millions of Features in Large Language Models
ICLR 2025Rejected
通讯5
Balancing Label Quantity and Quality for Scalable Elicitation
ICLR 2025Rejected
二作6
Understanding Gradient Descent through the Training Jacobian
ICLR 2025Withdrawn
一作