影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
54.39/100
前 6.7%
全站排名 #4,295
发表论文8 篇
平均评分
年均产出2.7 篇/年
David Lindner
研究方向
Saftey · Robustness · Evaluations · Alignment · Safety · Monitoring · AI Control · Scheming · Deception · Legibility · Faithfulness · Monitorability
17
Large language models can learn and generalize steganographic chain-of-thought under process supervision
NeurIPS 2025Poster
通讯9
MONA: Myopic Optimization with Non-myopic Approval Can Mitigate Multi-step Reward Hacking
ICML 2025Poster
三作24
Evaluating Frontier Models for Stealth and Situational Awareness
NeurIPS 2025Rejected
18
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
ICLR 2025Rejected
二作