影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
62.31/100
前 4.2%
全站排名 #2,700
发表论文8 篇
平均评分
年均产出2.7 篇/年
John Hughes
研究方向
Adversarial Robustness · Scalable Oversight · AI Safety · Automatic Speech Recognition · Self Supervised Learning
10
How Do Large Language Monkeys Get Their Power (Laws)?
ICML 2025Oral
三作12
Why Do Some Language Models Fake Alignment While Others Don't?
NeurIPS 2025Spotlight
二作19
Best-of-N Jailbreaking
NeurIPS 2025Poster
一作29
Looking Inward: Language Models Can Learn About Themselves by Introspection
ICLR 2025Poster
15
Failures to Find Transferable Image Jailbreaks Between Vision-Language Models
ICLR 2025Poster
14
Attacking Audio Language Models with Best-of-N Jailbreaking
ICLR 2025Rejected
一作