影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
62.7/100
前 4.1%
全站排名 #2,635
发表论文9 篇
平均评分
年均产出3.0 篇/年
10
All Code, No Thought: Language Models Struggle to Reason in Ciphered Language
ICLR 2026Poster
三作12
Steering Language Models with Weight Arithmetic
ICLR 2026Poster
二作24
Unsupervised Elicitation of Language Models
ICLR 2026Rejected
12
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
ICLR 2026Rejected
12
Why Do Some Language Models Fake Alignment While Others Don't?
NeurIPS 2025Spotlight
通讯15
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
NeurIPS 2025Poster
15
Quantifying Elicitation of Latent Capabilities in Language Models
NeurIPS 2025Poster
23
Do Unlearning Methods Remove Information from Language Model Weights?
ICLR 2025Rejected
二作