影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
91.62/100
前 0.5%
全站排名 #297
发表论文28 篇
平均评分
年均产出9.3 篇/年
Jacob Steinhardt
研究方向
theory · science · value learning · human-compatible AI · adversarial examples · security · robustness
9
Eliciting Language Model Behaviors with Investigator Agents
ICML 2025Poster
通讯15
Monitoring Latent World States in Language Models with Propositional Probes
ICLR 2025Spotlight
三作16
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
ICLR 2025Spotlight
三作17
Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision
ICLR 2025Spotlight
三作9
Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts
ICML 2025Poster
三作19
LLM Layers Immediately Correct Each Other
NeurIPS 2025Poster
通讯19
Interpreting the Second-Order Effects of Neurons in CLIP
ICLR 2025Poster
三作12
What Do Learning Dynamics Reveal About Generalization in LLM Mathematical Reasoning?
ICML 2025Poster
20
Language Models Learn to Mislead Humans via RLHF
ICLR 2025Poster
21
Which Attention Heads Matter for In-Context Learning?
ICML 2025Poster
二作17
Teaching LLMs to Decode Activations Into Natural Language
ICLR 2025Rejected
三作16
VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models
ICLR 2025Poster
11
Adversaries Can Misuse Combinations of Safe Models
ICML 2025Poster
三作21
Evaluating Model Robustness Against Unforeseen Adversarial Attacks
ICLR 2025Rejected
17
Which Attention Heads Matter for In-Context Learning?
ICLR 2025Rejected
二作6
Adversaries Can Misuse Combinations of Safe Models
ICLR 2025Rejected
三作26
Pre-Memorization Train Accuracy Reliably Predicts Generalization in LLM Reasoning
ICLR 2025Rejected
7
SmartBackdoor: Malicious Language Model Agents that Avoid Being Caught
ICLR 2025Withdrawn
通讯