影响力指数
91.62/100
前 0.5%
全站排名 #297
发表论文28
平均评分5.9
年均产出9.3 篇/年

Jacob Steinhardt

Assistant Professor@University of California Berkeley·OpenReview
研究方向

theory · science · value learning · human-compatible AI · adversarial examples · security · robustness

7.8
9

Eliciting Language Model Behaviors with Investigator Agents

ICML 2025Poster
通讯
7.5
15

Monitoring Latent World States in Language Models with Propositional Probes

ICLR 2025Spotlight
三作
7.5
16

Uncovering Gaps in How Humans and LLMs Interpret Subjective Language

ICLR 2025Spotlight
三作
7.3
17

Iterative Label Refinement Matters More than Preference Optimization under Weak Supervision

ICLR 2025Spotlight
三作
7.2
9

Extractive Structures Learned in Pretraining Enable Generalization on Finetuned Facts

ICML 2025Poster
三作
6.8
19

LLM Layers Immediately Correct Each Other

NeurIPS 2025Poster
通讯
6.8
19

Interpreting the Second-Order Effects of Neurons in CLIP

ICLR 2025Poster
三作
6.6
12

What Do Learning Dynamics Reveal About Generalization in LLM Mathematical Reasoning?

ICML 2025Poster
6.3
20

Language Models Learn to Mislead Humans via RLHF

ICLR 2025Poster
5.9
21

Which Attention Heads Matter for In-Context Learning?

ICML 2025Poster
二作
5.3
17

Teaching LLMs to Decode Activations Into Natural Language

ICLR 2025Rejected
三作
5.3
16

VibeCheck: Discover and Quantify Qualitative Differences in Large Language Models

ICLR 2025Poster
4.9
11

Adversaries Can Misuse Combinations of Safe Models

ICML 2025Poster
三作
4.8
21

Evaluating Model Robustness Against Unforeseen Adversarial Attacks

ICLR 2025Rejected
4.6
17

Which Attention Heads Matter for In-Context Learning?

ICLR 2025Rejected
二作
4.3
6

Adversaries Can Misuse Combinations of Safe Models

ICLR 2025Rejected
三作
4.3
26

Pre-Memorization Train Accuracy Reliably Predicts Generalization in LLM Reasoning

ICLR 2025Rejected
4.0
7

SmartBackdoor: Malicious Language Model Agents that Avoid Being Caught

ICLR 2025Withdrawn
通讯