影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
70.27/100
前 2.6%
全站排名 #1,676
发表论文12 篇
平均评分
年均产出4.0 篇/年
Erik Jones
研究方向
Automated Evaluation · Large Language Models · Red-teaming · Selective Classification · Distributional Robustness · Adversarial Attacks and Defenses · Natural Language Processing
13
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
ICLR 2026Rejected
10
Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded Outputs
ICLR 2026Poster
通讯16
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
ICLR 2026Rejected
5
Abstractive Red-Teaming of Language Model Character
ICLR 2026Withdrawn
通讯10
How Do Large Language Monkeys Get Their Power (Laws)?
ICML 2025Oral
16
Uncovering Gaps in How Humans and LLMs Interpret Subjective Language
ICLR 2025Spotlight
一作19
LLM Layers Immediately Correct Each Other
NeurIPS 2025Poster
三作19
Best-of-N Jailbreaking
NeurIPS 2025Poster
11
Adversaries Can Misuse Combinations of Safe Models
ICML 2025Poster
一作6
Adversaries Can Misuse Combinations of Safe Models
ICLR 2025Rejected
一作