影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
35.62/100
前 18.3%
全站排名 #11,760
发表论文9 篇
平均评分
年均产出3.0 篇/年
Dylan Hadfield-Menell
研究方向
Value Alignment · Inverse Reinforcement Learning · Preference Elicitation · Human-Robot Interaction · Sequential Decision Making · Planning · Motion Planning · Markov decision processes · Reinforcement Learning
21
Diverse Preference Learning for Capabilities and Alignment
ICLR 2025Poster
三作30
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
ICLR 2025Rejected
26
Altared Environments: The Role of Normative Infrastructure in AI Alignment
ICLR 2025Rejected
6
Inverse Prompt Engineering for Task-Specific LLM Safety
ICLR 2025Rejected
二作