影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
71.04/100
前 2.4%
全站排名 #1,573
发表论文16 篇
平均评分
年均产出5.3 篇/年
Samuel Marks
研究方向
interpretability · large language models · model editing
13
Steering Evaluation-Aware Language Models To Act Like They Are Deployed
ICLR 2026Poster
三作24
Liars' Bench: Evaluating Deception Detectors for AI Assistants
ICLR 2026Rejected
通讯14
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
ICLR 2026Rejected
24
Unsupervised Elicitation of Language Models
ICLR 2026Rejected
12
Inoculation Prompting: Instructing LLMs to misbehave at train-time improves test-time alignment
ICLR 2026Rejected
通讯11
Eliciting Secret Knowledge from Language Models
ICLR 2026Rejected
通讯12
Robustly Improving LLM Fairness in Realistic Settings via Interpretability
ICLR 2026Rejected
二作18
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
ICLR 2025Oral
一作27
NNsight and NDIF: Democratizing Access to Open-Weight Foundation Model Internals
ICLR 2025Poster
21
Erasing Conceptual Knowledge from Language Models
NeurIPS 2025Poster
三作10
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability
ICML 2025Poster
15
Erasing Conceptual Knowledge from Language Models
ICLR 2025Rejected
三作