影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
91.41/100
前 0.5%
全站排名 #307
发表论文23 篇
平均评分
年均产出7.7 篇/年
Yonatan Belinkov
研究方向
Biological foundation models · Emergent communication · Interpretability · robustness · deep learning · representation learning · machine translation · speech recognition · syntactic parsing · question answering
19
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
ICLR 2026Poster
17
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
ICLR 2026Poster
通讯20
Language Models Use Lookbacks to Track Beliefs
ICLR 2026Poster
12
Structured RAG for Answering Aggregative Questions
ICLR 2026Rejected
通讯5
Beyond Natural Language: Invented Communication in Vision-Language Models
ICLR 2026Withdrawn
18
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
ICLR 2025Oral
20
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
ICLR 2025Spotlight
三作21
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
NeurIPS 2025Poster
通讯22
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
COLM 2025Poster
二作22
Arithmetic Without Algorithms: Language Models Solve Math with a Bag of Heuristics
ICLR 2025Poster
通讯26
Inside-Out: Hidden Factual Knowledge in LLMs
COLM 2025Poster
16
CtD: Composition through Decomposition in Emergent Communication
ICLR 2025Poster
三作24
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
ICLR 2025Poster
通讯29
Jamba: Hybrid Transformer-Mamba Language Models
ICLR 2025Poster
12
MIB: A Mechanistic Interpretability Benchmark
ICML 2025Poster
通讯19
Distinguishing Ignorance from Error in LLM Hallucinations
ICLR 2025Withdrawn
通讯24
Context-aware Prompt Tuning: Advancing In-Context Learning with Adversarial Methods
ICLR 2025Rejected
三作