影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
77.25/100
前 1.7%
全站排名 #1,070
发表论文19 篇
平均评分
年均产出6.3 篇/年
Himabindu Lakkaraju
研究方向
Interpetability · Fairness · and Safety in Machine Learning · Causality · Counterfactual Inference
17
EvoLM: In Search of Lost Language Model Training Dynamics
NeurIPS 2025Oral
20
How Post-Training Reshapes LLMs: A Mechanistic View on Knowledge, Truthfulness, Refusal, and Confidence
COLM 2025Poster
24
More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness
ICLR 2025Oral
三作21
Inference-Time Reward Hacking in Large Language Models
NeurIPS 2025Spotlight
23
Measuring the Faithfulness of Thinking Drafts in Large Reasoning Models
NeurIPS 2025Poster
通讯18
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
ICLR 2025Poster
通讯26
Quantifying Generalization Complexity for Large Language Models
ICLR 2025Poster
21
Towards Unifying Interpretability and Control: Evaluation via Intervention
ICLR 2025Rejected
通讯27
On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
ICLR 2025Rejected
通讯29
Weak-to-Strong Trustworthiness: Eliciting Trustworthiness with Weak Supervision
ICLR 2025Rejected
通讯21
Generalized Group Data Attribution
ICLR 2025Rejected
通讯