影响力指数
96.87/100
前 0.2%
全站排名 #101
发表论文45
平均评分5.7
年均产出15.0 篇/年

Sham M. Kakade

Full Professor@Harvard University·美国·OpenReview
研究方向

Optimization · Machine Learning

7.0
13

Cognitive models can reveal interpretable value trade-offs in language models

ICLR 2026Poster
6.5
17

The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton

ICLR 2026Poster
三作
6.5
18

Any-Order Flexible Length Masked Diffusion

ICLR 2026Poster
5.0
15

Fine-Tuning Masked Diffusion for Provable Self-Correction

ICLR 2026Rejected
5.0
17

Seesaw: Accelerating Training by Balancing Batch Size and Learning Rate Scheduling

ICLR 2026Poster
通讯
5.0
11

Adam or Gauss-Newton? — A Comparative Study In Terms of Basis Alignment and SGD Noise

ICLR 2026Rejected
通讯
4.7
16

In Good GRACES: Principled Teacher Selection for Knowledge Distillation

ICLR 2026Poster
4.5
15

Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning

ICLR 2026Rejected
三作
4.5
10

Parameter-Efficient Reinforcement Learning using Prefix Optimization

ICLR 2026Poster
三作
4.4
23

The Emergence of Complex Behavior in Large-Scale Ecological Environments

ICLR 2026Rejected
4.0
13

LOTION: Smoothing the Optimization Landscape for Quantized Training

ICLR 2026Rejected
通讯
3.5
14

Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants

ICLR 2026Rejected
通讯
3.5
23

Understanding the Design Space and Cross-Modality Transfer for Vision-Language Models

ICLR 2026Rejected
3.3
10

A Mechanistic Analysis of Low-Precision Instabilities in Microscaling Formats

ICLR 2026Withdrawn
3.3
11

Random Scaling of Emergent Capabilities

ICLR 2026Rejected
3.3
5

Selective Underfitting in Diffusion Models

ICLR 2026Rejected
7.8
11

Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions

ICML 2025Oral
7.8
17

EvoLM: In Search of Lost Language Model Training Dynamics

NeurIPS 2025Oral
7.0
10

Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining

COLM 2025Poster
三作
7.0
9

The Role of Sparsity for Length Generalization in LLMs

ICML 2025Poster
7.0
17

Interpreting the linear structure of vision-language model embedding spaces

COLM 2025Poster
7.0
23

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

ICLR 2025Oral
7.0
6

Mixture of Parrots: Experts improve memorization more than reasoning

ICLR 2025Poster
6.8
32

Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond

ICLR 2025Poster
通讯
6.8
34

How Does Critical Batch Size Scale in Pre-training?

ICLR 2025Poster
通讯
6.8
18

Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems

ICLR 2025Poster
6.6
32

Eliminating Position Bias of Language Models: A Mechanistic Approach

ICLR 2025Poster
6.3
14

A New Perspective on Shampoo's Preconditioner

ICLR 2025Poster
6.3
26

SOAP: Improving and Stabilizing Shampoo using Adam for Language Modeling

ICLR 2025Poster
通讯
6.0
19

Deconstructing What Makes a Good Optimizer for Autoregressive Language Models

ICLR 2025Poster
通讯
5.5
11

Universal Length Generalization with Turing Programs

ICML 2025Poster
三作
5.3
25

Multi-Agent Reinforcement Learning from Human Feedback: Data Coverage and Algorithmic Techniques

ICLR 2025Rejected
4.7
10

Universal length generalization with Turing Programs

ICLR 2025Rejected
通讯
4.5
16

Soup to go: mitigating forgetting during continual learning with model averaging

ICLR 2025Rejected