影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
96.87/100
前 0.2%
全站排名 #101
发表论文45 篇
平均评分
年均产出15.0 篇/年
13
Cognitive models can reveal interpretable value trade-offs in language models
ICLR 2026Poster
17
The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
ICLR 2026Poster
三作18
Any-Order Flexible Length Masked Diffusion
ICLR 2026Poster
15
Fine-Tuning Masked Diffusion for Provable Self-Correction
ICLR 2026Rejected
17
Seesaw: Accelerating Training by Balancing Batch Size and Learning Rate Scheduling
ICLR 2026Poster
通讯11
Adam or Gauss-Newton? — A Comparative Study In Terms of Basis Alignment and SGD Noise
ICLR 2026Rejected
通讯16
In Good GRACES: Principled Teacher Selection for Knowledge Distillation
ICLR 2026Poster
15
Discovering Hierarchical Latent Capabilities of Language Models via Causal Representation Learning
ICLR 2026Rejected
三作10
Parameter-Efficient Reinforcement Learning using Prefix Optimization
ICLR 2026Poster
三作23
The Emergence of Complex Behavior in Large-Scale Ecological Environments
ICLR 2026Rejected
13
LOTION: Smoothing the Optimization Landscape for Quantized Training
ICLR 2026Rejected
通讯14
Connections between Schedule-Free Optimizers, AdEMAMix, and Accelerated SGD Variants
ICLR 2026Rejected
通讯23
Understanding the Design Space and Cross-Modality Transfer for Vision-Language Models
ICLR 2026Rejected
10
A Mechanistic Analysis of Low-Precision Instabilities in Microscaling Formats
ICLR 2026Withdrawn
11
Random Scaling of Emergent Capabilities
ICLR 2026Rejected
5
Selective Underfitting in Diffusion Models
ICLR 2026Rejected
11
Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions
ICML 2025Oral
17
EvoLM: In Search of Lost Language Model Training Dynamics
NeurIPS 2025Oral
10
Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
COLM 2025Poster
三作9
The Role of Sparsity for Length Generalization in LLMs
ICML 2025Poster
17
Interpreting the linear structure of vision-language model embedding spaces
COLM 2025Poster
23
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
ICLR 2025Oral
6
Mixture of Parrots: Experts improve memorization more than reasoning
ICLR 2025Poster
32
Flash Inference: Near Linear Time Inference for Long Convolution Sequence Models and Beyond
ICLR 2025Poster
通讯34
How Does Critical Batch Size Scale in Pre-training?
ICLR 2025Poster
通讯18
Follow My Instruction and Spill the Beans: Scalable Data Extraction from Retrieval-Augmented Generation Systems
ICLR 2025Poster
32
Eliminating Position Bias of Language Models: A Mechanistic Approach
ICLR 2025Poster
14
A New Perspective on Shampoo's Preconditioner
ICLR 2025Poster
26
SOAP: Improving and Stabilizing Shampoo using Adam for Language Modeling
ICLR 2025Poster
通讯19
Deconstructing What Makes a Good Optimizer for Autoregressive Language Models
ICLR 2025Poster
通讯11
Universal Length Generalization with Turing Programs
ICML 2025Poster
三作25
Multi-Agent Reinforcement Learning from Human Feedback: Data Coverage and Algorithmic Techniques
ICLR 2025Rejected
10
Universal length generalization with Turing Programs
ICLR 2025Rejected
通讯16
Soup to go: mitigating forgetting during continual learning with model averaging
ICLR 2025Rejected