影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
94.2/100
前 0.3%
全站排名 #201
发表论文29 篇
平均评分
年均产出9.7 篇/年
Martin Jaggi
研究方向
Optimization · Distributed Training · Large Language Models · Collaborative Learning
15
Gradient-Normalized Smoothness for Optimization with Approximate Hessians
ICLR 2026Poster
二作12
Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining
ICLR 2026Poster
通讯14
Weight Decay may matter more than µP for Learning Rate Transfer in Practice
ICLR 2026Poster
11
Benchmarking Optimizers for Large Language Model Pretraining
ICLR 2026Rejected
三作15
A Split-Client Approach to Second-Order Optimization
ICLR 2026Rejected
二作13
$\alpha$-LoRA: Effective Fine-Tuning via Base Model Rescaling
ICLR 2026Rejected
通讯6
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference
ICLR 2026Rejected
13
FineWeb2: One Pipeline to Scale Them All — Adapting Pre-Training Data Processing to Every Language
COLM 2025Poster
11
Can Performant LLMs Be Ethical? Quantifying the Impact of Web Crawling Opt-Outs
COLM 2025Poster
26
Effective Interplay between Sparsity and Quantization: From Theory to Practice
ICLR 2025Spotlight
18
GRAPE: Optimize Data Mixture for Group Robust Multi-target Adaptive Pretraining
NeurIPS 2025Poster
三作15
Attention with Markov: A Curious Case of Single-layer Transformers
ICLR 2025Spotlight
14
HyperINF: Unleashing the HyperPower of Schulz's Method for Data Influence Estimation
COLM 2025Poster
三作21
Intrinsic User-Centric Interpretability through Global Mixture of Experts
ICLR 2025Poster
16
URLs Help, Topics Guide: Understanding Metadata Utility in LLM Training
NeurIPS 2025Poster
三作29
Towards Fully FP8 GEMM LLM Training at Scale
NeurIPS 2025Poster
通讯7
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
ICML 2025Poster
通讯16
CoTFormer: A Chain of Thought Driven Architecture with Budget-Adaptive Computation Cost at Inference
ICLR 2025Poster
三作21
On-Device Collaborative Language Modeling via a Mixture of Generalists and Specialists
ICLR 2025Rejected
三作38
HyperINF: Unleashing the HyperPower of the Schulz's Method for Data Influence Estimation
ICLR 2025Rejected
三作