影响力指数
99.08/100
前 0.1%
全站排名 #18
发表论文48
平均评分6.0
年均产出16.0 篇/年

Taiji Suzuki

Full Professor@The University of Tokyo·OpenReview
研究方向

Deep learning theory · tensor analysis · stochastic optimization · Kernel method

8.7
13

Transformers Provably Solve Parity Efficiently with Chain of Thought

ICLR 2025Oral
二作
8.2
15

Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points

NeurIPS 2025Poster
三作
7.3
17

Weighted Point Set Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric

ICLR 2025Spotlight
二作
7.3
18

On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent

ICLR 2025Spotlight
7.3
15

State Size Independent Statistical Error Bound for Discrete Diffusion Models

NeurIPS 2025Poster
二作
7.0
28

Direct Distributional Optimization for Provable Alignment of Diffusion Models

ICLR 2025Poster
通讯
7.0
8

Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation

ICML 2025Poster
通讯
6.8
21

Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency

NeurIPS 2025Poster
三作
6.8
22

Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression

NeurIPS 2025Poster
6.8
19

Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel

NeurIPS 2025Poster
6.8
26

How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?

NeurIPS 2025Poster
通讯
6.7
12

Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression

ICLR 2025Poster
6.6
10

Nonlinear transformers can perform inference-time feature learning

ICML 2025Poster
通讯
6.4
23

From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers

NeurIPS 2025Spotlight
6.3
11

Provable In-Context Vector Arithmetic via Retrieving Task Concepts

ICML 2025Poster
通讯
6.3
8

On the Role of Label Noise in the Feature Learning Process

ICML 2025Poster
通讯
6.3
15

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars

COLM 2025Poster
通讯
6.0
16

Flow matching achieves almost minimax optimal convergence

ICLR 2025Poster
二作
5.8
22

State Space Models are Provably Comparable to Transformers in Dynamic Token Selection

ICLR 2025Poster
二作
5.6
36

Quantifying Memory Utilization with Effective State-Size

ICLR 2025Rejected
5.6
29

Label Noise Gradient Descent Improves Generalization in the Low SNR Regime

ICLR 2025Rejected
通讯
5.5
12

Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble

ICML 2025Poster
通讯
5.5
14

Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning

ICML 2025Poster
通讯
5.3
19

The Role of Label Noise in the Feature Learning Process

ICLR 2025Rejected
通讯
4.9
13

Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models

ICML 2025Poster
二作
4.9
9

Quantifying Memory Utilization with Effective State-Size

ICML 2025Poster