影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
99.08/100
前 0.1%
全站排名 #18
发表论文48 篇
平均评分
年均产出16.0 篇/年
Taiji Suzuki
研究方向
Deep learning theory · tensor analysis · stochastic optimization · Kernel method
12
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
ICLR 2026Rejected
二作11
Inference-time Alignment with Rewards in Anisotropic Besov Spaces: Superiority of Neural Networks over Linear Estimators
ICLR 2026Rejected
二作11
Consistency Is Not Always Correct: Towards Understanding the Role of Exploration in Post-Training Reasoning
ICLR 2026Desk Rejected
通讯12
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
ICLR 2026Rejected
三作15
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
ICLR 2026Poster
二作15
Test time training enhances in-context learning of nonlinear functions
ICLR 2026Rejected
二作21
Provable Benefit of Curriculum in Transformer Tree-Reasoning Post-Training
ICLR 2026Rejected
通讯13
Transformers Provably Solve Parity Efficiently with Chain of Thought
ICLR 2025Oral
二作15
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
NeurIPS 2025Poster
三作17
Weighted Point Set Embedding for Multimodal Contrastive Learning Toward Optimal Similarity Metric
ICLR 2025Spotlight
二作18
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
ICLR 2025Spotlight
15
State Size Independent Statistical Error Bound for Discrete Diffusion Models
NeurIPS 2025Poster
二作28
Direct Distributional Optimization for Provable Alignment of Diffusion Models
ICLR 2025Poster
通讯8
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
ICML 2025Poster
通讯21
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
NeurIPS 2025Poster
三作22
Trained Mamba Emulates Online Gradient Descent in In-Context Linear Regression
NeurIPS 2025Poster
19
Generalization Bound of Gradient Flow through Training Trajectory and Data-dependent Kernel
NeurIPS 2025Poster
26
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
NeurIPS 2025Poster
通讯12
Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
ICLR 2025Poster
10
Nonlinear transformers can perform inference-time feature learning
ICML 2025Poster
通讯23
From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
NeurIPS 2025Spotlight
11
Provable In-Context Vector Arithmetic via Retrieving Task Concepts
ICML 2025Poster
通讯8
On the Role of Label Noise in the Feature Learning Process
ICML 2025Poster
通讯15
When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
COLM 2025Poster
通讯16
Flow matching achieves almost minimax optimal convergence
ICLR 2025Poster
二作22
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
ICLR 2025Poster
二作36
Quantifying Memory Utilization with Effective State-Size
ICLR 2025Rejected
29
Label Noise Gradient Descent Improves Generalization in the Low SNR Regime
ICLR 2025Rejected
通讯12
Propagation of Chaos for Mean-Field Langevin Dynamics and its Application to Model Ensemble
ICML 2025Poster
通讯14
Mixture of Experts Provably Detect and Learn the Latent Cluster Structure in Gradient-Based Learning
ICML 2025Poster
通讯19
The Role of Label Noise in the Feature Learning Process
ICLR 2025Rejected
通讯13
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
ICML 2025Poster
二作9
Quantifying Memory Utilization with Effective State-Size
ICML 2025Poster