影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
77.09/100
前 1.7%
全站排名 #1,075
发表论文29 篇
平均评分
年均产出9.7 篇/年
Haitao Mi
研究方向
Agent · Reinforcement learning · Large Language Models · Natural Language Processing · Dialogue System · Machine Translation
21
R-Zero: Self-Evolving Reasoning LLM from Zero Data
ICLR 2026Poster
18
The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context
ICLR 2026Poster
20
DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
ICLR 2026Poster
15
THE END OF MANUAL DECODING: TOWARDS TRULY END-TO-END LANGUAGE MODELS
ICLR 2026Poster
29
Vision-SR1: Self-Rewarding Vision-Language Model via Reasoning Decomposition and Multi-Reward Policy Optimization
ICLR 2026Poster
24
DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains
ICLR 2026Poster
31
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
ICLR 2026Poster
14
InComeS: Integrating Compression and Selection Mechanisms into LLMs for Efficient Model Editing
ICLR 2026Withdrawn
17
On the Evolution of Language Models without Labels: Majority Drives Selection, Novelty Promotes Variation
ICLR 2026Rejected
21
WebAggregator: Scaling Complex Logical Information Aggregation for Web Agents Foundation Models
ICLR 2026Withdrawn
16
DeepTheorem: Advancing LLM Reasoning for Theorem Proving Through Natural Language and Reinforcement Learning
ICLR 2026Rejected
19
One Token to Fool LLM-as-a-Judge
ICLR 2026Rejected
5
CLUE: Non-parametric Verification from Experience via Hidden-State Clustering
ICLR 2026Withdrawn
5
VOGUE: Guiding Exploration with Visual Uncertainty Improves Multimodal Reasoning
ICLR 2026Withdrawn
18
Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models
NeurIPS 2025Spotlight
22
Improving LLM General Preference Alignment via Optimistic Online Mirror Descent
NeurIPS 2025Spotlight
28
Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards
NeurIPS 2025Poster
23
Two Experts Are All You Need for Steering Thinking: Reinforcing Cognitive Effort in MoE Reasoning Models Without Additional Training
NeurIPS 2025Poster
31
MPS-Prover: Advancing Stepwise Theorem Proving by Multi-Perspective Search and Data Curation
NeurIPS 2025Poster
22
UniGist: Towards General and Hardware-aligned Sequence-level Long Context Compression
NeurIPS 2025Poster
20
DOTS: Learning to Reason Dynamically in LLMs via Optimal Reasoning Trajectories Search
ICLR 2025Poster
三作6
Iterative Nash Policy Optimization: Aligning LLMs with General Preferences via No-Regret Learning
ICLR 2025Oral
26
The First Few Tokens Are All You Need: An Efficient and Effective Unsupervised Prefix Fine-Tuning Method for Reasoning Models
NeurIPS 2025Poster
11
Do NOT Think That Much for 2+3=? On the Overthinking of Long Reasoning Models
ICML 2025Poster
14
HDFlow: Enhancing LLM Complex Problem-Solving with Hybrid Thinking and Dynamic Workflows
ICLR 2025Rejected
二作11
Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
ICML 2025Rejected
4
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models
ICLR 2025Withdrawn