影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
44.71/100
前 11.1%
全站排名 #7,178
发表论文9 篇
平均评分
年均产出3.0 篇/年
10
RESTRAIN: From Spurious Votes to Signals — Self-Training RL with Self-Penalization
ICLR 2026Poster
通讯16
Hybrid Reinforcement: when reward is sparse, better to be dense
ICLR 2026Poster
13
The Era of Real-World Human Interaction: RL from User Conversations
ICLR 2026Rejected
二作6
Bridging Offline and Online Reinforcement Learning for LLMs
ICLR 2026Rejected
13
CoT-Self-Instruct: Building high-quality synthetic prompts data for reasoning and non-reasoning tasks
ICLR 2026Rejected
通讯