影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
55.11/100
前 6.4%
全站排名 #4,131
发表论文11 篇
平均评分
年均产出3.7 篇/年
10
RESTRAIN: From Spurious Votes to Signals — Self-Training RL with Self-Penalization
ICLR 2026Poster
14
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
ICLR 2026Poster
三作16
Hybrid Reinforcement: when reward is sparse, better to be dense
ICLR 2026Poster
通讯15
Jointly Reinforcing Diversity and Quality in Language Model Generations
ICLR 2026Rejected
三作6
Adaptive Decoding via Latent Preference Optimization
ICLR 2026Rejected
三作6
Bridging Offline and Online Reinforcement Learning for LLMs
ICLR 2026Rejected
5
Diverse Preference Optimization
ICLR 2026Rejected
13
CoT-Self-Instruct: Building high-quality synthetic prompts data for reasoning and non-reasoning tasks
ICLR 2026Rejected
一作