影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
96.88/100
前 0.2%
全站排名 #100
发表论文42 篇
平均评分
年均产出14.0 篇/年
13
Relative Scaling Laws for LLMs
ICLR 2026Desk Rejected
三作15
Pre-training under infinite compute
ICLR 2026Oral
三作14
Reinforcement Learning for Machine Learning Engineering Agents
ICLR 2026Poster
三作10
WorldGym: World Model as An Environment for Policy Evaluation
ICLR 2026Poster
16
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
ICLR 2026Poster
11
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
ICLR 2026Poster
21
Fantastic Pretraining Optimizers and Where to Find Them
ICLR 2026Poster
通讯24
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
ICLR 2026Poster
22
MLE-Smith: Scaling MLE Tasks with Automated Multi-agent Pipeline
ICLR 2026Poster
16
SpecEval: Evaluating Model Adherence to Behavior Specifications
ICLR 2026Rejected
通讯6
Curating High Quality Pretraining Data for Language Models via Compression Ratios
ICLR 2026Rejected
通讯6
Scheduling data improves fine-tuning data efficiency
ICLR 2026Withdrawn
二作16
UQ: Assessing Language Models on Unsolved Questions
ICLR 2026Rejected
4
AHELM: A Holistic Evaluation of Audio-Language Models
ICLR 2026Withdrawn
通讯5
RoboReward: A Dataset and Benchmark for Vision-Language Reward Models in Robotics
ICLR 2026Withdrawn
5
Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
ICLR 2026Withdrawn
18
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
ICLR 2025Oral
通讯9
Eliciting Language Model Behaviors with Investigator Agents
ICML 2025Poster
15
AIR-BENCH 2024: A Safety Benchmark based on Regulation and Policies Specified Risk Categories
ICLR 2025Spotlight
19
Blackbox Model Provenance via Palimpsestic Membership Inference
NeurIPS 2025Spotlight
通讯11
Reliable and Efficient Amortized Model-based Evaluation
ICML 2025Poster
三作13
Language Models May Verbatim Complete Text They Were Not Explicitly Trained On
ICML 2025Spotlight
31
Audits Under Resource, Data, and Access Constraints: Scaling Laws For Less Discriminatory Alternatives
NeurIPS 2025Poster
26
Model Equality Testing: Which Model is this API Serving?
ICLR 2025Poster
二作34
Reliable and Efficient Amortized Model-based Evaluation
ICLR 2025Rejected
三作17
On the Entropy Calibration of Language Models
NeurIPS 2025Poster
三作13
Independence Tests for Language Models
ICML 2025Spotlight
通讯54
BioDiscoveryAgent: An AI Agent for Designing Genetic Perturbation Experiments
ICLR 2025Poster
29
AutoBencher: Towards Declarative Benchmark Construction
ICLR 2025Poster
12
Auditing Prompt Caching in Language Model APIs
ICML 2025Poster
20
Instruction Following without Instruction Tuning
ICLR 2025Rejected
通讯15
Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape View
ICLR 2025Poster
23
Independence Tests for Language Models
ICLR 2025Rejected
通讯24
VideoAgent: Self-Improving Video Generation
ICLR 2025Rejected
5
On the Entropy Calibration of Language Models
ICLR 2025Withdrawn
三作