影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
87.53/100
前 0.8%
全站排名 #488
发表论文35 篇
平均评分
年均产出11.7 篇/年
Pengfei Liu
研究方向
Alignment in Large Language Model · Evaluation · benchmark · Pretraining Model
25
SR-Scientist: Scientific Equation Discovery With Agentic AI
ICLR 2026Poster
三作23
MegaScience: Pushing the Frontiers of Open Post-Training Datasets for Science Reasoning
ICLR 2026Rejected
三作18
InnovatorBench: Evaluating Agents’ Ability to Conduct Innovative AI Research
ICLR 2026Poster
通讯24
ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of AI Research
ICLR 2026Rejected
通讯17
GeneVLM: Automated Parsing Executable Digital Gene from a Single Image
ICLR 2026Withdrawn
通讯16
LIMI: Less is More for Agency
ICLR 2026Rejected
通讯15
Efficient Agent Training for Computer Use
ICLR 2026Poster
三作20
Proximal Supervised Fine-Tuning
ICLR 2026Poster
通讯13
DatasetResearch: Benchmarking Agent Systems for Demand-Driven Dataset Discovery
ICLR 2026Rejected
通讯30
One RL to See Them All: Visual Triple Unified Reinforcement Learning
ICLR 2026Rejected
21
Discovering Architectures via an Evolutionary Agentic Framework
ICLR 2026Rejected
通讯5
Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
ICLR 2026Withdrawn
20
Attention Localization Through Separator Tokens: Unlocking Long Numerical Sequence Processing in LLMs
ICLR 2026Withdrawn
11
ARGO: Asynchronous Rollout with Human Guidance for Research Agent Optimization
ICLR 2026Withdrawn
通讯7
Deep Cognition: A Multi-Agent Framework for Collaborative Research with Real-Time Cognitive Oversight
ICLR 2026Withdrawn
通讯11
One Sample to Rule Them All: Extreme Data Efficiency in RL Scaling
ICLR 2026Withdrawn
20
Weak-to-Strong Preference Optimization: Stealing Reward from Weak Aligned Model
ICLR 2025Spotlight
13
Programming Every Example: Lifting Pre-training Data Quality Like Experts at Scale
ICML 2025Poster
通讯16
Progress or Regress? Self-Improvement Reversal in Post-training
ICLR 2025Poster
三作18
On Evaluating LLM Alignment by Evaluating LLMs as Judges
NeurIPS 2025Poster
二作21
LIMOPro: Reasoning Refinement for Efficient and Effective Test-time Scaling
NeurIPS 2025Poster
通讯9
RepoST: Scalable Repository-Level Coding Environment Construction with Sandbox Testing
COLM 2025Poster
13
LIMO: Less is More for Reasoning
COLM 2025Poster
通讯35
Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale
ICLR 2025Rejected
通讯19
OmniBal: Towards Fast Instruction-Tuning for Vision-Language Models via Omniverse Computation Balance
ICML 2025Poster
17
FacTool: Factuality Detection in Generative AI -- A Tool Augmented Framework for Multi-Task and Multi-Domain Scenarios
COLM 2025Poster
通讯29
BeHonest: Benchmarking Honesty in Large Language Models
ICLR 2025Rejected
通讯15
OMNIBAL: TOWARDS FAST INSTRUCT-TUNING FOR VISION-LANGUAGE MODELS VIA OMNIVERSE COMPUTATION BALANCE
ICLR 2025Withdrawn
13
CodeBenchGen: Creating Scalable Execution-based Code Generation Benchmarks
ICLR 2025Rejected
8
TOMVALLEY: EVALUATING THE THEORY OF MIND REASONING OF LLMS IN REALISTIC SOCIAL CONTEXT
ICLR 2025Withdrawn
通讯