影响力指数
95.83/100
前 0.2%
全站排名 #146
发表论文64
平均评分5.3
年均产出21.3 篇/年

Fei Huang

Senior Research Director@Alibaba Group US·美国·OpenReview
研究方向

Machine Learning · Machine Translation · Natural Language Processing · machine translation · natural language processing · information extraction · dialogue system

7.0
24

Adaptive Social Learning via Mode Policy Optimization for Language Agents

ICLR 2026Poster
6.5
28

Perception-Aware Policy Optimization for Multimodal Reasoning

ICLR 2026Poster
6.5
27

WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research

ICLR 2026Poster
6.0
17

Demystifying Deep Search: A Holistic Evaluation with Hint-free Multi-Hop Questions and Factorised Metrics

ICLR 2026Poster
6.0
15

SPELL: Self-Play Reinforcement Learning for Evolving Long-Context Language Models

ICLR 2026Poster
通讯
6.0
22

Scaling Generalist Data-Analytic Agents

ICLR 2026Poster
6.0
22

SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference Optimization

ICLR 2026Poster
通讯
5.5
20

SASFT: Sparse Autoencoder-guided Supervised Finetuning to Mitigate Unexpected Code-Switching in LLMs

ICLR 2026Poster
5.5
21

Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis

ICLR 2026Poster
5.5
20

OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents

ICLR 2026Poster
通讯
5.5
10

AgentFold: Long-Horizon Web Agents with Proactive Context Folding

ICLR 2026Poster
5.5
23

WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent

ICLR 2026Poster
5.3
15

MaskSearch: Towards Scalable Agentic Pre-Training for Search-Enhanced Reasoning

ICLR 2026Rejected
5.0
16

Overcoming Joint Intractability with Lossless Hierarchical Speculative Decoding

ICLR 2026Oral
二作
5.0
24

Agentic Reinforcement Learning with Implicit Step Rewards

ICLR 2026Poster
5.0
20

Efficient Multimodal Planning Agent for Visual Question-Answering

ICLR 2026Rejected
5.0
14

ZeroSearch: Incentivize the Search Capability of LLMs without Searching

ICLR 2026Desk Rejected
5.0
20

Repurposing Synthetic Data for Fine-grained Search Agent Supervision

ICLR 2026Poster
5.0
17

WebSailor-V2: Bridging the Chasm to Proprietary Agents via Synthetic Data and Scalable Reinforcement Learning

ICLR 2026Poster
4.7
21

P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling

ICLR 2026Oral
4.7
16

Scaling Agents via Continual Pre-training

ICLR 2026Poster
4.5
13

AutoHete: An Automatic and Efficient Heterogeneous Training System for LLMs

ICLR 2026Rejected
4.5
28

A Simple "Motivation" Can Enhance Reinforcement Finetuning of Large Reasoning Models

ICLR 2026Poster
4.4
26

Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning

ICLR 2026Withdrawn
4.0
15

mGRPO: Unlocking LLM Reasoning through Multilingual Thinking

ICLR 2026Rejected
4.0
20

MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

ICLR 2026Rejected
4.0
27

ReSum: Unlocking Long-Horizon Search Intelligence via Context Summarization

ICLR 2026Rejected
4.0
13

Towards General Agentic Intelligence via Environment Scaling

ICLR 2026Withdrawn
4.0
33

IterResearch: Rethinking Long-Horizon Agents with Interaction Scaling

ICLR 2026Poster
3.5
5

UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning

ICLR 2026Withdrawn
3.5
24

MARS: Optimizing Dual-System Deep Research via Multi-Agent Reinforcement Learning

ICLR 2026Rejected
通讯
3.0
5

DynamicBench: Evaluating Real-Time Report Generation in Large Language Models

ICLR 2026Withdrawn
2.5
5

End-to-End QA Construction Pipeline for Continual Pre-training of Large Language Models

ICLR 2026Withdrawn
2.0
5

TimeHC-RL: Temporal‑aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence

ICLR 2026Withdrawn
-1

WebSailor: Navigating Super-human Reasoning for Web Agent

ICLR 2026Desk Rejected
7.8
26

Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding

NeurIPS 2025Spotlight
7.8
18

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning

NeurIPS 2025Poster
7.8
22

Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation

NeurIPS 2025Poster
7.2
15

Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning

ICML 2025Poster
7.0
18

On the Role of Attention Heads in Large Language Model Safety

ICLR 2025Oral
6.8
21

VLM-R³: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

NeurIPS 2025Poster
6.8
28

OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-time Emotional Speech Synthesis

NeurIPS 2025Poster
6.8
22

WebDancer: Towards Autonomous Information Seeking Agency

NeurIPS 2025Poster
6.8
14

StructRAG: Boosting Knowledge Intensive Reasoning of LLMs via Inference-time Hybrid Information Structurization

ICLR 2025Poster
6.8
38

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

ICLR 2025Poster
6.6
11

ConText: Driving In-context Learning for Text Removal and Segmentation

ICML 2025Poster
6.4
8

Benchmarking Agentic Workflow Generation

ICLR 2025Poster
6.3
26

Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent

ICLR 2025Poster
5.8
43

MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

ICLR 2025Rejected
5.7
22

Extend Model Merging from Fine-Tuned to Pre-Trained Large Language Models via Weight Disentanglement

ICLR 2025Rejected
5.5
34

Advancing Language Multi-Agent Learning with Credit Re-Assignment for Interactive Environment Generalization

COLM 2025Poster
5.0
37

ExploraCoder: Advancing code generation for multiple unseen APIs via planning and chained exploration

ICLR 2025Withdrawn
5.0
22

In-Context Transfer Learning: Demonstration Synthesis by Transferring Similar Tasks

ICLR 2025Rejected
4.9
13

LaRA: Benchmarking Retrieval-Augmented Generation and Long-Context LLMs – No Silver Bullet for LC or RAG Routing

ICML 2025Poster
4.8
11

Exploiting Presentative Feature Distributions for Parameter-Efficient Continual Learning of Large Language Models

ICML 2025Poster
4.3
4

Exploring Knowledge Boundaries in Large Language Models for Retrieval Judgment

ICLR 2025Withdrawn
通讯
4.3
10

Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?

ICLR 2025Rejected
3.7
5

Enabling Weak LLMs to Judge Response Reliability via Meta Ranking

ICLR 2025Rejected
3.0
6

Enhancing Multi-Agent Learning in Real-World Interactive Environments through Process Reward Decomposition

ICLR 2025Rejected