影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
95.7/100
前 0.2%
全站排名 #153
发表论文72 篇
平均评分
年均产出24.0 篇/年
Caiming Xiong
研究方向
text summarization · Dialogue learning · self-supervised learning · deep learning for image classification · segmentation · question-answering · memory network · active learning · active clustering · image classification · action recognition
15
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
ICLR 2026Poster
10
Agentic Confidence Calibration
ICLR 2026Desk Rejected
二作14
Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels
ICLR 2026Poster
10
Learning to Reason over Continuous Tokens with Reinforcement Learning
ICLR 2026Poster
18
Entropy-Based Block Pruning for Efficient Large Language Models
ICLR 2026Poster
13
SWERank: Software Issue Localization with Code Ranking
ICLR 2026Poster
17
MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers
ICLR 2026Rejected
20
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
ICLR 2026Poster
16
CoAct-1: Computer-using Multi-agent System with Coding Actions
ICLR 2026Poster
通讯14
GTA1: GUI Test-time Scaling Agent
ICLR 2026Poster
14
Test-Time Adaptation for LLM Agents via Environment Interaction
ICLR 2026Poster
通讯16
WALT: Web Agents that Learn Tools
ICLR 2026Poster
28
SCUBA: Salesforce Computer Use Benchmark
ICLR 2026Poster
20
Scalable Chain of Thoughts via Elastic Reasoning
ICLR 2026Poster
通讯25
TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models
ICLR 2026Poster
13
Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation
ICLR 2026Poster
30
UniDoc-Bench: A Unified Benchmark for Document-Centric Multimodal RAG
ICLR 2026Withdrawn
30
SSR: Socratic Self-Refine for Large Language Model Reasoning
ICLR 2026Rejected
19
UserRL: Training Interactive User-Centric Agent via Reinforcement Learning
ICLR 2026Rejected
6
Enabling Tool Use of Reasoning Models Without Verifiable Reward via SFT-RL Loop
ICLR 2026Rejected
三作5
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
ICLR 2026Withdrawn
18
Fractured Chain-of-Thought Reasoning
ICLR 2026Rejected
通讯12
ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning
ICLR 2026Rejected
5
BLIP3-o: A Family of Fully Open Unified Multimodal Models—Architecture, Training and Dataset
ICLR 2026Withdrawn
19
GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness
ICLR 2026Rejected
6
Synthesizing Agentic Data for Web Agent Training with Progressive Difficulty Enhancement
ICLR 2026Withdrawn
5
LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering
ICLR 2026Desk Rejected
15
MAS-Zero: Designing Multi-Agent Systems with Zero Supervision
ICLR 2026Withdrawn
5
Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math
ICLR 2026Withdrawn
19
Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
ICLR 2025Oral
50
ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement
ICLR 2025Oral
22
AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials
ICLR 2025Spotlight
29
Automatic Curriculum Expert Iteration for Reliable LLM Reasoning
ICLR 2025Poster
23
DyMU: Dynamic Merging and Virtual Unmerging for Efficient Variable-Length VLMs
NeurIPS 2025Poster
三作23
ThinK: Thinner Key Cache by Query-Driven Pruning
ICLR 2025Spotlight
26
SiReRAG: Indexing Similar and Related Information for Multihop Reasoning
ICLR 2025Poster
16
GReaTer: Gradients Over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
ICLR 2025Poster
9
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
ICML 2025Poster
通讯31
CodeXEmbed: A Generalist Embedding Model Family for Multilingual and Multi-task Code Retrieval
COLM 2025Poster
27
BingoGuard: LLM Content Moderation Tools with Risk Levels
ICLR 2025Poster
18
Bridging the Data Provenance Gap Across Text, Speech, and Video
ICLR 2025Poster
20
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
NeurIPS 2025Poster
9
Reward-Guided Speculative Decoding for Efficient LLM Reasoning
ICML 2025Poster
通讯29
BLIP-3-Video: You Only Need 32 Tokens to Represent a Video Even in VLMs
ICLR 2025Rejected
22
Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents
ICLR 2025Poster
通讯13
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
ICML 2025Poster
21
LAM Simulator: Advancing Large Action Model Training for Agent via Online Exploration and Feedback Simulation
ICLR 2025Rejected
36
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
ICLR 2025Poster
39
Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
ICLR 2025Rejected
47
Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
ICLR 2025Rejected
通讯18
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
ICLR 2025Rejected
38
GIFT-Eval: A Benchmark for General Time Series Forecasting Model Evaluation
ICLR 2025Rejected
22
Trust but Verify: Programmatic VLM Evaluation in the Wild
ICLR 2025Rejected
33
Direct Judgement Preference Optimization
ICLR 2025Rejected
33
UniTST: Effectively Modeling Inter-Series and Intra-Series Dependencies for Multivariate Time Series Forecasting
ICLR 2025Rejected
12
Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
ICML 2025Poster
19
MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
ICLR 2025Rejected
7
MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs
ICLR 2025Withdrawn
8
Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
ICLR 2025Rejected
4
Expanding the Web, Smaller Is Better: A Comprehensive Study in Post-training
ICLR 2025Withdrawn