影响力指数
95.7/100
前 0.2%
全站排名 #153
发表论文72
平均评分5.3
年均产出24.0 篇/年

Caiming Xiong

Research Scientist@Salesforce Research·美国·OpenReview
研究方向

text summarization · Dialogue learning · self-supervised learning · deep learning for image classification · segmentation · question-answering · memory network · active learning · active clustering · image classification · action recognition

6.7
15

Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains

ICLR 2026Poster
6.0
10

Agentic Confidence Calibration

ICLR 2026Desk Rejected
二作
6.0
14

Webscale-RL: Automated Data Pipeline for Scaling RL Data to Pretraining Levels

ICLR 2026Poster
5.5
10

Learning to Reason over Continuous Tokens with Reinforcement Learning

ICLR 2026Poster
5.5
18

Entropy-Based Block Pruning for Efficient Large Language Models

ICLR 2026Poster
5.5
13

SWERank: Software Issue Localization with Code Ranking

ICLR 2026Poster
5.5
17

MCP-Universe: Benchmarking Large Language Models with Real-World Model Context Protocol Servers

ICLR 2026Rejected
5.5
20

LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild

ICLR 2026Poster
5.5
16

CoAct-1: Computer-using Multi-agent System with Coding Actions

ICLR 2026Poster
通讯
5.5
14

GTA1: GUI Test-time Scaling Agent

ICLR 2026Poster
5.3
14

Test-Time Adaptation for LLM Agents via Environment Interaction

ICLR 2026Poster
通讯
5.0
16

WALT: Web Agents that Learn Tools

ICLR 2026Poster
4.8
28

SCUBA: Salesforce Computer Use Benchmark

ICLR 2026Poster
4.7
20

Scalable Chain of Thoughts via Elastic Reasoning

ICLR 2026Poster
通讯
4.7
25

TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models

ICLR 2026Poster
4.5
13

Scaling Knowledge Graph Construction through Synthetic Data Generation and Distillation

ICLR 2026Poster
4.5
30

UniDoc-Bench: A Unified Benchmark for Document-Centric Multimodal RAG

ICLR 2026Withdrawn
4.5
30

SSR: Socratic Self-Refine for Large Language Model Reasoning

ICLR 2026Rejected
4.5
19

UserRL: Training Interactive User-Centric Agent via Reinforcement Learning

ICLR 2026Rejected
4.0
6

Enabling Tool Use of Reasoning Models Without Verifiable Reward via SFT-RL Loop

ICLR 2026Rejected
三作
4.0
5

Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math

ICLR 2026Withdrawn
4.0
18

Fractured Chain-of-Thought Reasoning

ICLR 2026Rejected
通讯
4.0
12

ToolLibGen: Scalable Automatic Tool Creation and Aggregation for LLM Reasoning

ICLR 2026Rejected
4.0
5

BLIP3-o: A Family of Fully Open Unified Multimodal Models—Architecture, Training and Dataset

ICLR 2026Withdrawn
3.5
19

GUI-KV: Efficient GUI Agents via KV Cache with Spatio-Temporal Awareness

ICLR 2026Rejected
3.5
6

Synthesizing Agentic Data for Web Agent Training with Progressive Difficulty Enhancement

ICLR 2026Withdrawn
3.5
5

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

ICLR 2026Desk Rejected
3.0
15

MAS-Zero: Designing Multi-Agent Systems with Zero Supervision

ICLR 2026Withdrawn
2.5
5

Reasoning Curriculum: Bootstrapping Broad LLM Reasoning from Math

ICLR 2026Withdrawn
8.0
19

Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows

ICLR 2025Oral
7.5
50

ReGenesis: LLMs can Grow into Reasoning Generalists via Self-Improvement

ICLR 2025Oral
7.3
22

AgentTrek: Agent Trajectory Synthesis via Guiding Replay with Web Tutorials

ICLR 2025Spotlight
7.0
29

Automatic Curriculum Expert Iteration for Reliable LLM Reasoning

ICLR 2025Poster
6.8
23

DyMU: Dynamic Merging and Virtual Unmerging for Efficient Variable-Length VLMs

NeurIPS 2025Poster
三作
6.8
23

ThinK: Thinner Key Cache by Query-Driven Pruning

ICLR 2025Spotlight
6.8
26

SiReRAG: Indexing Similar and Related Information for Multihop Reasoning

ICLR 2025Poster
6.7
16

GReaTer: Gradients Over Reasoning Makes Smaller Language Models Strong Prompt Optimizers

ICLR 2025Poster
6.6
9

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

ICML 2025Poster
通讯
6.5
31

CodeXEmbed: A Generalist Embedding Model Family for Multilingual and Multi-task Code Retrieval

COLM 2025Poster
6.5
27

BingoGuard: LLM Content Moderation Tools with Risk Levels

ICLR 2025Poster
6.5
18

Bridging the Data Provenance Gap Across Text, Speech, and Video

ICLR 2025Poster
6.4
20

Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning

NeurIPS 2025Poster
6.3
9

Reward-Guided Speculative Decoding for Efficient LLM Reasoning

ICML 2025Poster
通讯
6.3
29

BLIP-3-Video: You Only Need 32 Tokens to Represent a Video Even in VLMs

ICLR 2025Rejected
6.3
22

Diversity Empowers Intelligence: Integrating Expertise of Software Engineering Agents

ICLR 2025Poster
通讯
6.1
13

Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators

ICML 2025Poster
6.0
21

LAM Simulator: Advancing Large Action Model Training for Agent via Online Exploration and Feedback Simulation

ICLR 2025Rejected
5.8
36

FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"

ICLR 2025Poster
5.5
39

Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts

ICLR 2025Rejected
5.5
47

Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

ICLR 2025Rejected
通讯
5.3
18

JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking

ICLR 2025Rejected
5.3
38

GIFT-Eval: A Benchmark for General Time Series Forecasting Model Evaluation

ICLR 2025Rejected
5.0
22

Trust but Verify: Programmatic VLM Evaluation in the Wild

ICLR 2025Rejected
5.0
33

Direct Judgement Preference Optimization

ICLR 2025Rejected
5.0
33

UniTST: Effectively Modeling Inter-Series and Intra-Series Dependencies for Multivariate Time Series Forecasting

ICLR 2025Rejected
4.9
12

Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts

ICML 2025Poster
4.5
19

MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases

ICLR 2025Rejected
4.2
7

MathHay: An Automated Benchmark for Long-Context Mathematical Reasoning in LLMs

ICLR 2025Withdrawn
3.8
8

Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding

ICLR 2025Rejected
3.7
4

Expanding the Web, Smaller Is Better: A Comprehensive Study in Post-training

ICLR 2025Withdrawn