影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
90.97/100
前 0.5%
全站排名 #328
发表论文26 篇
平均评分
年均产出8.7 篇/年
Xuandong Zhao
研究方向
Machine Learning · Natural Language Processing · AI Safety
10
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
ICLR 2026Poster
20
Learning to Reason without External Rewards
ICLR 2026Poster
一作21
In-Context Watermarks for Large Language Models
ICLR 2026Poster
二作19
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
ICLR 2026Poster
三作12
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
ICLR 2026Rejected
二作15
InfoSynth: Information-Guided Benchmark Synthesis for LLMs
ICLR 2026Rejected
三作16
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
ICLR 2026Rejected
三作16
PromptArmor: An Essential Baseline for Prompt Injection Defenses
ICLR 2026Rejected
5
Confidence-Guided MCTS for Efficient Long-Horizon Web Agent Tasks
ICLR 2026Rejected
二作54
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
COLM 2025Poster
24
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
ICLR 2025Poster
23
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
COLM 2025Poster
22
An Undetectable Watermark for Generative Image Models
ICLR 2025Poster
二作16
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
ICLR 2025Poster
一作35
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
NeurIPS 2025Poster
二作27
Multimodal Situational Safety
ICLR 2025Poster
三作11
Weak-to-Strong Jailbreaking on Large Language Models
ICML 2025Poster
一作11
Improving LLM Safety Alignment with Dual-Objective Optimization
ICML 2025Poster
一作41
Weak-to-Strong Jailbreaking on Large Language Models
ICLR 2025Rejected
一作12
DIS-CO: Discovering Copyrighted Content in VLMs Training Data
ICML 2025Poster
二作6
ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World
ICLR 2025Withdrawn
通讯9
Efficiently Identifying Watermarked Segments in Mixed-Source Texts
ICLR 2025Withdrawn
一作