影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
98.34/100
前 0.1%
全站排名 #41
发表论文62 篇
平均评分
年均产出20.7 篇/年
25
CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale
ICLR 2026Oral
通讯17
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
ICLR 2026Poster
16
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
ICLR 2026Rejected
19
DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle
ICLR 2026Poster
21
AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
ICLR 2026Rejected
21
In-Context Watermarks for Large Language Models
ICLR 2026Poster
20
Learning to Reason without External Rewards
ICLR 2026Poster
通讯22
Investigating the Link Between Representational Similarity and Model Interactions
ICLR 2026Rejected
14
RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?
ICLR 2026Poster
通讯24
Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
ICLR 2026Poster
20
Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
ICLR 2026Poster
通讯15
Usefulness-driven Learning of Formal Mathematics
ICLR 2026Rejected
通讯16
RepIt: Steering Language Models with Concept-Specific Refusal Vectors
ICLR 2026Poster
31
Scaling Agent Learning via Experience Synthesis
ICLR 2026Poster
19
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
ICLR 2026Poster
通讯20
VERINA: Benchmarking Verifiable Code Generation
ICLR 2026Poster
通讯11
SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers
ICLR 2026Rejected
二作12
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
ICLR 2026Rejected
三作19
Can Aha Moments Be Fake? Identifying True and Decorative Thinking Steps in Chain-of-Thought
ICLR 2026Rejected
通讯14
MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them
ICLR 2026Rejected
通讯19
RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents
ICLR 2026Poster
16
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
ICLR 2026Rejected
通讯15
InfoSynth: Information-Guided Benchmark Synthesis for LLMs
ICLR 2026Rejected
通讯5
WebGuard: Building a Generalizable Guardrail for Web Agents
ICLR 2026Withdrawn
12
Federated Agent Reinforcement Learning
ICLR 2026Rejected
通讯16
PromptArmor: An Essential Baseline for Prompt Injection Defenses
ICLR 2026Rejected
通讯4
AgentXploit: End-to-End Red-Teaming for AI Agents Powdered by Multi-Agent Systems
ICLR 2026Withdrawn
通讯8
CTIArena: Benchmarking LLM Knowledge and Reasoning across Heterogeneous Cyber Threat Intelligence
ICLR 2026Rejected
27
Capturing the Temporal Dependence of Training Data Influence
ICLR 2025Oral
二作18
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
NeurIPS 2025Poster
通讯30
Data Shapley in One Training Run
ICLR 2025Oral
三作54
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
COLM 2025Poster
15
AIR-BENCH 2024: A Safety Benchmark based on Regulation and Policies Specified Risk Categories
ICLR 2025Spotlight
24
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
ICLR 2025Poster
通讯25
An Illusion of Progress? Assessing the Current State of Web Agents
COLM 2025Poster
23
LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage
COLM 2025Poster
通讯22
An Undetectable Watermark for Generative Image Models
ICLR 2025Poster
三作35
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
NeurIPS 2025Poster
三作27
Multimodal Situational Safety
ICLR 2025Poster
11
GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning
ICML 2025Poster
29
GuardAgent: Safeguard LLM Agent by a Guard Agent via Knowledge-Enabled Reasoning
ICLR 2025Rejected
34
Tamper-Resistant Safeguards for Open-Weight LLMs
ICLR 2025Poster
20
KnowHalu: Multi-Form Knowledge Enhanced Hallucination Detection
ICLR 2025Rejected
29
KnowData: Knowledge-Enabled Data Generation for Improving Multimodal Models
ICLR 2025Rejected
19
AutoScale: Automatic Prediction of Compute-optimal Data Compositions for Training LLMs
ICLR 2025Rejected
11
Improving LLM Safety Alignment with Dual-Objective Optimization
ICML 2025Poster
通讯26
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
COLM 2025Poster
13
Assessing the Knowledge-intensive Reasoning Capability of Large Language Models with Realistic Benchmarks Generated Programmatically at Scale
ICLR 2025Rejected
通讯14
MultiTrust: Enhancing Safety and Trustworthiness of Large Language Models from Multiple Perspectives
ICLR 2025Rejected
二作20
SecCodePLT: A Unified Platform for Evaluating the Security of Code GenAI
ICLR 2025Rejected
通讯7
Can Editing LLMs Inject Harm?
ICLR 2025Rejected
5
Which Network is Trojaned? Increasing Trojan Evasiveness for Model-Level Detectors
ICLR 2025Withdrawn
15
IDS-Agent: An LLM Agent for Explainable Intrusion Detection in IoT Networks
ICLR 2025Rejected