影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
85.51/100
前 0.9%
全站排名 #572
发表论文23 篇
平均评分
年均产出7.7 篇/年
10
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
ICLR 2026Poster
三作14
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against LLM Jailbreaks and Prompt Injections
ICLR 2026Rejected
通讯16
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
ICLR 2026Poster
三作18
Auditing Agents for Adversarial Fine-tuning Detection
ICLR 2026Desk Rejected
三作16
Poisoning Attacks on LLMs Require a Near-constant Number of Poison Samples
ICLR 2026Rejected
8
AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example Defenses
ICML 2025Oral
一作20
Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI
ICLR 2025Spotlight
三作33
Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
ICLR 2025Poster
25
Scalable Extraction of Training Data from Aligned, Production Language Models
ICLR 2025Poster
三作30
On Evaluating the Durability of Safeguards for Open-Weight LLMs
ICLR 2025Poster
三作7
Exploring and Mitigating Adversarial Manipulation of Voting-Based Leaderboards
ICML 2025Oral
21
AutoAdvExBench: Benchmarking Autonomous Exploitation of Adversarial Example Defenses
ICLR 2025Rejected
一作22
IF-Guide: Influence Function-Guided Detoxification of LLMs
NeurIPS 2025Poster
三作32
Evaluating Privacy Risks of Parameter-Efficient Fine-Tuning
ICLR 2025Rejected
二作28
Persistent Pre-training Poisoning of LLMs
ICLR 2025Poster
5
Certified Robustness to Clean-label Poisoning Using Diffusion Denoising
ICLR 2025Withdrawn
二作9
Stealing User Prompts from Mixture-of-Experts Models
ICLR 2025Rejected
通讯