影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
63.61/100
前 3.8%
全站排名 #2,472
发表论文14 篇
平均评分
年均产出7.0 篇/年
20
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
ICLR 2026Poster
二作24
Propaganda AI: An Analysis of Semantic Divergence in Large Language Models
ICLR 2026Poster
三作10
Where Did It Go Wrong? Attributing Undesirable LLM Behaviors via Representation Gradient Tracing
ICLR 2026Poster
三作5
AgenticPA: Toward Automated and Large-Scale Prompt Attacks on LLMs
ICLR 2026Withdrawn
20
AutoBackdoor: Automating Backdoor Attacks via LLM Agents
ICLR 2026Rejected
一作21
Memory Injection Attacks on LLM Agents via Query-Only Interaction
NeurIPS 2025Poster
9
X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
ICML 2025Poster
三作11
CROW: Eliminating Backdoors from Large Language Models via Internal Consistency Regularization
ICML 2025Poster
三作20
Detecting Backdoor Samples in Contrastive Language Image Pretraining
ICLR 2025Poster
三作48
BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
ICLR 2025Poster
5
Adversarial Suffixes May Be Features Too!
ICLR 2025Withdrawn
三作5
AnyAttack: Self-supervised Generation of Targeted Adversarial Attacks for Vision-Language Models
ICLR 2025Withdrawn
33
TUAP: Targeted Universal Adversarial Perturbations for CLIP
ICLR 2025Rejected
三作5
Do Influence Functions Work on Large Language Models?
ICLR 2025Withdrawn
三作