影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
40.42/100
前 14.2%
全站排名 #9,110
发表论文9 篇
平均评分
年均产出4.5 篇/年
17
AdvPrefix: An Objective for Nuanced LLM Jailbreaks
NeurIPS 2025Poster
通讯7
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
ICLR 2025Rejected
28
Persistent Pre-training Poisoning of LLMs
ICLR 2025Poster
三作26
Gradient-based Jailbreak Images for Multimodal Fusion Models
ICLR 2025Rejected
10
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
ICML 2025Poster
7
Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
ICLR 2025Rejected
三作