影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
59.48/100
前 4.9%
全站排名 #3,182
发表论文9 篇
平均评分
年均产出4.5 篇/年
Andy Zhou
研究方向
large language models · AI agents · AI security
15
AIR-BENCH 2024: A Safety Benchmark based on Regulation and Policies Specified Risk Categories
ICLR 2025Spotlight
三作24
MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models
ICLR 2025Poster
19
AutoRedTeamer: Autonomous Red Teaming with Lifelong Attack Integration
NeurIPS 2025Poster
一作34
Tamper-Resistant Safeguards for Open-Weight LLMs
ICLR 2025Poster
36
GUARD: Guideline Upholding Test through Adaptive Role-play and Jailbreak Diagnostics for LLMs
ICLR 2025Rejected
16
AutoRedTeamer: An Autonomous Red Teaming Agent Against Language Models
ICLR 2025Rejected
一作21
Robust Prompt Optimization for Defending Language Models Against Jailbreaking Attacks
NeurIPS 2024Spotlight
一作18
Jailbreaking Large Language Models Against Moderation Guardrails via Cipher Characters
NeurIPS 2024Poster
二作30
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
ICLR 2024Rejected
一作