影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
58.31/100
前 5.2%
全站排名 #3,373
发表论文7 篇
平均评分
年均产出2.3 篇/年
Alexander Panfilov
研究方向
jailbreaking attacks on LLMs
22
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
ICLR 2026Poster
一作19
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
ICLR 2026Poster
二作25
Capability-Based Scaling Trends for LLM-Based Red-Teaming
ICLR 2026Poster
一作15
ASIDE: Architectural Separation of Instructions and Data in Language Models
ICLR 2026Poster
三作