影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
76.56/100
前 1.7%
全站排名 #1,112
发表论文14 篇
平均评分
年均产出4.7 篇/年
Maksym Andriushchenko
研究方向
adversarial robustness · alignment · AI safety · generalization
19
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
ICLR 2026Poster
22
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
ICLR 2026Poster
25
Capability-Based Scaling Trends for LLM-Based Red-Teaming
ICLR 2026Poster
三作32
Monitoring Decomposition Attacks with Lightweight Sequential Monitors
ICLR 2026Poster
19
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
ICLR 2025Poster
一作41
Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
ICLR 2025Poster
一作21
Does Refusal Training in LLMs Generalize to the Past Tense?
ICLR 2025Poster
一作32
Critical Influence of Overparameterization on Sharpness-aware Minimization
ICLR 2025Rejected
三作20
Is In-Context Learning Sufficient for Instruction Following in LLMs?
ICLR 2025Poster
二作