影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
-/100
发表论文4 篇
平均评分
年均产出4.0 篇/年
Yushi Yang
研究方向
Large Language Models · Mechanistic Interpretability · Technical AI Safety · LLM Post-training · LLM agents
21
Agentic reinforcement learning for search is unsafe
ICLR 2026Rejected
一作17
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
ICLR 2026Rejected
二作12
To Distill or Not to Distill: Knowledge Transfer Undermines Safety of LLMs
ICLR 2026Withdrawn
二作5
Collapse of Irrelevant Representations (CIR) Ensures Robust and Non-Disruptive LLM Unlearning
ICLR 2026Withdrawn
二作