影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
46.39/100
前 10.2%
全站排名 #6,546
发表论文18 篇
平均评分
年均产出6.0 篇/年
Yang Zhang
研究方向
Trustworthy Machine Leanring · AI Safety · Machine Learning Security · Memes · Social Network Analysis · Online Hate and Misinformation
16
Excessive Reasoning Attack on Reasoning LLMs
ICLR 2026Rejected
通讯22
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
ICLR 2026Rejected
通讯23
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
ICLR 2026Rejected
通讯6
Boosting Safety Alignment in LLMs with Response Shortcuts
ICLR 2026Rejected
5
IAAgent: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
ICLR 2026Withdrawn
通讯6
SOS! Soft Prompt Attack Against Open-Source Large Language Models
ICLR 2026Withdrawn
三作5
Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
ICLR 2026Withdrawn
通讯25
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
NeurIPS 2025Poster
45
Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
NeurIPS 2025Poster
通讯11
The Ripple Effect: On Unforeseen Complications of Backdoor Attacks
ICML 2025Poster
通讯22
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
ICLR 2025Poster
5
ACE: Attack Combo Enhancement Against Machine Learning Models
ICLR 2025Withdrawn
通讯