影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
69.66/100
前 2.7%
全站排名 #1,745
发表论文28 篇
平均评分
年均产出9.3 篇/年
Michael Backes
研究方向
Trustworthy AI · Security and Privacy
24
Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models
ICLR 2026Oral
16
Excessive Reasoning Attack on Reasoning LLMs
ICLR 2026Rejected
三作22
Benchmark of Benchmarks: Unpacking Influence and Code Repository Quality in LLM Safety Benchmarks
ICLR 2026Rejected
三作25
TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models
ICLR 2026Poster
16
SafeReview: Building a Robust Deep Review Assistant Against Prompt Injection
ICLR 2026Rejected
23
Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs
ICLR 2026Rejected
5
PRIVDISTIL: A Unified Framework for Accurate and Differentially Private Model Compression
ICLR 2026Rejected
6
Boosting Safety Alignment in LLMs with Response Shortcuts
ICLR 2026Rejected
三作5
IAAgent: Autonomous Inference Attacks Against ML Services With LLM-Based Agents
ICLR 2026Withdrawn
6
SOS! Soft Prompt Attack Against Open-Source Large Language Models
ICLR 2026Withdrawn
二作25
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
NeurIPS 2025Poster
三作45
Adjacent Words, Divergent Intents: Jailbreaking Large Language Models via Task Concurrency
NeurIPS 2025Poster
三作22
SaLoRA: Safety-Alignment Preserved Low-Rank Adaptation
ICLR 2025Poster
三作9
Provably Cost-Sensitive Adversarial Defense via Randomized Smoothing
ICML 2025Poster
三作13
Efficient and Privacy-Preserving Soft Prompt Transfer for LLMs
ICML 2025Poster
19
Captured by Captions: On Memorization and its Mitigation in CLIP Models
ICLR 2025Poster
30
POST: A Framework for Privacy of Soft-prompt Transfer
ICLR 2025Rejected
5
ACE: Attack Combo Enhancement Against Machine Learning Models
ICLR 2025Withdrawn