影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
88.52/100
前 0.7%
全站排名 #434
发表论文20 篇
平均评分
年均产出6.7 篇/年
Prateek Mittal
研究方向
Secure &Trustworthy Cyberspace · Adversarial Machine Learning · Privacy-Preserving Machine Learning
17
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
ICLR 2026Poster
通讯16
Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
ICLR 2026Poster
20
MURMUR: Using cross-user chatter to break collaborative language agents
ICLR 2026Rejected
4
Red-Teaming NSFW Image Classifiers as Text-to-Image Safeguards
ICLR 2026Withdrawn
29
Safety Alignment Should be Made More Than Just a Few Tokens Deep
ICLR 2025Oral
27
Capturing the Temporal Dependence of Training Data Influence
ICLR 2025Oral
30
Data Shapley in One Training Run
ICLR 2025Oral
二作20
ReliabilityRAG: Effective and Provably Robust Defense for RAG-based Web-Search
NeurIPS 2025Poster
24
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
ICLR 2025Poster
通讯21
Privacy Auditing of Large Language Models
ICLR 2025Poster
通讯30
On Evaluating the Durability of Safeguards for Open-Weight LLMs
ICLR 2025Poster
11
Adapting to Evolving Adversaries with Regularized Continual Robust Training
ICML 2025Poster
22
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
ICLR 2025Poster
15
Certifiably Robust RAG against Retrieval Corruption Attacks
ICLR 2025Rejected
通讯6
Lottery Ticket Adaptation: Mitigating Destructive Interference in LLMs
ICLR 2025Withdrawn
通讯