影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
53.42/100
前 7%
全站排名 #4,480
发表论文23 篇
平均评分
年均产出7.7 篇/年
18
SecMCP: Quantifying Conversation Drift in MCP via Latent Polytope
ICLR 2026Rejected
11
Untargeted Jailbreak Attack
ICLR 2026Withdrawn
16
Explainable Token-level Noise Filtering for LLM Fine-tuning Datasets
ICLR 2026Poster
9
BadReward: Clean-Label Poisoning of Reward Models in Text-to-Image RLHF
ICLR 2026Withdrawn
16
Dynamic Target Attack
ICLR 2026Rejected
5
Your Large Reasoning Models Can Be Safer on Its Own
ICLR 2026Withdrawn
通讯5
Module-Aware Parameter-Efficient Machine Unlearning on Transformers
ICLR 2026Withdrawn
4
CARP: Causal Alignment of Reward Models via Response-to-Prompt Prediction
ICLR 2026Withdrawn
5
Towards Evaluation for Real-World LLM Unlearning
ICLR 2026Withdrawn
5
CollabMask: Explainable Neuron Collaboration with Gradient Masks for LLM Fine-Tuning
ICLR 2026Withdrawn
28
Towards Resilient Safety-driven Unlearning for Diffusion Models against Downstream Fine-tuning
NeurIPS 2025Poster
25
WMCopier: Forging Invisible Watermarks on Arbitrary Images
NeurIPS 2025Poster
25
Taught Well Learned Ill: Towards Distillation-conditional Backdoor Attack
NeurIPS 2025Poster
19
JudgeRail: Harnessing Open-Source LLMs for Fast Harmful Text Detection with Judicial Prompting and Logit Rectification
ICLR 2025Rejected
49
REFINE: Inversion-Free Backdoor Defense via Model Reprogramming
ICLR 2025Poster
28
Structure-Aware Parameter-Efficient Machine Unlearning on Transformer Models
ICLR 2025Rejected
5
Don’t Say No: Jailbreaking LLM by Suppressing Refusal
ICLR 2025Withdrawn
6
Prompt-Consistency Image Generation (PCIG): A Unified Framework Integrating LLMs, Knowledge Graphs, and Controllable Diffusion Models
ICLR 2025Withdrawn
三作36
Towards False-claim-resistant Model Ownership Verification via Targeted Fingerprint
ICLR 2025Rejected
7
A Comprehensive Deepfake Detector Assessment Platform
ICLR 2025Rejected
5
Certified PEFTSmoothing: Parameter-Efficient Fine-Tuning with Randomized Smoothing
ICLR 2025Withdrawn