影响力指数
95.06/100
前 0.3%
全站排名 #169
发表论文64
平均评分5.2
年均产出21.3 篇/年

Pin-Yu Chen

Principal Researcher@International Business Machines·美国·OpenReview
研究方向

adversarial machine learning · trustworthy machine learning · adversarial robustness · machine learning and security · AI safety and alignment

7.8
23

Shape it Up! Restoring LLM Safety during Finetuning

NeurIPS 2025Poster
二作
7.5
25

When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers

ICLR 2025Oral
7.2
32

TabWak: A Watermark for Tabular Diffusion Models

ICLR 2025Spotlight
6.8
28

Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

ICLR 2025Poster
6.5
25

Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis

ICLR 2025Poster
三作
6.4
21

Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search

NeurIPS 2025Poster
6.0
21

Revisiting Mode Connectivity in Neural Networks with Bezier Surface

ICLR 2025Poster
二作
6.0
23

Large Language Models can Become Strong Self-Detoxifiers

ICLR 2025Poster
二作
6.0
28

CoP: Agentic Red-teaming for Large Language Models using Composition of Principles

NeurIPS 2025Poster
二作
5.8
22

SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection

ICLR 2025Poster
二作
5.5
30

ADAPT: Adaptive Prompt Tuning for Pre-Trained Vision-Language Models

ICLR 2025Rejected
三作
5.5
25

Test Time Augmentations are Worth One Million Images for Out-of-Distribution Detection

ICLR 2025Rejected
三作
5.5
7

DAG-Jailbreak: Enhancing Black-box Jailbreak Attacks and Defenses through DAG Dependency Analysis

ICLR 2025Rejected
5.3
19

Your Task May Vary: A Systematic Understanding of Alignment and Safety Degradation when Fine-tuning LLMs

ICLR 2025Rejected
5.3
49

REFINE: Inversion-Free Backdoor Defense via Model Reprogramming

ICLR 2025Poster
5.0
17

Sparse Gradient Compression for Fine-Tuning Large Language Models

ICLR 2025Withdrawn
通讯
5.0
32

Language Models Are Good Tabular Learners

ICLR 2025Rejected
4.8
15

Data-Driven Lipschitz Continuity: A Cost-Effective Approach to Improve Adversarial Robustness

ICLR 2025Rejected
二作
4.5
5

Defensive Prompt Patch: A Robust and Generalizable Defense of Large Language Models against Jailbreak Attacks

ICLR 2025Withdrawn
三作
4.3
5

SONAR: A Synthetic AI-Audio Detection Framework and Benchmark

ICLR 2025Withdrawn
二作
4.0
32

Benchmarking LLMs on Safety Issues in Scientific Labs

ICLR 2025Rejected
3.8
5

Visual Prompting Reimagined: The Power of Activation Prompts

ICLR 2025Withdrawn
3.7
5

Breaking Free: Hacking Diffusion Models for Generating Adversarial Examples and Bypassing Safety Guardrails

ICLR 2025Rejected
三作
3.5
5

GRE Score: Generative Risk Evaluation for Large Language Models

ICLR 2025Withdrawn
三作