影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
95.06/100
前 0.3%
全站排名 #169
发表论文64 篇
平均评分
年均产出21.3 篇/年
Pin-Yu Chen
研究方向
adversarial machine learning · trustworthy machine learning · adversarial robustness · machine learning and security · AI safety and alignment
22
Large Reasoning Models Learn Better Alignment from Flawed Thinking
ICLR 2026Rejected
26
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
ICLR 2026Poster
19
Detective SAM: Adaptive AI-Image Forgery Localization
ICLR 2026Poster
19
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
ICLR 2026Rejected
25
TrustGen: A Platform of Dynamic Benchmarking on the Trustworthiness of Generative Foundation Models
ICLR 2026Poster
6
Induction Head Implementation Across Diverse Transformer Weight Constructions
ICLR 2026Rejected
16
Emergent Deceptive Behaviors in Reward-Optimizing LLMs
ICLR 2026Desk Rejected
17
Aegis: Towards Governance, Integrity, and Security of AI Voice Agents
ICLR 2026Rejected
二作5
SpeechWakBench: How Well do Large Language Models Speak with Watermarks?
ICLR 2026Withdrawn
15
Intermediate Representations are Strong Training-Free AI-Generated Image Detectors
ICLR 2026Rejected
二作10
Attention Slipping: A Mechanistic Understanding of Jailbreak Attacks and Defenses in LLMs
ICLR 2026Rejected
二作4
Diversity Boosts AI-Generated Text Detection
ICLR 2026Withdrawn
二作35
Patching LLM like Software: A Lightweight Method for improving existing policy in Large Language Models
ICLR 2026Rejected
9
OjaKV: Context-Aware Online Low-Rank KV Cache Compression with Oja’s Rule
ICLR 2026Withdrawn
通讯4
CarBoN: Calibrated Best-of-N Sampling Improves Test-time Reasoning
ICLR 2026Withdrawn
二作23
Shape it Up! Restoring LLM Safety during Finetuning
NeurIPS 2025Poster
二作25
When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
ICLR 2025Oral
32
TabWak: A Watermark for Tabular Diffusion Models
ICLR 2025Spotlight
28
Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge
ICLR 2025Poster
25
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
ICLR 2025Poster
三作21
Adaptive Distraction: Probing LLM Contextual Robustness with Automated Tree Search
NeurIPS 2025Poster
21
Revisiting Mode Connectivity in Neural Networks with Bezier Surface
ICLR 2025Poster
二作23
Large Language Models can Become Strong Self-Detoxifiers
ICLR 2025Poster
二作28
CoP: Agentic Red-teaming for Large Language Models using Composition of Principles
NeurIPS 2025Poster
二作22
SEAL: Safety-enhanced Aligned LLM Fine-tuning via Bilevel Data Selection
ICLR 2025Poster
二作30
ADAPT: Adaptive Prompt Tuning for Pre-Trained Vision-Language Models
ICLR 2025Rejected
三作25
Test Time Augmentations are Worth One Million Images for Out-of-Distribution Detection
ICLR 2025Rejected
三作7
DAG-Jailbreak: Enhancing Black-box Jailbreak Attacks and Defenses through DAG Dependency Analysis
ICLR 2025Rejected
19
Your Task May Vary: A Systematic Understanding of Alignment and Safety Degradation when Fine-tuning LLMs
ICLR 2025Rejected
49
REFINE: Inversion-Free Backdoor Defense via Model Reprogramming
ICLR 2025Poster
17
Sparse Gradient Compression for Fine-Tuning Large Language Models
ICLR 2025Withdrawn
通讯32
Language Models Are Good Tabular Learners
ICLR 2025Rejected
15
Data-Driven Lipschitz Continuity: A Cost-Effective Approach to Improve Adversarial Robustness
ICLR 2025Rejected
二作5
Defensive Prompt Patch: A Robust and Generalizable Defense of Large Language Models against Jailbreak Attacks
ICLR 2025Withdrawn
三作5
SONAR: A Synthetic AI-Audio Detection Framework and Benchmark
ICLR 2025Withdrawn
二作32
Benchmarking LLMs on Safety Issues in Scientific Labs
ICLR 2025Rejected
5
Visual Prompting Reimagined: The Power of Activation Prompts
ICLR 2025Withdrawn
5
Breaking Free: Hacking Diffusion Models for Generating Adversarial Examples and Bypassing Safety Guardrails
ICLR 2025Rejected
三作5
GRE Score: Generative Risk Evaluation for Large Language Models
ICLR 2025Withdrawn
三作