影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
94.87/100
前 0.3%
全站排名 #174
发表论文35 篇
平均评分
年均产出11.7 篇/年
17
Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
ICLR 2026Poster
24
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
ICLR 2026Poster
16
Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
ICLR 2026Poster
通讯18
Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance
ICLR 2026Poster
通讯18
CASPO: Confidence-aware Step-wise Preference Optimization for Reliable Reasoning in Large Language Models
ICLR 2026Rejected
通讯16
Inference-Time Personalized Safety Control via Paired Difference-in-Means Intervention
ICLR 2026Poster
二作16
Data Valuation and Selection in a Federated Model Marketplace
ICLR 2026Rejected
二作19
ARTS: Alleviating Hallucinations in Large Vision–Language Models via Redundancy-Aware Token Selection
ICLR 2026Rejected
三作6
AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Object Detection Training
ICLR 2026Withdrawn
5
Distilling Reasoning into Student LLMs: Local Naturalness for Selecting Teacher Data
ICLR 2026Withdrawn
三作27
Capturing the Temporal Dependence of Training Data Influence
ICLR 2025Oral
通讯30
Data Shapley in One Training Run
ICLR 2025Oral
通讯15
AIR-BENCH 2024: A Safety Benchmark based on Regulation and Policies Specified Risk Categories
ICLR 2025Spotlight
23
LLM Can be a Dangerous Persuader: Empirical Study of Persuasion Safety in Large Language Models
COLM 2025Poster
24
SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal
ICLR 2025Poster
21
Data-Centric Human Preference with Rationales for Direct Preference Alignment
COLM 2025Poster
通讯42
Mind Control through Causal Inference: Predicting Clean Images from Poisoned Data
ICLR 2025Poster
31
LLMs Can Plan Only If We Tell Them
ICLR 2025Poster
二作21
Probing Hidden Knowledge Holes in Unlearned LLMs
NeurIPS 2025Poster
通讯12
LLMs Can Reason Faster Only If We Let Them
ICML 2025Poster
13
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
ICML 2025Poster
通讯19
AutoScale: Automatic Prediction of Compute-optimal Data Compositions for Training LLMs
ICLR 2025Rejected
通讯26
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
COLM 2025Poster
通讯22
LLM Spark: Critical Thinking Evaluation of Large Language Models
ICLR 2025Rejected
37
Data-Centric Human Preference Optimization with Rationales
ICLR 2025Rejected
通讯15
SCOPE: Scalable and Adaptive Evaluation of Misguided Safety Refusal in LLMs
ICLR 2025Rejected
通讯18
Fast and Noise-Robust Diffusion Solvers for Inverse Problems: A Frequentist Approach
ICLR 2025Rejected
9
CONCORD: Concept-informed Diffusion for Dataset Distillation
ICLR 2025Withdrawn
三作