影响力指数
98.34/100
前 0.1%
全站排名 #41
发表论文62
平均评分5.4
年均产出20.7 篇/年

Dawn Song

Full Professor@University of California Berkeley·美国·OpenReview
研究方向

Security · AI

7.0
25

CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale

ICLR 2026Oral
通讯
6.5
17

Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice

ICLR 2026Poster
6.0
16

SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs

ICLR 2026Rejected
6.0
19

DevOps-Gym: Benchmarking AI Agents in Software DevOps Cycle

ICLR 2026Poster
5.6
21

AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond

ICLR 2026Rejected
5.5
21

In-Context Watermarks for Large Language Models

ICLR 2026Poster
5.5
20

Learning to Reason without External Rewards

ICLR 2026Poster
通讯
5.5
22

Investigating the Link Between Representational Similarity and Model Interactions

ICLR 2026Rejected
5.2
14

RL Grokking Recipe: How Does RL Unlock and Transfer New Algorithms in LLMs?

ICLR 2026Poster
通讯
5.2
24

Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation

ICLR 2026Poster
5.0
20

Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence

ICLR 2026Poster
通讯
5.0
15

Usefulness-driven Learning of Formal Mathematics

ICLR 2026Rejected
通讯
5.0
16

RepIt: Steering Language Models with Concept-Specific Refusal Vectors

ICLR 2026Poster
5.0
31

Scaling Agent Learning via Experience Synthesis

ICLR 2026Poster
4.7
19

AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents

ICLR 2026Poster
通讯
4.7
20

VERINA: Benchmarking Verifiable Code Generation

ICLR 2026Poster
通讯
4.5
11

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

ICLR 2026Rejected
二作
4.5
12

Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs

ICLR 2026Rejected
三作
4.5
19

Can Aha Moments Be Fake? Identifying True and Decorative Thinking Steps in Chain-of-Thought

ICLR 2026Rejected
通讯
4.5
14

MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them

ICLR 2026Rejected
通讯
4.5
19

RedCodeAgent: Automatic Red-teaming Agent against Diverse Code Agents

ICLR 2026Poster
4.0
16

Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs

ICLR 2026Rejected
通讯
4.0
15

InfoSynth: Information-Guided Benchmark Synthesis for LLMs

ICLR 2026Rejected
通讯
3.5
5

WebGuard: Building a Generalizable Guardrail for Web Agents

ICLR 2026Withdrawn
3.5
12

Federated Agent Reinforcement Learning

ICLR 2026Rejected
通讯
3.5
16

PromptArmor: An Essential Baseline for Prompt Injection Defenses

ICLR 2026Rejected
通讯
3.3
4

AgentXploit: End-to-End Red-Teaming for AI Agents Powdered by Multi-Agent Systems

ICLR 2026Withdrawn
通讯
2.7
8

CTIArena: Benchmarking LLM Knowledge and Reasoning across Heterogeneous Cyber Threat Intelligence

ICLR 2026Rejected
8.0
27

Capturing the Temporal Dependence of Training Data Influence

ICLR 2025Oral
二作
7.8
18

Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations

NeurIPS 2025Poster
通讯
7.5
30

Data Shapley in One Training Run

ICLR 2025Oral
三作
7.5
54

Assessing Judging Bias in Large Reasoning Models: An Empirical Study

COLM 2025Poster
7.5
15

AIR-BENCH 2024: A Safety Benchmark based on Regulation and Policies Specified Risk Categories

ICLR 2025Spotlight
7.0
24

MMDT: Decoding the Trustworthiness and Safety of Multimodal Foundation Models

ICLR 2025Poster
通讯
6.8
25

An Illusion of Progress? Assessing the Current State of Web Agents

COLM 2025Poster
6.8
23

LeakAgent: RL-based Red-teaming Agent for LLM Privacy Leakage

COLM 2025Poster
通讯
6.5
22

An Undetectable Watermark for Generative Image Models

ICLR 2025Poster
三作
6.4
35

Scalable Best-of-N Selection for Large Language Models via Self-Certainty

NeurIPS 2025Poster
三作
6.4
27

Multimodal Situational Safety

ICLR 2025Poster
6.1
11

GuardAgent: Safeguard LLM Agents via Knowledge-Enabled Reasoning

ICML 2025Poster
6.0
29

GuardAgent: Safeguard LLM Agent by a Guard Agent via Knowledge-Enabled Reasoning

ICLR 2025Rejected
5.8
34

Tamper-Resistant Safeguards for Open-Weight LLMs

ICLR 2025Poster
5.7
20

KnowHalu: Multi-Form Knowledge Enhanced Hallucination Detection

ICLR 2025Rejected
5.5
29

KnowData: Knowledge-Enabled Data Generation for Improving Multimodal Models

ICLR 2025Rejected
5.5
19

AutoScale: Automatic Prediction of Compute-optimal Data Compositions for Training LLMs

ICLR 2025Rejected
5.5
11

Improving LLM Safety Alignment with Dual-Objective Optimization

ICML 2025Poster
通讯
5.3
26

AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs

COLM 2025Poster
5.3
13

Assessing the Knowledge-intensive Reasoning Capability of Large Language Models with Realistic Benchmarks Generated Programmatically at Scale

ICLR 2025Rejected
通讯
5.0
14

MultiTrust: Enhancing Safety and Trustworthiness of Large Language Models from Multiple Perspectives

ICLR 2025Rejected
二作
5.0
20

SecCodePLT: A Unified Platform for Evaluating the Security of Code GenAI

ICLR 2025Rejected
通讯
4.4
7

Can Editing LLMs Inject Harm?

ICLR 2025Rejected
3.5
5

Which Network is Trojaned? Increasing Trojan Evasiveness for Model-Level Detectors

ICLR 2025Withdrawn
3.0
15

IDS-Agent: An LLM Agent for Explainable Intrusion Detection in IoT Networks

ICLR 2025Rejected