影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
82.33/100
前 1.2%
全站排名 #753
发表论文17 篇
平均评分
年均产出5.7 篇/年
He He
研究方向
natural language processing · large language models
20
Is it Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort
ICLR 2026Oral
通讯16
Jailbreak Transferability Emerges from Shared Representations
ICLR 2026Poster
三作17
Measuring LLM Novelty As The Frontier Of Original And High-Quality Output
ICLR 2026Poster
通讯24
LLM Watermark Evasion via Bias Inversion
ICLR 2026Rejected
二作32
Monitoring Decomposition Attacks with Lightweight Sequential Monitors
ICLR 2026Poster
通讯24
Unsupervised Elicitation of Language Models
ICLR 2026Rejected
18
Predicting Empirical AI Research Outcomes with Language Models
NeurIPS 2025Poster
16
Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats
ICLR 2025Poster
23
Transformers Struggle to Learn to Search
ICLR 2025Poster
通讯16
Reasoning Models Know When They’re Right: Probing Hidden States for Self-Verification
COLM 2025Poster
通讯20
Language Models Learn to Mislead Humans via RLHF
ICLR 2025Poster
15
Hyperparameter Loss Surfaces Are Simple Near their Optima
COLM 2025Poster
二作21
An Asymptotic Theory of Random Search for Hyperparameters in Deep Learning
ICLR 2025Rejected
二作