影响力指数
96.86/100
前 0.2%
全站排名 #102
发表论文53
平均评分5.3
年均产出17.7 篇/年

Yang Yu

Professor@Nanjing University·中国·OpenReview
研究方向

reinforcement learning · derivative-free optimization · ensemble learning

7.8
20

Focus-Then-Reuse: Fast Adaptation in Visual Perturbation Environments

NeurIPS 2025Poster
通讯
6.8
19

Adaptable Safe Policy Learning from Multi-task Data with Constraint Prioritized Decision Transformer

NeurIPS 2025Poster
通讯
6.6
13

LLM-Assisted Semantically Diverse Teammate Generation for Efficient Multi-agent Coordination

ICML 2025Poster
通讯
6.5
32

Efficient Multi-agent Offline Coordination via Diffusion-based Trajectory Stitching

ICLR 2025Poster
通讯
6.4
25

Uncertainty-Sensitive Privileged Learning

NeurIPS 2025Poster
三作
6.3
26

On the Optimization Landscape of Low Rank Adaptation Methods for Large Language Models

ICLR 2025Poster
通讯
6.3
9

Behavior-Regularized Diffusion Policy Optimization for Offline Reinforcement Learning

ICML 2025Poster
6.0
14

Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement Learning

ICLR 2025Poster
通讯
6.0
16

Q-Adapter: Customizing Pre-trained LLMs to New Preferences with Forgetting Mitigation

ICLR 2025Poster
6.0
8

Any-step Dynamics Model Improves Future Predictions for Online and Offline Reinforcement Learning

ICLR 2025Poster
通讯
6.0
36

SOO-Bench: Benchmarks for Evaluating the Stability of Offline Black-Box Optimization

ICLR 2025Poster
通讯
5.5
37

Safe Multi-task Pretraining with Constraint Prioritized Decision Transformer

ICLR 2025Rejected
通讯
5.5
13

Improving Reward Model Generalization from Adversarial Process Enhanced Preferences

ICML 2025Poster
通讯
5.5
62

LLMOPT: Learning to Define and Solve General Optimization Problems from Scratch

ICLR 2025Poster
通讯
5.5
7

Controlling Large Language Model with Latent Action

ICML 2025Poster
通讯
5.5
26

Multi-Agent Imitation by Learning and Sampling from Factorized Soft Q-Function

NeurIPS 2025Poster
通讯
5.5
7

Learning to Reuse Policies in State Evolvable Environments

ICML 2025Poster
通讯
5.4
30

Learning View-invariant World Models for Visual Robotic Manipulation

ICLR 2025Poster
通讯
5.2
19

Hindsight Preference Learning for Offline Preference-based Reinforcement Learning

ICLR 2025Rejected
4.8
7

Haland: Human-AI Coordination via Policy Generation from Language-guided Diffusion

ICLR 2025Rejected
通讯
3.7
4

Boosting Offline Multi-Objective Reinforcement Learning via Preference Conditioned Diffusion Models

ICLR 2025Withdrawn
通讯
3.0
9

Learning Generalizable Environment Models via Discovering Superposed Causal Relationships

ICLR 2025Rejected
3.0
5

Diffusion-Guided Safe Policy Optimization From Cost-Label-Free Offline Dataset

ICLR 2025Withdrawn
通讯
-1

Whale-X: Learning Scalable Embodied World Models with Enhanced Generalizability

ICLR 2025Withdrawn
通讯