影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
97.55/100
前 0.1%
全站排名 #71
发表论文45 篇
平均评分
年均产出15.0 篇/年
Min-hwan Oh
研究方向
Bandit Algorithms · Reinforcement Learning
14
Convergence of Muon with Newton-Schulz
ICLR 2026Poster
二作16
Peng's Q($\lambda$) for Conservative Value Estimation in Offline Reinforcement Learning
ICLR 2026Poster
二作18
KL-Regularization Is Sufficient in Contextual Bandits and RLHF
ICLR 2026Withdrawn
二作24
Optimal Batched (Generalized) Linear Contextual Bandit Algorithm
ICLR 2026Rejected
二作14
Offline Preference-Based Value Optimization
ICLR 2026Poster
二作15
Diversified Multinomial Logit Contextual Bandits
ICLR 2026Poster
三作13
Batched Stochastic Matching Bandits
ICLR 2026Rejected
二作13
Blessings of Many Good Arms in Multi-Objective Linear Bandits
ICLR 2026Withdrawn
二作15
Follow-the-Perturbed-Leader for Decoupled Bandits: Best-of-Both-Worlds and Practicality
ICLR 2026Withdrawn
三作12
Sample-Efficient Pruning Model Selection via Lasso
ICLR 2026Withdrawn
三作4
Inverse GFlowNets for Generative Imitation Learning
ICLR 2026Withdrawn
27
Exploration via Feature Perturbation in Contextual Bandits
NeurIPS 2025Spotlight
二作23
True Impact of Cascade Length in Contextual Cascading Bandits
NeurIPS 2025Poster
三作23
Infrequent Exploration in Linear Bandits
NeurIPS 2025Poster
二作22
Tractable Multinomial Logit Contextual Bandits with Non-Linear Utilities
NeurIPS 2025Poster
三作21
Revisiting Follow-the-Perturbed-Leader with Unbounded Perturbations in Bandit Problems
NeurIPS 2025Poster
通讯11
Improved Online Confidence Bounds for Multinomial Logistic Bandits
ICML 2025Poster
二作25
Thompson Sampling for Multi-Objective Linear Contextual Bandit
NeurIPS 2025Poster
三作27
Minimax Optimal Reinforcement Learning with Quasi-Optimism
ICLR 2025Poster
二作17
Adversarial Policy Optimization for Offline Preference-based Reinforcement Learning
ICLR 2025Poster
二作15
Oracle-Efficient Combinatorial Semi-Bandits
NeurIPS 2025Poster
三作29
Preference-based Reinforcement Learning beyond Pairwise Comparisons: Benefits of Multiple Options
NeurIPS 2025Poster
三作16
ADAM Optimization with Adaptive Batch Selection
ICLR 2025Poster
二作16
Dynamic Assortment Selection and Pricing with Censored Preference Feedback
ICLR 2025Poster
二作29
EUGens: Efficient, Unified and General Dense Layers
NeurIPS 2025Poster
17
Lasso Bandit with Compatibility Condition on Optimal Arm
ICLR 2025Poster
三作11
Combinatorial Reinforcement Learning with Preference Feedback
ICML 2025Poster
二作13
Symmetry-Aware GFlowNets
ICML 2025Poster
三作30
GFlowNets Need Automorphism Correction for Unbiased Graph Generation
ICLR 2025Rejected
三作32
Combinatorial Reinforcement Learning with Preference Feedback
ICLR 2025Rejected
二作10
Optimal and Practical Batched Linear Bandit Algorithm
ICML 2025Poster
二作21
Magnituder Layers for Implicit Neural Representations in 3D
ICLR 2025Rejected
通讯27
Linear Bandits with Partially Observable Features
ICLR 2025Rejected
通讯23
Stochastic Matching Bandits under Preference Feedback
ICLR 2025Withdrawn
二作9
Linear Bandits with Partially Observable Features
ICML 2025Poster
通讯9
Neural Dynamic Pricing: Provable and Practical Efficiency
ICLR 2025Withdrawn
通讯14
Mostly Exploration-free Algorithms for Multi-Objective Linear Bandits
ICLR 2025Withdrawn
二作-1
Coordinated Exploration in Distributed Reinforcement Learning
ICLR 2025Withdrawn
二作