影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
74.7/100
前 1.9%
全站排名 #1,242
发表论文26 篇
平均评分
年均产出8.7 篇/年
Souradip Chakraborty
研究方向
Uncertainty & Bayesian methods for Foundational Models · Reinforcement Learning from Human Feedback · Large Language Models & Generative Model Alignment · Deep Reinforcement Learning · Bayesian Optimization · Multimodal learning and Image Captioning · Representation Learning · Language Models and Information Retrieval
20
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach
ICLR 2026Poster
二作14
Repair Aware Forgetting: An Iterative Approach to Unlearning in T2I Diffusion Models
ICLR 2026Desk Rejected
15
Post-training Large Language Models for Diverse High-Quality Responses
ICLR 2026Poster
二作27
TEST-TIME SCALING IN DIFFUSION LLMS VIA HIDDEN SEMI-AUTOREGRESSIVE EXPERTS
ICLR 2026Poster
14
Cut the Overcredit: Precision First Process Rewards for Reasoning LLMs
ICLR 2026Rejected
二作26
Multi-Level Multi-Turn RL Outperforms GRPO: Reasoning with Textual Feedback
ICLR 2026Rejected
三作29
TRAM: Test-time Risk Adaptation with Mixture of Agents
ICLR 2026Rejected
三作25
Mitigating Reward Hacking in Inference-Time Alignment of T2I Diffusion Models via Distributional Regularization
ICLR 2026Rejected
11
HEART: Emotionally-driven test-time scaling of Language Models
ICLR 2026Rejected
13
A Principled Approach to Chain-of-Thought Monitorability in Reasoning Models
ICLR 2026Withdrawn
二作4
SafeThink: A Key to Safety in Multi-Modal Large Reasoning Models
ICLR 2026Withdrawn
三作23
Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning Models
NeurIPS 2025Poster
二作21
On the Global Optimality of Policy Gradient Methods in General Utility Reinforcement Learning
NeurIPS 2025Poster
二作31
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
ICLR 2025Poster
一作11
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
ICML 2025Poster
三作22
SAIL: Self-improving Efficient Online Alignment of Large Language Models
ICLR 2025Rejected
二作8
Hierarchical Preference Optimization: Learning to achieve goals via feasible subgoals prediction
ICLR 2025Withdrawn
二作8
Aligning Large Language Models With Preference Privacy
ICLR 2025Rejected
30
LIAR: Leveraging Inverse Alignment to Jailbreak LLMs in Seconds
ICLR 2025Rejected
二作26
On the Sample Complexity of a Policy Gradient Algorithm with Occupancy Approximation for General Utility Reinforcement Learning
ICLR 2025Rejected
二作5
DIPPER: Direct Preference Optimization for Primitive-Enabled Hierarchical Reinforcement Learning
ICLR 2025Withdrawn
二作5
AIME: AI System Optimization via Multiple LLM Evaluators
ICLR 2025Withdrawn
二作