影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
84.59/100
前 1%
全站排名 #626
发表论文33 篇
平均评分
年均产出11.0 篇/年
Shafiq Joty
研究方向
Large Language Models · Reinforcement Learning · Deep Learning
15
Foundational Automatic Evaluators: Scaling Multi-Task Generative Evaluator Training for Reasoning-Centric Domains
ICLR 2026Poster
通讯13
SWERank: Software Issue Localization with Code Ranking
ICLR 2026Poster
通讯20
LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild
ICLR 2026Poster
通讯18
References Improve LLM Alignment in Non-Verifiable Domains
ICLR 2026Poster
22
Variation in Verification: Understanding Verification Dynamics in Large Language Models
ICLR 2026Poster
通讯13
On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization
ICLR 2026Poster
通讯12
Gradually Compacting Large Language Models for Reasoning Like a Boiling Frog
ICLR 2026Rejected
6
Enabling Tool Use of Reasoning Models Without Verifiable Reward via SFT-RL Loop
ICLR 2026Rejected
通讯16
WAFER-QA: Evaluating Vulnerabilities of Agentic Workflows with Agent-as-Judge
ICLR 2026Rejected
通讯5
Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
ICLR 2026Withdrawn
通讯6
Synthesizing Agentic Data for Web Agent Training with Progressive Difficulty Enhancement
ICLR 2026Withdrawn
通讯15
MAS-Zero: Designing Multi-Agent Systems with Zero Supervision
ICLR 2026Withdrawn
通讯17
Preference Optimization for Reasoning with Pseudo Feedback
ICLR 2025Spotlight
31
CodeXEmbed: A Generalist Embedding Model Family for Multilingual and Multi-task Code Retrieval
COLM 2025Poster
三作20
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning
NeurIPS 2025Poster
29
The Emergence of Abstract Thought in Large Language Models Beyond Any Language
NeurIPS 2025Poster
13
Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
ICML 2025Poster
通讯37
Beyond In-Context Learning: Enhancing Long-form Generation of Large Language Models via Task-Inherent Attribute Guidelines
ICLR 2025Rejected
36
FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
ICLR 2025Poster
通讯18
JudgeRank: Leveraging Large Language Models for Reasoning-Intensive Reranking
ICLR 2025Rejected
二作33
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction
ICLR 2025Rejected
通讯33
Direct Judgement Preference Optimization
ICLR 2025Rejected
通讯12
Investigating Factuality in Long-Form Text Generation: The Roles of Self-Known and Self-Unknown
ICLR 2025Rejected
三作5
Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks
ICLR 2025Withdrawn
4
Expanding the Web, Smaller Is Better: A Comprehensive Study in Post-training
ICLR 2025Withdrawn
通讯