影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
42.98/100
前 12.3%
全站排名 #7,946
发表论文11 篇
平均评分
年均产出3.7 篇/年
16
TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them
ICLR 2026Poster
22
Deep Literature Survey Automation with an Iterative Workflow
ICLR 2026Rejected
37
Rethinking LLM Evaluation: Can We Evaluate LLMs with 200× Less Data?
ICLR 2026Poster
21
AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
ICLR 2026Withdrawn
21
RLAR: An Agentic Reward System for Multi-task Reinforcement Learning on Large Language Models
ICLR 2026Withdrawn
二作15
Detecting Variant Contamination in LLMs via Variance of Generation Distribution
ICLR 2026Withdrawn
通讯5
PROBE: Benchmarking Reasoning Paradigm Overfitting in Large Language Models
ICLR 2026Withdrawn
二作14
SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models
ICLR 2025Poster
三作27
NovelQA: Benchmarking Question Answering on Documents Exceeding 200K Tokens
ICLR 2025Poster
一作27
RedHat: Towards Reducing Hallucination in Essay Critiques with Large Language Models
ICLR 2025Rejected
二作