影响力指数
论文质量、代表作、近期表现、广度与样本量置信度综合计算
92.09/100
前 0.4%
全站排名 #277
发表论文47 篇
平均评分
年均产出15.7 篇/年
Zhaoxiang Zhang
Full Professor@Institute of automation, Chinese academy of science, Chinese Academy of Sciences·中国·OpenReview
研究方向
Embodied Intelligence · Agent Learning · Autonomous driving · biological-inspired model · brain-inspired computing · deep neural networks · obejct detection · semantic segmentation · action recognition · person re-identification · scene analysis and understanding
13
Unified Vision-Language-Action Model
ICLR 2026Poster
通讯31
IF-VidCap: Can Video Caption Models Follow Instructions?
ICLR 2026Poster
11
YuE: Scaling Open Foundation Models for Long-Form Music Generation
ICLR 2026Poster
16
Uniform Discrete Diffusion with Metric Path for Video Generation
ICLR 2026Poster
16
DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
ICLR 2026Poster
通讯15
ResGen: Residual Diffusion Model for LiDAR-based Point Cloud Generation
ICLR 2026Rejected
通讯29
SpectraLLM: Uncovering the Ability of LLMs for Molecule Structure Elucidation from Multi-Spectra
ICLR 2026Poster
通讯16
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Method
ICLR 2026Poster
通讯16
Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
ICLR 2026Poster
通讯15
GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation
ICLR 2026Rejected
通讯23
MLLM-CL: Continual Learning for Multimodal Large Language Models
ICLR 2026Rejected
通讯32
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
ICLR 2026Poster
通讯26
HiPO: Hybrid Policy Optimization for Dynamic Reasoning in LLMs
ICLR 2026Withdrawn
44
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
ICLR 2026Poster
45
From Genomic Whispers to Therapeutics: Multi-Resolution Transcriptome-Guided Diffusion Models for Drug Design and Screening
ICLR 2026Withdrawn
17
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
ICLR 2026Withdrawn
5
VGGT-X: When VGGT Meets Dense Novel View Synthesis
ICLR 2026Withdrawn
通讯5
Scaling Autonomous Driving Safety with Synthetic Data
ICLR 2026Withdrawn
5
SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where
ICLR 2026Withdrawn
5
LVCap-Eval: Towards Holistic Long Video Caption Evaluation for Multimodal LLMs
ICLR 2026Withdrawn
通讯-1
IconBank: Mining Hard Samples via Visual Concepts for Data-Efficient GUI Grounding
ICLR 2026Withdrawn
通讯19
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
NeurIPS 2025Spotlight
21
Enhancing End-to-End Autonomous Driving with Latent World Model
ICLR 2025Poster
22
DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving
NeurIPS 2025Poster
通讯23
TC-Light: Temporally Coherent Generative Rendering for Realistic World Transfer
NeurIPS 2025Poster
通讯18
CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes
ICLR 2025Poster
通讯22
McEval: Massively Multilingual Code Evaluation
ICLR 2025Poster
27
FreeVS: Generative View Synthesis on Free Driving Trajectory
ICLR 2025Poster
通讯30
Reconstructive Visual Instruction Tuning
ICLR 2025Poster
通讯28
MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models
ICLR 2025Poster
19
OmniBench: Towards The Future of Universal Omni-Language Models
ICLR 2025Rejected
28
MIO: A Foundation Model on Multimodal Tokens
ICLR 2025Rejected
20
POC: Preventing the Over-Collapse of Classes for Class-Incremental Learning
ICLR 2025Rejected
30
Towards Flexible and Controllable Unknown Rejection
ICLR 2025Rejected
三作23
AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs
ICLR 2025Rejected
通讯20
HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
ICLR 2025Withdrawn
4
FlexDrive: Toward Trajectory Flexibility in Driving Scene Reconstruction and Rendering
ICLR 2025Withdrawn
三作5
UI-Pro: A Hidden Recipe for Building Vision-Language Models for GUI Grounding
ICLR 2025Withdrawn
通讯