影响力指数
92.09/100
前 0.4%
全站排名 #277
发表论文47
平均评分5.3
年均产出15.7 篇/年

Zhaoxiang Zhang

Full Professor@Institute of automation, Chinese academy of science, Chinese Academy of Sciences·中国·OpenReview
研究方向

Embodied Intelligence · Agent Learning · Autonomous driving · biological-inspired model · brain-inspired computing · deep neural networks · obejct detection · semantic segmentation · action recognition · person re-identification · scene analysis and understanding

6.5
13

Unified Vision-Language-Action Model

ICLR 2026Poster
通讯
6.0
31

IF-VidCap: Can Video Caption Models Follow Instructions?

ICLR 2026Poster
6.0
11

YuE: Scaling Open Foundation Models for Long-Form Music Generation

ICLR 2026Poster
5.5
16

Uniform Discrete Diffusion with Metric Path for Video Generation

ICLR 2026Poster
5.5
16

DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving

ICLR 2026Poster
通讯
5.3
15

ResGen: Residual Diffusion Model for LiDAR-based Point Cloud Generation

ICLR 2026Rejected
通讯
5.0
29

SpectraLLM: Uncovering the Ability of LLMs for Molecule Structure Elucidation from Multi-Spectra

ICLR 2026Poster
通讯
5.0
16

Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Method

ICLR 2026Poster
通讯
5.0
16

Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs

ICLR 2026Poster
通讯
4.7
15

GSPlane: Concise and Accurate Planar Reconstruction via Structured Representation

ICLR 2026Rejected
通讯
4.5
23

MLLM-CL: Continual Learning for Multimodal Large Language Models

ICLR 2026Rejected
通讯
4.5
32

FeatureBench: Benchmarking Agentic Coding for Complex Feature Development

ICLR 2026Poster
通讯
4.5
26

HiPO: Hybrid Policy Optimization for Dynamic Reasoning in LLMs

ICLR 2026Withdrawn
4.5
44

OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs

ICLR 2026Poster
4.0
45

From Genomic Whispers to Therapeutics: Multi-Resolution Transcriptome-Guided Diffusion Models for Drug Design and Screening

ICLR 2026Withdrawn
4.0
17

MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents

ICLR 2026Withdrawn
3.5
5

VGGT-X: When VGGT Meets Dense Novel View Synthesis

ICLR 2026Withdrawn
通讯
3.5
5

Scaling Autonomous Driving Safety with Synthetic Data

ICLR 2026Withdrawn
3.5
5

SIG-Chat: Spatial Intent-Guided Conversational Gesture Generation Involving How, When and Where

ICLR 2026Withdrawn
3.5
5

LVCap-Eval: Towards Holistic Long Video Caption Evaluation for Multimodal LLMs

ICLR 2026Withdrawn
通讯
-1

IconBank: Mining Hard Samples via Visual Concepts for Data-Efficient GUI Grounding

ICLR 2026Withdrawn
通讯
9.1
19

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation

NeurIPS 2025Spotlight
7.0
21

Enhancing End-to-End Autonomous Driving with Latent World Model

ICLR 2025Poster
6.8
22

DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving

NeurIPS 2025Poster
通讯
6.8
23

TC-Light: Temporally Coherent Generative Rendering for Realistic World Transfer

NeurIPS 2025Poster
通讯
6.5
18

CityGaussianV2: Efficient and Geometrically Accurate Reconstruction for Large-Scale Scenes

ICLR 2025Poster
通讯
6.5
22

McEval: Massively Multilingual Code Evaluation

ICLR 2025Poster
5.8
27

FreeVS: Generative View Synthesis on Free Driving Trajectory

ICLR 2025Poster
通讯
5.8
30

Reconstructive Visual Instruction Tuning

ICLR 2025Poster
通讯
5.8
28

MTU-Bench: A Multi-granularity Tool-Use Benchmark for Large Language Models

ICLR 2025Poster
5.8
19

OmniBench: Towards The Future of Universal Omni-Language Models

ICLR 2025Rejected
5.5
28

MIO: A Foundation Model on Multimodal Tokens

ICLR 2025Rejected
5.3
20

POC: Preventing the Over-Collapse of Classes for Class-Incremental Learning

ICLR 2025Rejected
5.3
30

Towards Flexible and Controllable Unknown Rejection

ICLR 2025Rejected
三作
5.0
23

AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs

ICLR 2025Rejected
通讯
4.8
20

HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models

ICLR 2025Withdrawn
4.7
4

FlexDrive: Toward Trajectory Flexibility in Driving Scene Reconstruction and Rendering

ICLR 2025Withdrawn
三作
4.3
5

UI-Pro: A Hidden Recipe for Building Vision-Language Models for GUI Grounding

ICLR 2025Withdrawn
通讯