M.S. Student, Beihang University

Pengyiang Liu刘彭逸昂

I am a master's student at the CoLab, Institute of Artificial Intelligence, Beihang University, advised by Prof. Si Liu.

My research asks how multimodal models can keep up with the world as it unfolds: maintaining and updating an internal world state over continuous video streams, and reasoning over long-horizon visual evidence without losing grounding. I build benchmarks and agents for streaming and long-form video understanding.

Night view over Beijing from an observation deck

News

Publications

* denotes equal contribution.

SVCBench overview: process-level streaming counting evaluation Overview
SVCBench taxonomy: object and event counting categories Taxonomy

ECCV 2026

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

Pengyiang Liu, Zhongyue Shi, Hongye Hao, Qi Fu, Xueting Bi, Siwei Zhang, Xiaoyang Hu, Zitian Wang, Linjiang Huang, Si Liu

Repositions counting as a minimal, deterministically verifiable probe for how models maintain world state during video playback. 406 videos, 1,000 streaming QA pairs, and 4,576 timeline query points across 8 subcategories reveal that current multimodal LLMs nearly fail on long-horizon state maintenance such as periodic event counting.

OVO-S-Bench overview: four levels of streaming spatial intelligence Overview
OVO-S-Bench taxonomy examples across four levels of spatial understanding Taxonomy

EMNLP 2026 Main Oral Outstanding Award Nominee

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

Yifei Li*, Pengyiang Liu*, Yuhang Zang, Zhongyue Shi, Qi Fu, Hongye Hao, Jiwen Lu

Constrains models to the video prefix available at query time and evaluates streaming spatial intelligence across four levels, from instantaneous egocentric perception to global topological mapping. Mapping the environment topology emerges as the dominant bottleneck, with frontier models far behind humans.

TRACE overview: evidence closure loop over anchored visual evidence Overview
TRACE method: evidence growth trajectory and trajectory reconciliation Method

EMNLP 2026 Main Oral

TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding

Pengyiang Liu*, Junbo Niu*, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu

A long-video reasoning agent that keeps the answer path grounded in raw visual clips instead of text summaries. Evidence is organized as anchored, growing trajectories, and a trajectory reconciler decides when to stop by checking answer convergence, addressing the "when to stop" problem in long-horizon reasoning. Ships with VES-Bench, an evidence-coverage audit benchmark.

StaMina overview: maintaining counting state as a video scene changes Overview
StaMina method: state-conditioned transition and count-path learning Method

arXiv 2026

StaMina: When Should the Count Change? Learning State Maintenance for Causal Video Counting

Pengyiang Liu, Dongyue Lyu, Junbo Niu, Zhongyue Shi, Jiahao Xie, Si Liu

Learns state-conditioned updates for continuous video counting, separating visual recognition from the maintenance of visibility, persistent identities, and completed-event records. A differentiable recurrence organizes spatial queries and event annotations into count trajectories.

RTP overview: adaptive revision of visual plans after execution feedback Overview
RTP method: grounded facts and feedback-conditioned plan revision Method

arXiv 2026

RTP: Revisable Visual Plans for Closed-Loop World-Action Models

Pengyiang Liu, Junbo Niu, Wenhao Zheng, Xinchen Chen, Canyu Li, Zhongyue Shi, Jiahao Xie, Si Liu

Maintains a visual future as a persistent action condition and revises it after execution feedback. A learned revision bridge resumes an intermediate generation state, adapts its continuation to new observations, and lets an adaptive policy choose retention, revision, or fresh replanning.

Experience