M.S. Student, Beihang University

Pengyiang Liu刘彭逸昂

I am a master's student at the CoLab, Institute of Artificial Intelligence, Beihang University, advised by Prof. Si Liu.

My research asks how multimodal models can keep up with the world as it unfolds: maintaining and updating an internal world state over continuous video streams, and reasoning over long-horizon visual evidence without losing grounding. I build benchmarks and agents for streaming and long-form video understanding.

Night view over Beijing from an observation deck

News

Publications

* denotes equal contribution.

SVCBench teaser: streaming counting queries along the video timeline

ECCV 2026

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

Pengyiang Liu, Zhongyue Shi, Hongye Hao, Qi Fu, Xueting Bi, Siwei Zhang, Xiaoyang Hu, Zitian Wang, Linjiang Huang, Si Liu

Repositions counting as a minimal, deterministically verifiable probe for how models maintain world state during video playback. 406 videos, 1,000 streaming QA pairs, and 4,576 timeline query points across 8 subcategories reveal that current multimodal LLMs nearly fail on long-horizon state maintenance such as periodic event counting.

OVO-S-Bench overview: four levels of streaming spatial intelligence

EMNLP 2026 Main

OVO-S-Bench: A Hierarchical Benchmark for Streaming Spatial Intelligence in Multimodal LLMs

Yifei Li*, Pengyiang Liu*, Yuhang Zang, Zhongyue Shi, Qi Fu, Hongye Hao, Jiwen Lu

Constrains models to the video prefix available at query time and evaluates streaming spatial intelligence across four levels, from instantaneous egocentric perception to global topological mapping. Mapping the environment topology emerges as the dominant bottleneck, with frontier models far behind humans.

TRACE architecture: evidence closure loop over anchored visual evidence

EMNLP 2026 Main

TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding

Pengyiang Liu*, Junbo Niu*, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu

A long-video reasoning agent that keeps the answer path grounded in raw visual clips instead of text summaries. Evidence is organized as anchored, growing trajectories, and a trajectory reconciler decides when to stop by checking answer convergence, addressing the "when to stop" problem in long-horizon reasoning. Ships with VES-Bench, an evidence-coverage audit benchmark.

Experience