Joint semantic and geometric segmentation of videos with a stage model

Buyu Liu, Xuming He, S. Gould

Research output: Chapter in Book/Report/Conference proceedingConference contribution


We address the problem of geometric and semantic consistent video segmentation for outdoor scenes. With no assumption on camera movement, we jointly model the semantic-geometric class of spatio-temporal regions (supervoxels) and geometric scene layout in each frame. Our main contribution is to propose a stage scene model to efficiently capture the dependency between the semantic and geometric labels. We build a unified CRF model on supervoxel labels and stage parameters, and design an alternating inference algorithm to minimize the resulting energy function. We also extend smoothing based on hierarchical image segmentation to spatio-temporal setting and show it achieves better performance than a pairwise random field model. Our method is evaluated on the CamVid dataset and achieves state-of-the-art per-pixel as well as per-class accuracy in predicting both semantic and geometric labels.
Original languageEnglish
Title of host publicationIEEE Winter Conference on Applications of Computer Vision
PublisherInstitute of Electrical and Electronics Engineers (IEEE)
Number of pages8
ISBN (Electronic)978-1-4799-4985-4
Publication statusPublished - Mar 2014


Dive into the research topics of 'Joint semantic and geometric segmentation of videos with a stage model'. Together they form a unique fingerprint.

Cite this