STSyn-BEV: BEV segmentation from surround-view fisheye cameras via spatio-temporal synchronization

STSyn-BEV: сегментация BEV с Surround-View фишай-камер посредством пространственно-временной синхронизации
Ping Liu, Xuan Shao, Le Huang
2026-03-09

BEV segmentationFB-SSEM datasetPose-Sync encoderattention-based global reasoningconvolution-based structural refinementheterogeneous pathwaysmIoU improvementregion-level contrastive learningsemantic consistency supervisionspatio-temporal synchronizationstage-wise supervision decodersurround-view fisheye cameras
Abstract Bird’s-Eye-View (BEV) semantic segmentation is critical for environmental perception in autonomous driving. Surround-view fisheye camera systems are increasingly adopted to enlarge the perception range and eliminate blind spots. However, severe geometric distortions and frequent ego-motion make accurate spatio-temporal feature alignment across multiple views and timestamps challenging. Such misalignment often leads to semantic inconsistency and notable drops in BEV segmentation accuracy. Moreover, most existing methods overlook these alignment errors and apply semantic supervision only at the final output, resulting in suboptimal intermediate BEV representations. To address these challenges, we propose STSyn-BEV, a Spatio-Temporal Synchronized BEV segmentation framework for surround-view fisheye cameras. It comprises three key components: a Pose-Sync (pose-synchronized) encoder, a semantic consistency supervision module, and a stage-wise supervision decoder with heterogeneous pathways. First, the Pose-Sync encoder explicitly transforms multi-view fisheye features from previous poses and timestamps into a unified BEV space via geometric transformation, substantially improving geometric consistency and temporal alignment. Second, the semantic consistency supervision module applies region-level contrastive learning to aggregated BEV features, enhancing semantic discrimination particularly for long-tailed categories. Third, the deep supervised decoder employs heterogeneous pathways—attention-based for global semantic reasoning and convolution-based for fine-grained structural refinement—guided by stage-wise supervision, enabling improved BEV feature decoding without additional inference cost. Extensive experiments on the FB-SSEM dataset demonstrate that STSyn-BEV surpasses state-of-the-art fisheye image-based BEV segmentation methods, notably achieving a 6.25% mIoU improvement over the strongest fisheye-specific baseline.
1
A deep supervised decoder with heterogeneous pathways (attention-based for global semantics and convolution-based for structural refinement) and stage-wise supervision improves BEV feature decoding without extra inference cost.
2
A semantic consistency supervision module uses region-level contrastive learning on aggregated BEV features to enhance semantic discrimination, benefiting long-tailed categories.
3
On the FB-SSEM dataset, STSyn-BEV outperforms state-of-the-art fisheye image-based BEV segmentation methods, achieving a 6.25% mIoU improvement over the strongest fisheye-specific baseline.
4
STSyn-BEV is a spatio-temporal synchronized BEV segmentation framework designed for surround-view fisheye cameras to address geometric distortion and ego-motion misalignment.
5
The Pose-Sync encoder geometrically transforms multi-view fisheye features from previous poses and timestamps into a unified BEV space, substantially improving geometric consistency and temporal alignment.

Surround-view fisheye camera-based bird’s-eye-view (BEV) semantic segmentation system for autonomous driving

Spatio-temporal synchronization and semantic-consistency-enhanced BEV feature alignment and decoding (including pose-synchronized encoding, region-level contrastive supervision, and stage-wise heterogeneous decoding) to improve BEV segmentation accuracy

Publication Details
Publication Date
2026-03-09
Journal
Publisher
ISSN
Cited by
0
Access Type
Author Information
Authors
Ping Liu
Xuan Shao
Le Huang
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%