GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes

GOLD-BEV is a cross-view dataset of time-synchronized aerial and ego-centric sensor observations for dense semantic bird's-eye-view (BEV) mapping of dynamic road scenes. A helicopter-mounted camera captures high-resolution overhead RGB imagery while an instrumented vehicle simultaneously records a front-facing RGB camera, multiple LiDAR sensors, and GNSS/INS measurements. The synchronized aerial view provides a dense, geometrically align observation of the same dynamic scene seen from the vehicle.

The dataset contains 8,199 synchronized samples collected across approximately 60 km of urban, suburban, and highway driving. GOLD-BEV supports research on ego-to-BEV reconstruction, dense semantic BEV mapping, cross-view learning, and aerial supervision for dynamic road scenes.

Cross-view aerial supervision for BEV semantic mapping.
A helicopter-mounted camera records high-resolution overhead RGB imagery (a), time-synchronized with an instrumented car that captures a forward-facing RGB view and LiDAR sweeps (b). By geo-aligning the aerial imagery to the vehicle frame, we obtain BEV-aligned crops and dense semantic targets that supervise BEV map prediction from ego sensors (c).”
The examples show synchronized ground RGB and LiDAR observations
together with the corresponding vehicle-centered aerial crop a dense semantic BEV supervision. GOLD-BEV uses five semantic classes: road, sidewalk, building, vehicle, and vulnerable road user (VRU).

Please cite the following paper if you use GOLD-BEV in your work:

Niemeijer, J., Ben Zekri, A. E., Bahmanyar, R., Schmälzle, P. M., Chaabouni-Chouayakh, H., and Kurz, F., “GOLD-BEV: GrOund and aeriaL Data for Dense Semantic BEV Mapping of Dynamic Scenes,” arXiv preprint arXiv:2604.19411, 2026.