DroneSplat+: Semantics-Enhanced 3D Gaussian Splatting for
Robust 3D Reconstruction from In-the-Wild Drone Imagery
Jiadong Tang1     Yu Gao1     Dianyi Yang1      Haolin Yu1      Tianji Jiang1      Mengyin Fu1      Yufeng Yue1      Yi Yang1
1Beijing Institute of Technology    
Poster Image
Abstract

Drones have become indispensable for reconstructing in-the-wild scenes due to their remarkable agility. Recent advances in radiance-field representations have achieved photorealistic rendering, opening new avenues for 3D reconstruction from drone imagery. However, limited view constraints hinder reliable geometric recovery, while dynamic distractors in the wild violate the multi-view consistency in radiance fields. To address these challenges, we propose DroneSplat+, a novel semantics-enhanced framework for robust 3D reconstruction from in-the-wild drone imagery. DroneSplat+ couples multi-view stereo predictions with segmentation priors to impose semantic constraints that steer Gaussian optimization, enabling accurate geometry recovery under limited-view conditions. To suppress dynamics, DroneSplat+ introduces a dual-elimination strategy that combines adaptive local-global masking and label-based localization to precisely identify and remove dynamic distractors from static scenes. We also provide a drone-captured 3D reconstruction benchmark dataset encompassing both dynamic and static scenes for comprehensive evaluation. Experimental results show that our method outperforms both 3DGS and NeRF baselines in reconstructing in-the-wild drone imagery.
Framework

Pipeline Image
Given a few drone images of a wild scene, we reconstruct the scene with accurate geometry while eliminating dynamic distractors. We estimate a dense point cloud with multi-view stereo and annotate each point with a unique semantic label. Then we perform label unification based on feature correspondences. We initialize 3DGS from the point cloud with unified labels, and regularize Gaussians using an instance-aware grid and a segmentation-guided consistency optimization. In parallel, we compute normalized residuals and aggregate residuals within each mask to obtain instance-wise residuals. Based on per-iteration residual statistics, we adaptively set a local threshold to predict local masks. We also track high residual instances to produce global masks. We query the semantic labels of instances flagged as dynamic distractors and prune Gaussians who include these labels, ensuring robust distractor removal. Ultimately, we can obtain clean novel view renderings with accurate geometry.
DroneSplat Dataset

Limited-view Reconstruction Results

Road
FSGS
EAP-GS
DropGaussian
DroneSplat+
Plaza
FSGS
EAP-GS
DropGaussian
DroneSplat+
Garden
FSGS
EAP-GS
DropGaussian
DroneSplat+
Distractor Elimination Results

TangTian
GS-W
RobustSplat
SpotlessSplat
DroneSplat+
Simingshan
GS-W
RobustSplat
SpotlessSplat
DroneSplat+
Sculpture
GS-W
RobustSplat
SpotlessSplat
DroneSplat+