arxiv_cv 95% Match Research Paper Robotics engineers,Computer vision researchers,AR/VR developers,Autonomous systems researchers 1 week ago

Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting

computer-vision › 3d-vision

📄 Abstract

Abstract: Open-vocabulary querying in 3D space is crucial for enabling more intelligent perception in applications such as robotics, autonomous systems, and augmented reality. However, most existing methods rely on 2D pixel-level parsing, leading to multi-view inconsistencies and poor 3D object retrieval. Moreover, they are limited to static scenes and struggle with dynamic scenes due to the complexities of motion modeling. In this paper, we propose Segment then Splat, a 3D-aware open vocabulary segmentation approach for both static and dynamic scenes based on Gaussian Splatting. Segment then Splat reverses the long established approach of "segmentation after reconstruction" by dividing Gaussians into distinct object sets before reconstruction. Once reconstruction is complete, the scene is naturally segmented into individual objects, achieving true 3D segmentation. This design eliminates both geometric and semantic ambiguities, as well as Gaussian-object misalignment issues in dynamic scenes. It also accelerates the optimization process, as it eliminates the need for learning a separate language field. After optimization, a CLIP embedding is assigned to each object to enable open-vocabulary querying. Extensive experiments one various datasets demonstrate the effectiveness of our proposed method in both static and dynamic scenarios.

Authors (8)

Yiren Lu

Yunlai Zhou

Yiran Qiao

Chaoda Song

Tuo Liang

Jing Ma

+2 more

Submitted

March 28, 2025

arXiv Category

cs.CV

arXiv PDF

Key Contributions

Proposes Segment then Splat, a novel 3D-aware open-vocabulary segmentation approach using Gaussian Splatting that reverses traditional reconstruction order. By segmenting Gaussians into object sets *before* reconstruction, it achieves true 3D segmentation for both static and dynamic scenes, eliminating ambiguities.

Business Value

Enables more intelligent perception for robots and autonomous systems by allowing them to understand and segment 3D environments based on natural language descriptions. This is crucial for tasks like object manipulation, navigation, and human-robot interaction.

Paper Metadata

Innovation Type

novel method/framework

Deployment Feasibility

Moderate. Relies on Gaussian Splatting, which is computationally intensive, but offers a unified approach for static and dynamic scenes.

Limitations Addressed

Limitations of 2D pixel-level parsing for 3D tasks, multi-view inconsistencies, poor 3D object retrieval, and difficulties in handling dynamic scenes due to motion modeling complexities.

Technical Tags

open-vocabulary segmentation3D segmentationGaussian Splattingstatic scenesdynamic scenesobject retrievalmulti-view inconsistenciesmotion modeling3D-aware perceptionscene understanding

Research Topics

3D Computer VisionScene UnderstandingOpen-Vocabulary RecognitionRobotics PerceptionGaussian Splatting

Methods & Architectures

Segment then SplatGaussian Splattingobject set division before reconstruction3D-aware segmentation Segment then Splat

Applications & Tasks

robotics autonomous systems augmented reality 3D scene reconstruction scene understanding 2D pixel-level parsing limitations in 3Dmulti-view inconsistenciespoor 3D object retrievalstruggle with dynamic scenescomplexities of motion modeling open-vocabulary segmentation in 3D3D object retrievalsegmentation of static and dynamic scenes

Related Fields

Computer Vision3D ReconstructionRoboticsScene UnderstandingAugmented Reality

Keywords

3D segmentationopen-vocabularyGaussian Splattingdynamic scenesstatic scenesroboticsautonomous systemsscene understanding3D perceptionobject retrievalaugmented reality

Academic Context

#3D Computer Vision#Scene Understanding#Open-Vocabulary Recognition#Robotics Perception#Gaussian Splatting

Commercial Potential

Potential Products

3D perception systems for autonomous vehiclesRobotic manipulation planning toolsAR/VR environment understanding platforms

Target Industries

RoboticsAutomotiveLogisticsConstructionEntertainment (AR/VR)

Use Case Examples

Enabling a robot to identify and segment specific objects in a warehouseAllowing an AR system to understand and interact with dynamic 3D environmentsImproving 3D scene reconstruction for virtual reality

Competitive Edge

Offers a unified approach for both static and dynamic scenes, overcoming limitations of 2D methods and achieving true 3D segmentation by reversing the traditional reconstruction pipeline.

Production Readiness

Maturity Level

Research

View Full Paper Back to Papers