Papers
arxiv:2606.01573

VG^2GT: Voxel-Gaussian Splatting Visual Geometry Grounded Transformer

Published on Jun 3
Authors:
,
,
,
,
,

Abstract

Gaussian splatting has shown strong potential for 3D reconstruction and novel view synthesis. However, most existing methods require accurate camera parameters and per-scene optimization, while feed-forward methods with pixel-aligned Gaussian primitives often suffer from artifacts and non-uniform primitives. In this paper, we propose VG^2GT, a Voxel-Gaussian Splatting Visual Geometry-Grounded Transformer. VG^2GT leverages a frozen pretrained visual foundation model (VFM), incorporates a multi-scale differentiable voxel module to enhance geometric understanding, and directly splits and regresses Gaussian primitive parameters from voxel features. During training, depth maps are supervised through stochastic solid volume rendering, enabling geometrically accurate Gaussian scene reconstruction while keeping the visual foundation model fully frozen. This design enables VG^2GT to be seamlessly plugged into any patch-feature-based VFM, while substantially reducing the required training cost. VG^2GT outperforms current state-of-the-art methods on widely used DTU, Replica, TAT, and ScanNet datasets.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2606.01573
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2606.01573 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2606.01573 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2606.01573 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.