Researchers have introduced SILSA, a novel framework for generating high-resolution three-dimensional (3D) models that addresses limitations in current methods. Existing approaches often rely on voxel latents and multi-stage pipelines, which can fragment continuous surfaces, increase computational costs, and compromise topological consistency, particularly for thin or complex shapes. SILSA represents 3D shapes using compact sliding-window slice latents, a method that preserves cross-sectional continuity and supports single-stage generation.
The SILSA framework employs a Slice VAE to encode oriented surface samples into multi-axis slice latents. These latents are then used by a sparse volumetric decoder to reconstruct the 3D geometry. To ensure structural accuracy, SILSA incorporates slice-level topology supervision. This supervision aligns persistence diagrams and Betti transitions across neighboring slices, effectively maintaining topological correctness. A Volumetric Anchor Lattice component coordinates directional slice streams within a shared 3D workspace.
Experiments detailed in the paper demonstrate that SILSA significantly improves structural fidelity while reducing generation costs. The framework achieved an 8.7% increase in Peak Signal-to-Noise Ratio (PSNR), a 5.96 absolute point gain in coverage, and a 9.2% reduction in Betti error compared to the strongest baseline models. Furthermore, SILSA uses 70% fewer tokens than the next most compact baseline and over 98% fewer tokens than sparse or hierarchical tokenizers. This efficiency translates to a 40.4% reduction in training memory and a 58.5% decrease in inference time. Qualitative results also indicate improved preservation of thin structures, repeated components, and long-range connectivity within generated 3D models.
The development of SILSA is a step toward more efficient and topologically sound high-resolution 3D generation. By moving away from expensive voxel tokens and fragmented multi-stage processes, SILSA offers a more integrated and efficient approach to creating detailed 3D assets. The method's ability to maintain topological integrity is particularly beneficial for generating complex geometries that have previously posed challenges for automated systems.
