Universe Invedors Logo
Universe InvedorsAI · XR · GAMES · CLOUD
← Back to Insights
3D & Performance10 min readJanuary 18, 2025

Optimising Large GLB Models for Web-Based 3D Experiences

A practical guide to Meshopt compression, KTX2 textures, geometry quantisation and LOD strategies for delivering high-quality 3D on the web without destroying performance.

GLBThree.jsWebGLPerformanceCompression

Introduction

Delivering high-quality 3D models on the web requires careful optimisation. Large GLB files can destroy load times, crash mobile browsers, and provide poor user experience. This article provides practical techniques for optimising GLB models while maintaining visual quality.

We will cover Meshopt compression for geometry, KTX2 texture compression, geometry quantisation, LOD (Level of Detail) strategies, and runtime optimisation techniques. These techniques have been tested in production with models ranging from architectural visualisations to game assets.

Meshopt Compression

Meshopt compression is the single most effective optimisation for GLB geometry. It can reduce file size by 50-80% with no visual quality loss. The compression works by optimising vertex order for better compression, quantising vertex attributes, and applying general-purpose compression.

How Meshopt Works

Meshopt compression reorders vertices to improve cache locality during compression, quantises floating-point vertex attributes to reduce precision where visual impact is minimal, and applies LZ-style compression to the resulting data. The decompression happens in the browser using JavaScript, so there is a small runtime cost but it is typically outweighed by the network savings.

Implementation

To use Meshopt compression, you need to apply it during the GLB export process. Tools like glTF-Transform provide command-line interfaces for Meshopt compression. The basic command is: gltf-transform input.glb output.glb --meshopt. You can adjust compression levels to balance file size against decompression time.

In the browser, you need to include the Meshopt decoder and configure your loader to use it. For Three.js, this means loading the MeshoptDecoder and setting the GLTFLoader meshoptDecoder property. The decoder is small (~20KB gzipped) and only needs to be loaded once.

Results

In our testing, Meshopt compression reduced a 50MB architectural model to 12MB—a 76% reduction. Load time dropped from 8 seconds to 2 seconds on a typical mobile connection. The visual quality was identical to the uncompressed version. For models with simple geometry but high vertex count, the savings can be even larger.

KTX2 Texture Compression

Textures often account for 70-90% of GLB file size. KTX2 texture compression with Basis Universal can reduce texture size by 5-10x while maintaining quality. KTX2 is a container format that can hold multiple texture compression formats, and Basis Universal provides a universal compression format that works across all platforms.

Why KTX2 Over Traditional Formats

Traditional texture formats like JPEG and PNG are designed for 2D images, not GPU textures. They do not compress well in GPU memory and require transcoding at runtime. KTX2 with Basis Universal provides GPU-native compression that stays compressed in memory, reducing both download size and memory usage.

Compression Settings

Basis Universal provides multiple compression quality levels. For most web applications, we use quality level 128 (out of 255) which provides good visual quality with significant size reduction. For textures where quality is critical (like UI elements or hero assets), we use higher quality levels up to 200. The quality setting is per-texture, so you can optimise selectively.

We also enable texture filtering and mipmaps in the KTX2 export. Mipmaps are critical for performance at distance—they prevent shimmering and reduce texture bandwidth when objects are far from the camera.

Browser Support

KTX2 is supported in all modern browsers through the Basis Universal transcoder. The transcoder converts the universal Basis format to the native GPU format (ASTC, ETC2, S3TC, etc.) at load time. This transcode step adds a small overhead but is typically faster than downloading larger textures in traditional formats.

Geometry Quantisation

Geometry quantisation reduces the precision of vertex attributes to save space. Modern GPUs can handle lower precision without visible quality loss for most use cases.

Position Quantisation

Vertex positions are typically stored as 32-bit floats. For most models, 16-bit or even 12-bit precision is sufficient. The key is to quantise relative to the model bounding box rather than absolute coordinates. This ensures that precision is used where it matters most—within the model itself rather than in world space.

Tools like glTF-Transform can automatically quantise positions during export. The default settings work well for most models. For architectural models where precision is critical for alignment, you might need to use higher precision for specific meshes.

Normal and UV Quantisation

Normals and UV coordinates can also be quantised. Normals can be stored as 16-bit values or even octahedral encoding for further compression. UV coordinates typically need less precision than positions—12-bit is often sufficient. The quantisation parameters should be tested visually to ensure no artifacts appear.

Index Buffer Optimisation

Index buffers can be optimised by using smaller data types when possible. If a model has fewer than 65535 vertices, indices can be stored as 16-bit values instead of 32-bit, halving the index buffer size. Tools can automatically detect when this is safe and apply the optimisation.

Level of Detail (LOD) Strategies

LOD reduces detail for objects that are far from the camera or small on screen. This is critical for performance with large scenes.

LOD Generation

LOD meshes can be generated automatically using tools like Simplygon or manually created by artists. Automatic LOD generation uses algorithms to reduce polygon count while preserving silhouette and important details. We typically generate 3-5 LOD levels per model, with each level reducing polygon count by approximately 50%.

LOD Switching

LOD switching should be based on screen-space size rather than absolute distance. This ensures consistent quality across different screen resolutions and field-of-view settings. The switch distance should include a hysteresis buffer to prevent rapid switching when the object is near the threshold.

For Three.js, we use the LOD object which handles automatic switching based on distance. We configure the distance thresholds based on testing—too aggressive LOD switching causes visible popping, too conservative switching wastes performance.

Texture LOD

Textures should also have LOD through mipmaps. KTX2 automatically includes mipmaps when exported correctly. At runtime, the GPU automatically selects the appropriate mipmap level based on distance. This reduces texture bandwidth and prevents aliasing artifacts.

Runtime Optimisation

Beyond file size optimisation, runtime techniques improve performance and user experience.

Progressive Loading

Progressive loading shows a low-quality version of the model immediately and progressively improves it. This can be implemented by loading a low-poly LOD first, then loading higher LODs in the background. Users see something quickly rather than waiting for the full model to load.

Instancing

For repeated objects like trees, chairs, or architectural elements, use instancing. Instancing renders multiple copies of the same mesh with a single draw call, dramatically reducing CPU overhead. Three.js supports instancing through InstancedMesh. The memory savings are also significant—only one copy of the geometry is stored in memory.

Frustum and Occlusion Culling

Frustum culling skips objects outside the camera view. Occlusion culling skips objects hidden behind other objects. Three.js handles frustum culling automatically. For occlusion culling, we use techniques like portal culling for architectural models or software occlusion culling for complex scenes.

Material Simplification

Complex materials with multiple textures and shader effects are expensive. Where possible, simplify materials—combine textures into atlases, use fewer shader passes, and avoid expensive effects like real-time shadows on mobile devices. PBR materials can be approximated with simpler materials for distant objects.

Testing and Validation

Optimisation requires testing to ensure quality is maintained. We test on multiple devices and network conditions.

Visual Quality Testing

We compare optimised models side-by-side with originals at various zoom levels and lighting conditions. We look for artifacts like texture banding, geometry popping at LOD switches, and normal map errors. We also test with different camera angles to catch issues that only appear from certain viewpoints.

Performance Testing

We measure frame rate, GPU memory usage, and load time on target devices. For mobile, we test on mid-range and low-end devices, not just flagship phones. We simulate slow network connections using Chrome DevTools to ensure load times are acceptable even on poor connections.

A/B Testing

For production applications, we sometimes A/B test different optimisation levels with real users. This helps us understand the trade-off between quality and performance from a user perspective. Users often prefer slightly lower quality if it means significantly faster load times.

Conclusion

Optimising GLB models for the web requires a multi-faceted approach. Meshopt compression for geometry, KTX2 for textures, quantisation for precision reduction, LOD for distance-based detail, and runtime optimisations for performance. The key is to apply these techniques systematically and test thoroughly.

The results are worth the effort—models that load 5-10x faster, use less memory, and run smoothly on mobile devices. This enables web-based 3D experiences that compete with native applications in quality while maintaining the accessibility of the web.