Universe Invedors Logo
Universe InvedorsAI · XR · GAMES · CLOUD
← Back to Insights
XR Development12 min readMarch 15, 2025

Building Production-Grade XR Training for Meta Quest

A detailed walkthrough of the architecture, AI systems and optimisation strategies behind deploying enterprise VR training applications on Meta Quest at scale.

UnityMeta QuestXRAI NPCTraining

Introduction

Enterprise VR training has moved beyond experimental pilots to production deployments at scale. When building training platforms for Meta Quest that will be used by hundreds or thousands of trainees, the requirements shift dramatically from prototype to production. This article details the architecture, AI systems, and optimisation strategies we used to deploy a production-grade VR training platform for aviation security personnel.

The platform includes 13 training modules, AI-driven passengers with dynamic behaviour, real-time voice interaction, After Action Review with AI-generated feedback, adaptive difficulty, and multiplayer team coordination. This is not a demo—it is a production system used daily by security officers.

System Architecture Overview

The architecture is designed around three core principles: scalability across multiple concurrent training sessions, AI-driven dynamic content generation, and comprehensive assessment and feedback systems.

Unity + OpenXR Foundation

We built the VR client using Unity with OpenXR to ensure cross-platform compatibility across Meta Quest 2, Quest 3, and Windows VR. The XR Interaction Toolkit provided the foundation for hand tracking and physical interaction patterns. We structured the project with a clear separation between scene-specific content and shared systems—AI engine, scoring system, multiplayer networking, and audio pipelines exist as independent modules that can be loaded into any training scenario.

AI Passenger Engine

The core innovation is the AI Passenger Engine. Instead of fixed NPCs with scripted dialogue, passengers are generated dynamically by a GPT-powered system. Each passenger has a name, age, destination, travel history, personality type, stress level, and scenario-specific objective. The AI engine generates these attributes based on the training module and difficulty level, then uses them to drive conversation behaviour and decision-making throughout the scenario.

The engine maintains conversation state, remembers previous interactions within the session, and adjusts behaviour based on trainee decisions. A stressed passenger might become more agitated if the trainee uses harsh language. A nervous passenger might provide incomplete information unless the trainee uses appropriate de-escalation techniques. This dynamic behaviour prevents trainees from memorising patterns and forces genuine decision-making.

Scenario Engine with Adaptive Difficulty

The Scenario Engine manages the training flow across 5 difficulty levels. Each level adjusts multiple parameters: passenger stress levels, number of concurrent incidents, time pressure, complexity of threats, and required procedural precision. The engine monitors trainee performance in real-time and can inject additional challenges or provide guidance based on demonstrated competence.

For example, if a trainee consistently identifies threats correctly but struggles with communication procedures, the engine might increase the number of passengers requiring verbal interaction while maintaining threat complexity at an appropriate level. This adaptive approach keeps training challenging without overwhelming trainees.

Scoring and After Action Review

Every action a trainee takes is tracked and scored. The scoring engine evaluates identification accuracy, communication effectiveness, procedure compliance, escalation timing, and team coordination. Each decision point has associated point values and weights based on its importance to the training objectives.

After each scenario, an AI instructor generates a detailed After Action Review. This includes per-skill scores, a breakdown of mistakes with explanations, specific improvement recommendations, and comparison to previous performance. The AI instructor uses the complete action log to provide context-aware feedback rather than generic comments.

Multiplayer Coordination

For team-based scenarios, we implemented Photon multiplayer to support 4 concurrent officer roles: check-in, baggage, patrol, and incident commander. Officers coordinate via simulated VHF radio with realistic push-to-talk mechanics. The multiplayer system synchronises AI passenger state across all clients, ensuring that decisions made by one officer affect the scenario state for the entire team.

Voice Interaction Pipeline

Real-time voice interaction is critical for training communication skills. We implemented a pipeline using Speech-to-Text (STT) for trainee input, GPT for AI response generation, and Text-to-Speech (TTS) for passenger responses.

Speech-to-Text Integration

We evaluated multiple STT solutions including cloud-based APIs and on-device options. For production deployment, we selected a cloud-based solution with low latency and high accuracy for security terminology. The STT system is optimised for the specific vocabulary used in aviation security—threat types, procedure names, equipment terminology—to improve recognition accuracy.

The system handles background noise common in training environments through noise cancellation and confidence scoring. Low-confidence transcriptions trigger a clarification prompt from the AI passenger, maintaining training flow while ensuring accurate communication assessment.

GPT Integration and Prompt Engineering

The GPT integration required careful prompt engineering to balance realistic behaviour with training objectives. We developed a multi-layered prompt system: a base prompt establishing the passenger persona and scenario context, a dynamic prompt incorporating conversation history and current game state, and an assessment prompt that evaluates trainee communication quality.

The prompt system includes specific instructions for training objectives—for example, the AI knows when to withhold information to test questioning technique, when to escalate stress to test de-escalation, and when to provide hints if the trainee is struggling significantly. This ensures that AI behaviour serves training goals rather than just simulating conversation.

Text-to-Speech with Emotional Variation

Passenger responses use TTS with emotional variation based on stress level and personality. We selected voices that could convey appropriate emotional range—calm for professional interactions, agitated for stressed passengers, hesitant for nervous individuals. The TTS system adjusts pitch, speed, and intonation based on the emotional state generated by the AI engine.

Performance Optimisation

Running complex AI systems and multiplayer networking on mobile VR hardware requires aggressive optimisation. We applied optimisation across multiple layers.

Asset Optimisation

All 3D assets were optimised for Quest hardware with polygon budgets, texture compression using ASTC, and material simplification where visual quality could be maintained. We used LOD (Level of Detail) systems aggressively—distant objects use simplified models, and critical interaction areas maintain higher detail. Asset streaming ensures that only currently needed assets are in memory.

Rendering Optimisation

We implemented occlusion culling, frustum culling, and dynamic resolution scaling to maintain frame rate. The rendering pipeline uses single-pass instanced rendering for repeated objects like passengers and equipment. We optimised shaders for mobile GPU architecture, avoiding expensive operations and using mobile-friendly alternatives.

AI Pipeline Optimisation

The AI pipeline is optimised to minimise latency while maintaining quality. We batch GPT requests when possible, use streaming responses for faster perceived response time, and cache common response patterns. The STT and TTS systems use efficient audio processing to minimise CPU load. All network requests are optimised for the specific constraints of mobile VR headsets.

Memory Management

Memory is the most constrained resource on Quest hardware. We implemented aggressive memory management with object pooling for frequently instantiated objects, texture memory management with automatic unloading of unused assets, and careful management of audio buffer sizes. The system monitors memory usage and can trigger garbage collection or asset unloading if memory pressure increases.

Instructor Dashboard and Analytics

Instructors need comprehensive tools to manage training and assess trainee progress. We built a web-based dashboard that provides scenario configuration, live monitoring of training sessions, and detailed analytics.

Scenario Configuration

Instructors can configure training scenarios by selecting modules, setting difficulty levels, choosing passenger archetypes, and defining custom parameters. The dashboard provides a visual interface for scenario design with real-time preview of configuration changes. Instructors can save scenario templates for reuse and share configurations across training locations.

Live Monitoring

During training sessions, instructors can monitor multiple trainees simultaneously. The dashboard shows real-time status including current scenario stage, score progression, communication quality metrics, and any flagged issues. Instructors can send guidance messages to trainees or inject scenario modifications if needed for training purposes.

Analytics and Certification Tracking

The analytics system tracks complete training history for each trainee including scores across all modules, improvement over time, common mistake patterns, and certification status. Instructors can generate reports for individual trainees or aggregate analysis across groups. The system supports certification requirements with configurable passing thresholds and recertification schedules.

Deployment and Operations

Production deployment requires robust operations infrastructure. We implemented systems for app distribution, update management, and ongoing support.

App Distribution

We use Meta Quest for Business for enterprise app distribution, allowing controlled deployment to specific headsets and automatic updates. The app is configured for kiosk mode to prevent unauthorised usage and ensure consistent training environment. We implemented device management integration for remote support and monitoring.

Update Management

Updates are deployed through staged rollouts—we first release to a small test group, monitor for issues, then gradually expand to all devices. The update system includes data migration for preserving training history and configuration. We maintain backward compatibility where possible to minimise disruption during transitions.

Monitoring and Support

We implemented comprehensive monitoring including crash reporting, performance metrics, AI pipeline latency tracking, and usage analytics. The support system includes remote diagnostics, log collection, and the ability to push configuration changes without full app updates. This enables rapid response to issues and continuous improvement based on usage data.

Results and Impact

The production deployment has demonstrated significant impact on training effectiveness and operational efficiency. Training cycle time was reduced by 60% compared to traditional methods. Physical training infrastructure costs were eliminated. Consistent cross-site assessment became possible with standardised scenarios and AI evaluation. The adaptive difficulty turned the simulator from a one-time exercise into a continuous learning platform.

Trainee feedback has been positive, with particular appreciation for the realistic AI behaviour and comprehensive After Action Review. Instructors report that the system enables more efficient training delivery and better identification of individual trainee weaknesses.

Conclusion

Building production-grade XR training for Meta Quest requires attention to architecture, AI integration, performance optimisation, and operational infrastructure. The key lessons from this deployment are: design for scalability from the start, invest in AI prompt engineering for training objectives, optimise aggressively for mobile VR constraints, and build comprehensive instructor tools for training management.

As VR hardware continues to improve and AI capabilities advance, the potential for XR training will only increase. The architectural patterns and optimisation strategies described here provide a foundation for building production-grade training platforms that can scale across organisations and training domains.