Universe Invedors Logo
Universe InvedorsAI · XR · GAMES · CLOUD
← Back to Insights
AI & 3D14 min readOctober 22, 2024

Real-Time Avatar Lip Synchronisation with PyTorch and FastAPI

Building a sub-50ms phoneme-driven lip sync pipeline for 3D avatars — from audio analysis with PyTorch to blend-shape control in Unity via WebSocket.

PyTorchFastAPIUnityLip SyncAvatar

Introduction

Real-time lip synchronisation for 3D avatars is challenging because it requires low latency while maintaining visual quality. Users notice lip sync delays above 50ms, and poor sync breaks immersion. This article describes a production pipeline that achieves sub-50ms end-to-end latency using PyTorch for audio analysis, FastAPI for the backend, and Unity for avatar rendering.

The pipeline processes audio input, extracts phonemes, maps phonemes to viseme targets, and drives Unity blend-shapes in real-time. We will cover each component in detail, including optimisation techniques for low-latency operation.

System Architecture

The system consists of three main components: the Audio Processor (PyTorch), the Backend Server (FastAPI), and the Avatar Renderer (Unity). Components communicate via WebSocket for low-latency bidirectional communication.

Audio Flow

Audio is captured from the microphone at 16kHz, sent in small chunks (20-40ms) to the backend, processed to extract phonemes, and viseme targets are returned to Unity. The chunk size balances latency against processing overhead—smaller chunks reduce latency but increase overhead.

Component Responsibilities

The Audio Processor handles audio feature extraction and phoneme classification using a PyTorch model. The Backend Server manages WebSocket connections, coordinates processing, and handles client state. The Avatar Renderer receives viseme targets and applies them to blend-shapes with smoothing for natural motion.

Audio Processing with PyTorch

The core of the system is the audio processing pipeline that converts raw audio to phoneme classifications.

Feature Extraction

We extract MFCC (Mel-frequency cepstral coefficients) features from audio chunks. MFCCs capture the spectral envelope of audio and are well-suited for phoneme classification. We extract 13 MFCC coefficients plus energy and delta/delta-delta features, resulting in a 40-dimensional feature vector per frame. Feature extraction uses librosa for efficient computation.

Phoneme Classification Model

The classification model is a lightweight CNN architecture designed for low-latency inference. The model consists of three convolutional layers with batch normalisation and ReLU activations, followed by two LSTM layers for temporal context, and a final dense layer for phoneme classification. The entire model has under 100K parameters, enabling fast inference on CPU.

Model Optimisation

We optimise the model for inference speed using techniques like quantisation (reducing precision from FP32 to FP16), model pruning (removing less important weights), and ONNX export for faster inference. The model is pre-loaded in memory to avoid loading overhead during inference. We also implement batch processing when multiple audio chunks are available.

Phoneme to Viseme Mapping

Phonemes are mapped to visemes using a predefined mapping table. Different languages use different mappings—we support English and can extend to other languages. The mapping considers co-articulation effects where adjacent phonemes influence each other. We implement a simple co-articulation model that blends between visemes based on phoneme context.

FastAPI Backend

The FastAPI backend manages real-time communication and coordinates processing.

WebSocket Implementation

We use WebSocket for low-latency bidirectional communication. FastAPI's WebSocket support provides a clean API for managing connections. Each client gets a dedicated WebSocket connection with state management for audio buffering and processing coordination. We implement connection pooling and graceful reconnection handling.

Audio Buffering

Audio chunks are buffered to smooth jitter and ensure consistent processing. We implement a small buffer (2-3 chunks) that absorbs network latency variations. The buffer size is tunable based on network conditions—larger buffers for unreliable networks, smaller buffers for low-latency requirements.

Async Processing

FastAPI's async capabilities enable concurrent processing of multiple clients. We use asyncio for non-blocking I/O and CPU-bound tasks are offloaded to thread pools. The PyTorch model inference runs in a separate thread pool to avoid blocking the event loop. This architecture scales to handle multiple concurrent avatar sessions.

Latency Optimisation

We optimise for latency at every step: minimise audio chunk size while maintaining quality, use efficient serialisation (binary formats instead of JSON), pre-load models and avoid per-request allocations, and implement priority processing for real-time requests. We also measure end-to-end latency continuously and alert when it exceeds thresholds.

Unity Avatar Integration

The Unity component receives viseme targets and drives avatar blend-shapes.

Blend-Shape Setup

Avatars use standard blend-shapes for visemes (ARKit blend-shapes or custom viseme sets). We support both facial animation standards. The blend-shape setup includes phoneme-specific shapes for common visemes and additional shapes for expressions and co-articulation effects.

Viseme Smoothing

Raw viseme targets can appear jerky. We implement smoothing using interpolation between current and target viseme values. The smoothing factor is tunable—more smoothing for natural appearance, less for responsiveness. We also implement predictive smoothing that anticipates upcoming visemes based on phoneme context.

WebSocket Client

The Unity WebSocket client handles connection management, audio streaming, and viseme target reception. We implement automatic reconnection with exponential backoff. The client also handles audio capture from the microphone using Unity's audio system and streams chunks to the backend.

Fallback Behaviour

Network interruptions are inevitable. We implement fallback behaviour that uses the last known viseme targets with decay to neutral when connection is lost. This prevents the avatar from freezing in an unnatural expression. When connection is restored, the system smoothly transitions back to real-time sync.

Performance Optimisation

Achieving sub-50ms latency requires optimisation across the entire pipeline.

End-to-End Latency Breakdown

We measure latency at each stage: audio capture (5ms), network transmission (10-20ms), backend processing (10-15ms), network return (10-20ms), Unity application (5-10ms). The total is typically 40-70ms depending on network conditions. We optimise each stage to minimise its contribution.

Model Inference Optimisation

Beyond quantisation and pruning, we implement additional optimisations: use TorchScript for faster inference, enable CUDA if GPU is available, implement batched inference for multiple audio frames, and cache intermediate computations where possible. The model is compiled once at startup to avoid JIT overhead during inference.

Network Optimisation

We use binary protocol buffers for efficient serialisation instead of JSON. This reduces payload size and parsing overhead. We also implement compression for audio chunks using lightweight compression algorithms. WebSocket compression is enabled but we measure its impact as compression can sometimes add latency.

Unity Optimisation

In Unity, we minimise blend-shape updates to only what changed, use efficient data structures for viseme state, and apply blend-shape changes in LateUpdate to ensure they are applied after animation but before rendering. We also implement LOD for blend-shapes—use fewer blend-shapes for distant avatars.

Quality Assurance

Quality is critical for user acceptance of lip sync.

Testing Methodology

We test with diverse audio inputs including different speakers, accents, speaking rates, and background noise levels. We measure both objective metrics (phoneme classification accuracy, latency) and subjective quality (naturalness, sync accuracy). Human evaluation is essential for lip sync quality assessment.

Error Analysis

We analyse classification errors to identify patterns—certain phoneme pairs are commonly confused, specific speakers have lower accuracy, or certain audio conditions cause errors. This analysis informs model improvements and training data augmentation.

Continuous Improvement

We collect production data (with user consent) to continuously improve the model. Error cases are added to training data, and the model is periodically retrained. We also A/B test model improvements to ensure they actually improve user experience before deployment.

Deployment Considerations

Production deployment requires attention to scaling, monitoring, and reliability.

Scaling Strategy

The backend scales horizontally using container orchestration. Each instance can handle multiple concurrent sessions. We implement load balancing based on current load rather than simple connection count. The model is small enough that CPU-only instances are sufficient, keeping costs low.

Monitoring

We monitor latency percentiles (p50, p95, p99), error rates, connection counts, and resource utilisation. Alerts fire when latency exceeds thresholds or error rates spike. We also collect user feedback on sync quality to detect quality degradation.

Geographic Distribution

For global user's, we deploy backend instances in multiple regions. Users connect to the nearest region to minimise network latency. Regional instances share model updates through a central model repository, ensuring consistency across regions.

Conclusion

Building real-time lip sync requires careful attention to latency, quality, and reliability. The combination of PyTorch for efficient audio processing, FastAPI for low-latency backend, and Unity for avatar rendering provides a solid foundation. The key is optimisation at every stage and continuous measurement of actual performance.

The result is a system that delivers natural-looking lip sync with sub-50ms latency, enabling immersive avatar experiences in real-time applications. As audio processing models continue to improve and network infrastructure advances, the quality and capabilities of real-time avatar systems will only increase.