This article discusses architectural considerations for building a robust live captioning system, particularly focusing on handling client reconnections gracefully. It emphasizes separating concerns between real-time data flow, persistent session state, and durable transcripts to ensure data integrity and a reliable user experience.
Read original on Dev.to #architectureBuilding real-time features like live captions in healthtech sessions presents unique challenges, especially when dealing with unreliable client connections. The core problem addressed is how to ensure accurate caption delivery and placement, even if a client temporarily disconnects and then reconnects. The proposed solution involves a strong separation of concerns to prevent data inconsistencies.
The architecture advocates for three distinct systems, each with a specific responsibility:
Don't Overload the Real-time Channel
A common mistake is trying to make the real-time communication channel (e.g., WebSockets, pub/sub) responsible for presence, historical state, and durable records. This article highlights that a transient room should not become all three systems. Instead, let the real-time channel focus solely on fresh data delivery, and offload state management and durability to specialized services.
Presence information (who is currently connected) is crucial for UI decisions like showing active users, but it's insufficient for recovery. When a client reconnects, it cannot assume it received all prior events. The proposed robust recovery mechanism involves:
Each caption segment should carry essential metadata to ensure correct ordering and placement. A simple, vendor-neutral data contract is crucial:
from dataclasses import dataclass
@dataclass(frozen=True)
class CaptionSegment:
session_id: str
segment_id: str
sequence: int
start_ms: int
text: str
def validate(self) -> None:
if self.sequence < 0 or self.start_ms < 0:
raise ValueError("sequence and start_ms must be non-negative")
if not self.text.strip():
raise ValueError("caption text must not be empty")
The `sequence` provides a monotonic ordering for duplicate detection and ordering, while `start_ms` anchors the caption to the session timeline. Keeping segments short, ideally aligned with transcription stream boundaries and original media timestamps, enhances readability.