Synthesia, a video-generation startup valued at $4 billion, has expanded its enterprise product suite by introducing interactive AI avatars capable of conducting dynamic sales training and responding to real-time press inquiries. Operating from offices in London and New York, the digital avatar platform crossed $100 million in annual recurring revenue last year, establishing itself alongside competitors like D-ID, HeyGen, and Colossyan in the enterprise generative AI market.
Three core pillars of the enterprise product ecosystem
Synthesia builds three core product categories for enterprise customers: a standard video-creation platform, an agentic platform called Sessions, and an API platform. The traditional platform allows users to type scripts that classic avatars recite in generated videos.
Meanwhile, the Sessions product line offers interactive capabilities designed for employee surveys and roleplay scenarios, enabling workers to practice sales pitches with an AI avatar that responds and scores their delivery.
Multi-modality and flexible enterprise tech stacks
The underlying technical architecture combines multiple AI modalities to power these interactive models. According to company specifications, the tech stack integrates voice-to-text models to transcribe user speech, agentic language models to process text and execute tasks, text-to-voice models to generate audio responses, and proprietary video models to animate the avatar.
Enterprises can utilize Synthesia’s native voice and video models or select alternative providers such as Cartesia, ElevenLabs, Google, and OpenAI, while choosing between self-hosted cloud environments or Synthesia’s hosting services.
Virtual avatars enter corporate communications
The company deployed its technology for external communications when Alexandru Voica, head of corporate affairs at Synthesia, created an interactive virtual avatar trained to answer standard press questions about the startup’s operations.
Expanding on this approach during an office opening in New York, the firm produced personal and interactive digital twins for visiting journalists. The custom avatars utilize voice recordings and photo captures gathered in an on-site film studio, requiring explicit subject consent before generation.
Studio production and strict output constraints
Development of a custom interactive avatar takes several days and involves specific constraints to maintain deterministic outputs.

When tested against specific journalistic content—such as research regarding venture-backed startup fraud—the interactive avatar relies on designated training data to restrict responses strictly to the source material, redirecting off-topic inquiries back to the core subject matter.
Related reading