marketing artificial intelligence
PR Newswire
Published on : Aug 27, 2026
Authentic Interactions AI, the company behind Lookalike, StoryFile and Authentex, has partnered with speech-recognition provider Soniox to add real-time multilingual speech recognition to its digital-human products.
Under the agreement, Soniox's speech-recognition technology will provide the voice-understanding layer for Lookalike and StoryFile. The companies said the integration is intended to improve how digital humans understand conversations involving different languages, accents, mixed-language speech and challenging acoustic conditions.
The partnership illustrates a broader shift in conversational technology from scripted digital characters toward systems designed to support more natural, real-time interaction. For organizations deploying digital humans in museums, educational environments and other public spaces, speech recognition is a critical part of that experience because users cannot be expected to speak in controlled environments or use standardized language.
Soniox says its technology supports more than 60 languages and is designed for low-latency speech recognition. The company will now apply that capability to digital-human experiences developed by Authentic Interactions AI.
Digital humans combine several technologies, including speech recognition, conversational AI, language processing, response generation and visual representation. While much of the industry's attention has focused on the visual realism of virtual presenters, the ability to accurately understand spoken input can be equally important to the user experience.
A digital human that looks realistic but repeatedly misinterprets a visitor's question can quickly undermine the interaction.
That problem becomes more complicated in public installations. Museums and learning centers can contain background conversations, reverberation, music and other environmental noise. Visitors may also speak with regional accents, switch languages within a conversation or use informal phrasing.
Authentic Interactions said its Lookalike and StoryFile products have accumulated real-world conversational data through deployments, including museum exhibits. The company said those experiences influenced its decision to select Soniox for speech recognition.
The companies have not disclosed commercial terms or provided independent benchmark results comparing the combined system with competing speech-recognition platforms.
Soniox's support for more than 60 languages is one of the central elements of the partnership.
Multilingual speech technology has become increasingly important as conversational interfaces move beyond English-language applications. For museums and cultural institutions, supporting multiple languages can potentially allow a single installation to serve international visitors without requiring separate interfaces for each language.
For digital-human platforms, however, language coverage alone does not determine quality. Recognition accuracy across accents, code-switching, background noise and conversational speech can be more important than the number of supported languages.
Real-time performance also matters. In a conventional transcription workflow, a delay of several seconds may be acceptable. In a conversational digital-human application, noticeable latency can make the interaction feel unnatural because users expect a response immediately after speaking.
Soniox's integration therefore addresses two related requirements: understanding what users say and doing so quickly enough to sustain a conversational rhythm.
The partnership also reflects the industry's movement toward deploying conversational AI in physical environments rather than limiting it to websites and mobile applications.
StoryFile has focused on interactive video and conversational experiences, while Lookalike is designed around digital-human interactions. Their use in museums and learning environments creates a different technical challenge from a typical customer-service chatbot.
In a physical installation, users may not know how the system works or what language it expects. Multiple people may speak simultaneously, and microphones can capture environmental noise. The system therefore needs to cope with unpredictable input.
This is where real-world speech data can become valuable. Authentic Interactions said its products have been refined through thousands of conversations and hours of deployment experience. That data, combined with Soniox's speech-recognition technology, could provide a feedback loop for improving conversational performance.
The companies did not disclose the size, composition or methodology of the datasets used to evaluate the partnership, so claims about improved accuracy should be viewed as company statements rather than independently verified performance findings.
The market for conversational digital humans sits at the intersection of several technology categories: generative AI, speech recognition, synthetic media, conversational interfaces and virtual avatars.
Large technology companies and specialized vendors are competing across different layers of this stack. Speech-recognition providers focus on converting spoken language into machine-readable input, while digital-human platforms integrate speech with conversational intelligence and visual presentation.
The competitive differentiator is increasingly moving toward end-to-end interaction quality.
For organizations deploying these systems, a successful experience requires more than an attractive avatar. The platform must recognize users accurately, understand intent, generate appropriate responses, maintain context and deliver those responses with sufficiently low latency.
Multilingual capability adds another layer of complexity because performance can vary considerably across languages and acoustic conditions.
Soniox's partnership with Authentic Interactions therefore represents a vertical integration of two specialized capabilities rather than a standalone avatar announcement.
The integration could help digital-human deployments become more practical in environments where conventional touchscreen or text-based interfaces are less suitable.
Museums, visitor centers, educational institutions and cultural organizations can potentially use conversational interfaces to let visitors interact with historical figures, experts, exhibits or educational content using natural speech.
The bigger opportunity, however, is likely to depend on reliability rather than novelty. As conversational AI becomes more common, users will increasingly expect systems to understand ordinary speech without requiring repeated prompts or carefully worded questions.
That raises the importance of speech-recognition accuracy, latency, multilingual coverage and performance in noisy environments.
For Authentic Interactions, partnering with a specialized speech-recognition provider allows it to focus on the digital-human and conversational experience while using an external technology layer for speech input. For Soniox, the relationship provides another application environment in which real-time multilingual recognition can be tested against unpredictable human interaction.
If the combined system demonstrates measurable improvements in accuracy and responsiveness across real-world deployments, it could strengthen the case for digital humans as practical interfaces rather than primarily experiential demonstrations.
Get in touch with our MarTech Experts
Looking to publish a press release, guest article, interview or podcast? Connect with us.
GET FEATURED