Meta announced the acquisition of WaveForms AI, an A.I.A., a strategic initiative aimed at strengthening Meta ‘ s technical capacity in VR, AR and voice recognition. WaveForms AI focuses on the development of audio-linguistic models that recognize emotions and reproduce them in voice form, the technology of which is seen as a key step towards achieving the Voice Turing Test.

WaveForms AI was founded in December 2024 by Alexis Conneau, a former OpenAI researcher, and Coralie Lemaitre, a former Google advertising strategy expert, and received $40 million worth of seeder financing from Andressen Horowitz in just months of its existence, valued at $200 million.
WaveForms AI focuses on the development of audio technology that can sense knowledge and generate human voice, with the aim of allowing it to interact to the same level as real people. Conneau has led OpenAI ‘ s advanced GPT-4o voice model development, with technology that achieves low-relay, natural-flowing conversation experiences through direct processing of audio (rather than traditional voice-to-text). The mission of WaveForms is to integrate AI’s interactive “socio-emotional layers” into the voice, to be applied in areas such as virtual assistants, client services, education, etc.

Following this acquisition, the core team of WaveForms, including Conneau and Lemaitre, will join Meta ‘ s newly established Super-Intelligence Laboratory, headed by Alexandr Wang, former CEO of Scale AI. The Meta internal memo indicates that WaveForms technology will be used to enhance Meta ‘ s AI role, Meta AI ‘ s assistant, wearable equipment and audio content creation, promoting more natural, emotional voice interaction.
Meta’s acquisition strategy is closely linked to its long-term vision in the areas of AI and meta-cosm. In 2025, the CEO of Meta, Mark Zuckerberg, made it clear that AI was the company ‘ s top priority, and in 2024 the company invested more than $14 billion in AI infrastructure. The acquisition of WaveForms was another important move following the acquisition in July of PlayAI, a voice-based start-up company, which showed Meta’s centralized distribution of voice technology.
PlayAI focuses on low-delayed text transliteration (TTS) and multilingual voice generation, while WaveForms goes further on emotional recognition and speech synthesis, which together will provide a more immersive audio experience to the VR/AR ecology of Meta.

In the VR/AR area, audio is at the heart of the immersion experience. Meta started three audio-visual AI models (AViTAR, VIDA, VisualVoice) as early as 2022 to optimize voice enhancement and spatial audio matching in VR environments. For example, the Visual Acoustic Matching model can adapt sound to environmental images to match virtual scenes, such as the conversion of living room recordings into sound effects of forest camps.
The addition of WaveForms will further enhance the emotional expression of these models, make the voice of virtual assistants or digital characters more human, enhance the user experience of Meta Quest 3 and future AR glasses (such as the Oakley Smart Glasses, expected to be launched in 2025).