Meta Platforms has introduced an advanced AI model capable of translating and transcribing speech in numerous languages, potentially paving the way for seamless communication across language barriers.
The company, in a recent blog post, unveiled its SeamlessM4T AI model, designed to facilitate translations between text and speech across nearly 100 languages. Impressively, it also offers complete speech-to-speech translation for 35 languages, merging functionalities that were previously confined to distinct models.
Mark Zuckerberg, the CEO, envisions these tools as pivotal for enabling interactions among users worldwide within the metaverse, the interconnected realm of virtual worlds that Meta is banking on for its future.
The model is being made available to the public for non-commercial use, according to the blog post.
Throughout the year, the social media giant has released an array of AI models, most of which are free. Among these is Llama, a substantial language model that poses a noteworthy challenge to proprietary models offered by Google and OpenAI.
Zuckerberg contends that an open AI ecosystem benefits Meta, as the company finds more value in harnessing crowd-sourced efforts to develop consumer-centric tools for its social platforms rather than monetizing access to the models.
Yet, like others in the industry, Meta is grappling with legal concerns surrounding the usage of training data for model creation.
In July, comedian Sarah Silverman and two other authors filed copyright infringement suits against both Meta and OpenAI, alleging unauthorized utilization of their books as training data.
For the SeamlessM4T model, Meta’s researchers shared in a research paper that they compiled audio training data from 4 million hours of “raw audio” derived from publicly available web repositories, though the specific repository wasn’t mentioned.
Regarding text data, the research paper indicated that datasets from the previous year were used, drawing content from Wikipedia and affiliated website

