Google debuts SL2T, an AI model that’s designed to understand sign language

by · SiliconANGLE

Google DeepMind said today it wants to bring the artificial intelligence revolution to the estimated 70 million people across the world who are either deaf or hard of hearing with the launch of sign-language-to-text or SL2T.

In a blog post, Google’s AI researchers said SL2T is a multilingual translation model that’s making its debut on the Pixel 11 smartphone, where it powers a new “sign-to-text” dictation feature on Gboard and Live Transcribe. The model is initially capable of translating American Sign Language or ASL to English text, and is the first of its kind to be made available within a real-world consumer product, the company said.

AI’s ability to process human speech has progressed enormously in the last few years, to the point where anyone can dictate anything they like in dozens of global languages on a smartphone, tablet or personal computer. But the same isn’t true for those who rely on one of the more than 200 distinct sign languages because of their hearing disabilities. According to Google, it’s an audience that has been completely ignored by the AI industry, until now.

The SL2T model gives deaf and hard of hearing users the ability to interact with their smartphone using their native language. Just like speech AI makes it possible for users to talk to their device instead of typing, SL2T makes it possible for people to sign directly into their smartphone’s camera rather than tap away entering text.

It means users can now use sign language to perform dozens of different tasks, including web searches, drafting emails and text messages, editing documents and so on. They can also use SL2T to prompt Gemini to answer queries and perform actions on their behalf. The model is also available in Google’s Live Transcribe app, allowing users to sign their responses directly during face-to-face calls.

Google said SL2T was trained on more than 100,000 hours of sign language data spanning over 50 languages, with about a quarter of that dataset made up of ASL communications. While the initial release can only understand ASL, Google said it decided to train the model across multiple sign languages in order to learn the shared structural patterns across them. In this way, it can significantly outperform earlier sign language models.

To ease users’ privacy concerns, SL2T is powered by an on-device computer vision model called MediaPipe Holistic, which tracks the geometric pose locations across signer’s faces, hands, arms and torso. It then sends the coordinates of these locations to its cloud-based server, avoiding the need to upload any actual video, making it faster and more secure for users.

The model translates sequences directly into text, bypassing the intermediate text annotations known as “glosses.” By doing this, SL2T is better able to capture the non-manual expressions and spatial grammar structures that are characteristic of ASL, Google said. Other features include optimizations to reduce latency, hallucination prevention mechanisms for non-signing movements and support for both left-handed signers and one-handed signing, so users can interact with it while holding their smartphone in one hand.

Google’s performance claims are backed by solid data. SL2T achieved a score of 70 BLEURT on the FLEURS-ASL benchmark.

The release of SL2T is a key development for deaf communities globally, Google said. The development of AI models that can understand sign languages has been slow due to the unique challenges they present and also some common misconceptions. Unlike spoken languages, which map sequential sounds directly to text, sign languages have their own unique lexicons and grammar that must be translated in an entirely new way.

Another difficulty is that sign languages convey meaning through simultaneous movements of not just the hands, but also the head, face, arms and torso. But early attempts to develop sign languages focused on hand gestures only, rather than treating them as full-body visual languages.

Google signaled to the deaf and hard-of-hearing communities that it’s not going to leave them behind any longer. In addition to releasing SL2T, it has also established the AI Sign Language Advisory Committee in partnership with a number of deaf organizations and sign language experts to help guide the responsible deployment of the technology. The committee will also co-author a report alongside Google that outlines SL2T’s capabilities and its current limitations.

Going forward, Google plans to expand the model to cover additional sign languages and also develop models for sign language generation.

Image: Google DeepMind