Skip to content

Google unveils SL2T AI for sign-language-to-text translation

Person using sign language on video call with smartphone on stand beside open laptop on wooden desk.

Google has just unveiled sign-language-to-text, or SL2T, an AI system capable of translating sign language from video into text almost in real time. For now, the new feature is restricted to American Sign Language, although Google says support for other languages is coming soon.

Yesterday, Google introduced its new Google Pixel smartphones. Alongside the Pixel 11, its variants and new accessories, the company also announced several notable software developments. These include Tap to share, an Android feature that lets two people exchange contact details by bringing their smartphones close together.

Google also revealed new health technologies, including an algorithm that identifies respiratory emergencies and automatically contacts emergency services if the user can no longer respond. On the accessibility front, the company is launching its sign-language-to-text AI model, SL2T, which can convert sign language into text nearly in real time.

A sign-language translator

Put simply, this AI receives sign language through a video feed and produces the matching text. Google created the feature so that deaf and hard-of-hearing people can engage more naturally with its products, such as Gemini, instead of having to rely on a keyboard all the time.

The feature can also act as an interpreter between a deaf or hard-of-hearing person and someone who does not understand sign language. Its main drawback is that access to this new AI remains very limited for the time being.

Limited access for now

“SL2T enables sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language (ASL) to English,” Google says. However, it also promises to support more devices and more languages “soon”. Regarding languages, Google says it has already worked with 100,000 hours of data covering more than 50 different sign languages, with a quarter of that data in American Sign Language.

As for devices, the feature may require a certain amount of on-device computing power. SL2T does, in fact, operate partly locally. A local model turns the signer’s video into geometric coordinates. Those geometric points are then converted into text on Google’s servers. The benefit of this approach is improved privacy, since the video streams are not processed in the cloud.

70 million people affected

An estimated 70 million people worldwide are deaf or hard of hearing. According to Google, while AI’s ability to process spoken language has advanced rapidly, this “revolution” has not yet benefited people who need to use sign language.

“Just as hearing users can use voice dictation to speak instead of typing text, this feature enables deaf users to communicate in sign language on their phone wherever they would normally type text. You can use sign language to search the web, write messages or documents, and ask Gemini to answer your questions or carry out tasks,” explains the Mountain View company.

Google recently also shared data on how people use Gemini, which has just surpassed one billion monthly active users. The company reported a rise in the popularity of voice interactions with AI. Indeed, 63% of Gemini users speak directly to the AI. Some no longer use a keyboard at all to interact with it.

Comments

No comments yet. Be the first to comment!

Leave a Comment