Google DeepMind Brings Sign-Language-to-Text to Gboard and Live Transcribe, Says It’s for Low-Stakes Use

GOOGL GOOG

·

Google DeepMind said Wednesday it is rolling out a new sign-language-to-text system in Gboard and Live Transcribe, but the bigger story is how narrowly the company says it should be used. In a blog post and a co-authored advisory report published Aug. 12, Google and its partners described the tool as an accessibility feature for everyday, low-stakes communication, not for medical, legal or other situations where errors could carry serious consequences.

That framing matters because the feature puts sign input into mainstream phone software used by millions, aimed at Deaf and hard-of-hearing people. But Google is also attaching clear guardrails around privacy, accuracy and legal use. The advisory report says the product should not be used by third parties to avoid accessibility obligations or to replace certified human interpreters where those are required.

Google calls the system SL2T, short for sign-language-to-text, and describes it as a “massively multilingual” translation model. It is being integrated into Google’s Gboard keyboard and Live Transcribe, with initial availability first on Pixel 11 phones. At launch, the system supports American Sign Language to English, with more devices and additional languages planned at no extra cost. Google said the intended uses include typing and dictation, search queries, messaging, draft composition and informal one-on-one conversations such as ordering at a café.

In its blog post, Google DeepMind’s Sign Language Team said, “Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.” The company said the model was trained on more than 100,000 hours of data across more than 50 sign languages, with roughly a quarter of the data in ASL.

A key part of the deployment is how the system handles camera input. Google said SL2T does not take raw video as the model input. Instead, MediaPipe Holistic, a Google framework for tracking body movement, extracts 2D pose landmarks on the device. The deployed system uses 130 keypoints across the face, body and hands, and only those landmark coordinates are sent to Google’s servers for translation, while the original video is discarded. “To protect user privacy, SL2T sees sign language as a sequence of pose landmark locations rather than a raw camera feed,” the blog post said.

The AISLAC Joint Impact Report for SL2T 1.0, published the same day, added that “No logs of user inputs or system outputs are retained on Google’s side, unless explicitly authorized by the user in the context of a model evaluation study.” Google also said SL2T translates directly from landmark sequences into text, rather than relying on intermediate gloss annotations, a notation system sometimes used to represent signs in words.

On performance, Google said SL2T is “the most capable sign language translation model to date” on benchmarks including FLEURS-ASL. Google and the AISLAC report cited a BLEURT score of 70 on the FLEURS-ASL zero-shot sd-test. The company also said it worked to reduce streaming latency, limit hallucinations on non-signing inputs, improve fairness for left-handed signers and support one-handed signing so people can hold a phone while signing.

Just as important, Google says the outputs are meant to be drafts that users can review and edit. In Live Transcribe, the translated text is not automatically broadcast; users decide when text is shown or voiced. That design reflects the system’s limits as much as its capabilities.

Those limits are spelled out in the impact report. High-stakes settings are explicitly out of scope, including medical and legal contexts. Google and AISLAC also acknowledged that the system can make mistakes with rare signs, rapid fingerspelling, classifier depictions and tense. The report said 2D landmarks miss some information entirely, including tongue movement, depth and contact distinctions, and environmental context. It also assumes users can read and verify the English output.

The report frames the release as part of a broader accessibility toolkit, not a replacement for human support. “This release represents a step toward digital language equity and linguistic autonomy,” it said. “The SL2T 1.0 features add to the Deaf community’s portfolio of communication options, empowering signers to choose the modality that best fits their needs.”

That distinction also matters legally. Under U.S. accessibility rules, including the Americans with Disabilities Act, organizations may still need to provide qualified interpreters or other effective aids. The AISLAC report explicitly says SL2T 1.0 is not a substitute for those obligations.

Sign-language recognition and translation technology is not new; research groups and startups, including earlier Microsoft Kinect-based efforts, have worked on it for years. What stands out here is Google’s combination of a scaled multilingual model, a landmark-based privacy design, integration into widely used consumer apps and a formal advisory process that puts guardrails front and center.

Tags: #accessibility, #signlanguage, #google, #deepmind

Stocks: GOOGL GOOG