Google DeepMind has introduced its sign-language-to-text (SL2T) model, marking the first time sign language AI has been integrated into consumer phone features. The SL2T model, which debuted on Pixel 11 phones on August 12, 2026, powers sign-to-text dictation within Gboard and Live Transcribe. This functionality initially supports American Sign Language (ASL) to English translation.

The new feature allows ASL signers to dictate web searches, draft messages, create documents, or query Google's Gemini chatbot by signing to the phone's camera instead of typing. In Live Transcribe, users can respond in signed conversation rather than typing back and forth. Google reported that testers found signing in ASL to be faster and more natural than typing in English.

DeepMind developed SL2T as a breakthrough in the quality and generality of sign language translation. The model was trained on over 100,000 hours of signing data, encompassing more than 50 sign languages. Approximately a quarter of this training data was in ASL. DeepMind's experiments indicate that training across multiple languages, dialects, and proficiency levels resulted in a model that outperforms single-language systems.

The system employs an on-device model, MediaPipe Holistic, to track 130 key points on the face, body, and hands frame by frame. Only these geometric coordinates are sent off the device for translation, with the original camera feed discarded immediately to safeguard user privacy. From this coordinate stream, SL2T directly generates English text. This approach bypasses the intermediate annotation layer known as glosses, which prior sign language translation systems often relied upon. Glosses, which assign fixed labels to individual signs, can limit vocabulary and fail to capture the non-manual markers and spatial constructions that convey grammatical meaning in sign languages. Eliminating this bottleneck allows translation quality to improve with increased training data.

The release of SL2T includes a governance framework established by an advisory body, the AI Sign Language Advisory Committee. This committee, co-authored a joint impact report outlining the intended uses and limitations of the technology. The committee includes representatives from Deaf organizations such as the National Association of the Deaf, the World Federation of the Deaf, the Deaf Professional Arts Network, and Rochester Institute of Technology's National Technical Institute for the Deaf.

The report classifies SL2T 1.0 as an assistive draft-generation tool suitable for informal, low-stakes settings, including messaging, search, navigation, note-taking, and casual one-on-one conversations. It explicitly advises against using the tool for medical consultations, legal proceedings, police interactions, classroom instruction, job interviews, or government benefits determinations. The report also clarifies that the tool does not fulfill legal obligations for reasonable accommodations under the Americans with Disabilities Act or similar frameworks, and institutions must not use it as a substitute for certified human interpreters.

Users remain in control throughout the process. Both Gboard and Live Transcribe display a text preview that the user can review and edit before sending. In Live Transcribe, the Deaf user controls when translated text is shown to a conversation partner. The report acknowledges that this safeguard assumes functional English literacy to identify and correct errors.

The impact report also details the model's technical limitations. Accuracy can decrease in conditions such as low light, harsh backlighting, or when the signer is partially obscured. The model may also exhibit residual hallucination behavior, generating text when no one is signing, such as when a second person enters the frame or the signer pauses. Google states that the release version has significantly reduced these errors compared to pre-release builds tested by Deaf evaluators in July 2026, but has not eliminated them.

This initiative follows Google DeepMind's earlier announcement of SignGemma in May 2025, an open model designed to translate sign languages into spoken language text. SignGemma is intended to become part of the Gemma family of models, with an initial focus on American Sign Language (ASL) paired with English text output. DeepMind has been inviting feedback and participation from developers, researchers, and Deaf and hard-of-hearing communities for SignGemma's development.