Sign-Speak: Real-Time Continuous Indian Sign Language Translator
Keywords:
Assistive technology , Bi-LSTM, Deep learning, Gesture recognition, Indian Sign Language (ISL), MediaPipe holistic, Sequence learning, Sign language recognitionAbstract
Communication between deaf or hard-of-hearing individuals and people who are unfamiliar with Indian Sign Language (ISL) continues to present challenges in many everyday situations. Recent progress in artificial intelligence has made it possible to develop camera-based systems that recognize sign gestures without relying on wearable devices. This literature survey examines recent research on real-time ISL recognition, with particular attention to landmark-based feature extraction and sequence-learning models used for dynamic gesture interpretation. Four representative studies employing MediaPipe Holistic, Long Short-Term Memory (LSTM), Bidirectional LSTM (Bi-LSTM), Gated Recurrent Unit (GRU), and Transformer-based techniques are critically reviewed. Their methodologies, datasets, reported performance, strengths, and limitations are compared to identify current research trends and unresolved challenges. The analysis reveals that although existing approaches achieve promising recognition accuracy, many are constrained by limited vocabularies, small datasets, signer variability, and insufficient support for continuous sentence-level recognition. Based on these observations, this survey proposes a Sign-Speak framework that combines MediaPipe Holistic for extracting hand, facial, and body landmarks with a Bi-LSTM network for learning temporal gesture patterns. The recognized signs are intended to be converted into meaningful text and speech, enabling more accessible communication in real-world environments. The findings highlight the potential of vision-based deep learning systems to improve inclusive human-computer interaction while identifying opportunities for future research in multilingual translation, larger-scale datasets, and practical deployment.
References
A. Bhadouria, P. Bindal, N. Khare, D. Singh, and A. Verma, “LSTM-based recognition of sign language,” IC3-2024: Proceedings of the 2024 Sixteenth International Conference on Contemporary Computing, Oct 2024, pp. 508–514.
A. Tripathi, S. Makhloga, S. Singh, S. Semwal, and V. Tomar, “SLRMPCMC: Sign language recognition using MediaPipe and cross-model comparison,” 2024 International Conference on Electrical Electronics and Computing Technologies (ICEECT), Aug 2024, pp. 1–6.
M. Al-Qurishi, T. Khalid, and R. Souissi, “Deep learning for sign language recognition: Current techniques, benchmarks, and open issues,” IEEE Access, vol. 9, pp. 126917–126951, 2021
M. Geetha, N. Aloysius, D. A. Somasundaran, A. Raghunath, and P. Nedungadi, “Toward real-time recognition of continuous Indian Sign Language: A multi-modal approach using RGB and pose,” IEEE Access, vol. 13, pp. 60270–60283, Mar 2025.
P. Rawat, P. Kumar, V. K. Tamta, and A. Kumar, “A comprehensive approach to Indian Sign Language recognition: Leveraging LSTM and MediaPipe Holistic for dynamic and static hand gesture recognition,” EAI Endorsed Transactions on AI and Robotics, vol. 4, May 2025.
N. Anithadevi, S. Palanisamy, S. S. Rubini, and S. Shrestha, “MediaPipe-LSTM-enhanced framework for real-time dynamic sign language recognition in inclusive communication systems,” Engineering Reports, vol. 7, no. 7, Jul 2025.
B. Subramanian, B. Olimov, S. M. Naik, S. Kim, K.-H. Park, and J. Kim, “An integrated MediaPipe-optimized GRU model for Indian Sign Language recognition,” Scientific Reports, vol. 12, Jul 2022.
V. Ravikiran, “Real-time sign language recognition and translation using MediaPipe and LSTM-based deep learning,” International Journal of Computer Applications, vol. 187, no. 25, pp. 10–14, Jul. 2025.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation, vol. 9, no. 8, pp. 1735–1780, Nov 1997.
A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional LSTM networks,” Proceedings of the 2005 IEEE International Joint Conference on Neural Networks (IJCNN), vol. 4, pp. 2047–2052, Jul 2005.
R. Damdoo, P. Kumar, and R. Gogoi, “End-to-end sentence-level Indian Sign Language translation with ISH-NEWS dataset and transformer model,” Scientific Reports, Jul 2026.
P. Nedungadi, G. Dileep, M. Geetha, and R. Raman, “Vision transformer-powered conversational agent for real-time Indian Sign Language e-governance accessibility,” Scientific Reports, vol. 15, Sep 2025.
S. Patra and S. Samanta, “STARK: Spatio-temporal attention for representation of keypoints for continuous sign language recognition,” arXiv preprint arXiv:2603.16163, Mar 2026.