Automated Conversion of Arabic Sign Language to Text Based on the Hybrid CNN-LSTM Model to Enable Communication with the Deaf and Non-Speaking Individuals
DOI:
https://doi.org/10.22153/kej.2026.12.018Keywords:
Deep Learning; Sign language; CNN; LSTM; Deaf and muteAbstract
Communicating with the Deaf and non-speaking individuals and understanding sign language are complicated processes. With the continuous development in deep learning and automation of systems, however, solutions are available to overcome the difficulty of communicating with this segment of society and understanding sign language, especially in Arabic-speaking regions. This article proposes an automated sign language recognition system based on a hybrid convolutional neural network (CNN)–long short-term memory (LSTM) architecture, which represents two layers: CNN, which extracts spatial features from sign language motions, and LSTM, which records the temporal connections within gesture sequences. The model’s ability to provide smooth communication is demonstrated by evaluating its performance using a specially created Arabic Sign Language (ArSL) dataset, and then the model was tested and trained by using input videos recorded for three signers, each one of whom pronounced 50 different signs. We added three other signers to pronounce the days of the week in sign language, improving the accuracy of the pretraining. These videos served as added data to the ArSL dataset. The key point was extracted by using matplotlib, the key point values for training and testing were collected, and the data were preprocessed through feature extraction and batch normalisation. The signs were recognised in the Arabic language. The proposed model detects the sign language action and recognition according to the dataset training, which was achieved building and training by using the LSTM neural network and CNN. The accuracy rating, F-1 score, recall and precision of the proposed model for ArSL were 95%, 95%, 95% and 96%, respectively, with 126 epochs and batch normalisation equal to 30. Ultimately, the proposed model was able to recognise and detect ArSL accurately.
Downloads
References
[1] A. M. Mahmoud Ibrahim and H. H. Kamel, “Social welfare services as a mechanism to reduce social exclusion of the deaf and mute,” Egyptian Journal of Social Work, vol. 14, no. 1, pp. 35–56, 2022, doi: https://doi.org/10.21608/ejsw.2022.129263.1158.
[2] L. Sirch, L. Salvador, and A. Palese, “Communication difficulties experienced by deaf male patients during their in‐hospital stay: findings from a qualitative descriptive study,” Scand. J. Caring Sci., vol. 31, no. 2, pp. 368–377, 2017, doi: https://doi.org/10.1111/scs.12356.
[3] S. Bae, “A Study on the Difficulty of Communication through Sign Language in Non-English Speaking Countries,” Open J. Soc. Sci., vol. 11, no. 10, pp. 494–506, 2023, doi: https://doi.org/10.4236/jss.2023.1110028.
[4] T. Jamil, “Design and implementation of an intelligent system to translate arabic text into arabic sign language,” in 2020 IEEE Canadian Conference on Electrical and Computer Engineering (CCECE), IEEE, 2020, pp. 1–4. doi: https://doi.org/10.1109/CCECE47787.2020.9255774.
[5] S. K. Mahato and R. Jeya, “Convert sign language to text with CNN,” in AIP Conference Proceedings, AIP Publishing, 2024. doi: https://doi.org/10.1063/5.0217230.
[6] A. Baihan, A. I. Alutaibi, M. Alshehri, and S. K. Sharma, “Sign language recognition using modified deep learning network and hybrid optimization: a hybrid optimizer (HO) based optimized CNNSa-LSTM approach,” Sci. Rep., vol. 14, no. 1, p. 26111, 2024, doi: https://doi.org/10.1038/s41598-024-76174-7.
[7] A. A. I. Sidig, H. Luqman, S. Mahmoud, and M. Mohandes, “KArSL: Arabic sign language database,” ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP), vol. 20, no. 1, pp. 1–19, 2021, doi: https://doi.org/10.1145/3423420.
[8] H. Luqman and S. A. Mahmoud, “A machine translation system from Arabic sign language to Arabic,” Univers. Access Inf. Soc., vol. 19, no. 4, pp. 891–904, 2020, doi: https://doi.org/10.1007/s10209-019-00695-6.
[9] A. Khan, A. Sohail, U. Zahoora, and A. S. Qureshi, “A survey of the recent architectures of deep convolutional neural networks,” Artif. Intell. Rev., vol. 53, no. 8, pp. 5455–5516, 2020, doi: https://doi.org/10.1007/s10462-020-09825-6.
[10] V. R. Allugunti, “A machine learning model for skin disease classification using convolution neural network,” International Journal of Computing, Programming and Database Management, vol. 3, no. 1, pp. 141–147, 2022, doi: https://doi.org/10.33545/27076636.2022.v3.i1b.53.
[11] K. Zarzycki and M. Ławryńczuk, “Advanced predictive control for GRU and LSTM networks,” Inf. Sci. (N. Y)., vol. 616, pp. 229–254, 2022, doi: https://doi.org/10.1016/j.ins.2022.10.078 .
[12] B. C. Mateus, M. Mendes, J. T. Farinha, and A. M. Cardoso, “Anticipating future behavior of an industrial press using LSTM networks,” Applied Sciences, vol. 11, no. 13, p. 6101, 2021, doi: https://doi.org/10.3390/app11136101.
[13] S. M. S. Abdullah and A. M. Abdulazeez, “Facial expression recognition based on deep learning convolution neural network: A review,” Journal of Soft Computing and Data Mining, vol. 2, no. 1, pp. 53–65, 2021, doi: https://doi.org/10.30880/jscdm.2021.02.01.006.
[14] V. Passricha and R. K. Aggarwal, “A hybrid of deep CNN and bidirectional LSTM for automatic speech recognition,” Journal of Intelligent Systems, vol. 29, no. 1, pp. 1261–1274, 2019, doi: https://doi.org/10.1515/jisys-2018-0372.
[15] T. Lees et al., “Hydrological concept formation inside long short-term memory (LSTM) networks,” Hydrology and Earth System Sciences Discussions, vol. 2021, pp. 1–37, 2021, doi: https://doi.org/10.5194/hess-26-3079-2022,2022.
[16] B. Lindemann, B. Maschler, N. Sahlab, and M. Weyrich, “A survey on anomaly detection for technical systems using LSTM networks,” Comput. Ind., vol. 131, p. 103498, 2021, doi: https://doi.org/10.1016/j.compind.2021.103498.
[17] T. H. Noor et al., “Real-time arabic sign language recognition using a hybrid deep learning model,” Sensors, vol. 24, no. 11, p. 3683, 2024, doi: https://doi.org/10.3390/s24113683.
[18] K. D. Ismael and E. A. Saeed, “Controlling Mobile Robot Navigation Equipped with Computer Vision Perception for Text using OCR Technique,” in 2024 21st International Multi-Conference on Systems, Signals & Devices (SSD), IEEE, 2024, pp. 1–7. doi: https://doi.org/10.1109/SSD61670.2024.10549002.
[19] A. M. A. Moustafa et al., “Arabic sign language recognition systems: A systematic review,” Indian Journal of Computer Science and Engineering, vol. 15, pp. 1–18, 2024, doi: https://doi.org/10.21817/indjcse/2024/v15i1/241501008.
[20] M. A. Bencherif et al., “Arabic sign language recognition system using 2D hands and body skeleton data,” IEEE Access, vol. 9, pp. 59612–59627, 2021, doi: https://doi.org/10.1109/ACCESS.2021.3069714.
[21] Q. Bani Baker, N. Alqudah, T. Alsmadi, and R. Awawdeh, “Image‐Based Arabic Sign Language Recognition System Using Transfer Deep Learning Models,” Applied Computational Intelligence and Soft Computing, vol. 2023, no. 1, p. 5195007, 2023, doi: https://doi.org/10.1155/2023/5195007.
[22] S. Das, M. Chakraborty, and B. Purkayastha, “A review on sign language recognition (slr) system: Ml and dl for slr,” in 2021 IEEE International Conference on Intelligent Systems, Smart and Green Technologies (ICISSGT), IEEE, 2021, pp. 177–182. doi: https://doi.org/10.1109/ICISSGT52025.2021.00045.
[23] K. Amrutha and P. Prabu, “ML based sign language recognition system,” in 2021 International Conference on Innovative Trends in Information Technology (ICITIIT), IEEE, 2021, pp. 1–6. doi: https://doi.org/10.1109/ICITIIT51526.2021.9399594.
[24] S. Das, S. K. Biswas, and B. Purkayastha, “A deep sign language recognition system for Indian sign language,” Neural Comput. Appl., vol. 35, no. 2, pp. 1469–1481, 2023, doi: https://doi.org/10.1007/s00521-022-07840-y .
[25] A. Wali, R. Shariq, S. Shoaib, S. Amir, and A. A. Farhan, “Recent progress in sign language recognition: a review,” Mach. Vis. Appl., vol. 34, no. 6, p. 127, 2023, doi: https://doi.org/10.1007/s00138-023-01479-y .
[26] N. Ganatra and A. Patel, “A comprehensive study of deep learning architectures, applications and tools,” International Journal of Computer Sciences and Engineering, vol. 6, no. 12, pp. 701–705, 2018, doi: https://doi.org/10.26438/ijcse/v6i12.701705 .
[27] “mArSL-A multimodal manual and non-Manual Arabic sign language dataset.” Accessed: Apr. 27, 2025. [Online]. Available: https://hamzah-luqman.github.io/marsl/#
[28] B. Chen, T.-J. Chin, and M. Klimavicius, “Occlusion-robust object pose estimation with holistic representation,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022, pp. 2929–2939. doi: https://doi.org/10.1109/WACV51458.2022.00228 .
[29] A. K. Singh, V. A. Kumbhare, and K. Arthi, “Real-time human pose detection and recognition using mediapipe,” in International conference on soft computing and signal processing, Springer, 2021, pp. 145–154. doi: https://doi.org/10.1007/978-981-16-7088-6_12 .
[30] A. Chaikaew, “An applied holistic landmark with deep learning for Thai sign language recognition,” in 2022 37th International Technical Conference on Circuits/Systems, Computers and Communications (ITC-CSCC), IEEE, 2022, pp. 1046–1049. doi: https://doi.org/10.1109/ITC-CSCC55581.2022.9895052 .
[31] L. Chen, S. Li, Q. Bai, J. Yang, S. Jiang, and Y. Miao, “Review of image classification algorithms based on convolutional neural networks,” Remote Sens. (Basel)., vol. 13, no. 22, p. 4712, 2021, doi: https://doi.org/10.3390/rs13224712 .
[32] T. Kattenborn, J. Leitloff, F. Schiefer, and S. Hinz, “Review on Convolutional Neural Networks (CNN) in vegetation remote sensing,” ISPRS journal of photogrammetry and remote sensing, vol. 173, pp. 24–49, 2021, doi: https://doi.org/10.1016/j.isprsjprs.2020.12.010 .
[33] Z. Li, F. Liu, W. Yang, S. Peng, and J. Zhou, “A survey of convolutional neural networks: analysis, applications, and prospects,” IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 12, pp. 6999–7019, 2021, doi: https://doi.org/10.1109/TNNLS.2021.3084827 .
[34] A. Shah et al., “A comprehensive study on skin cancer detection using artificial neural network (ANN) and convolutional neural network (CNN),” Clinical eHealth, vol. 6, pp. 76–84, 2023, doi: https://doi.org/10.1016/j.ceh.2023.08.002 .
[35] S. Mascarenhas and M. Agarwal, “A comparison between VGG16, VGG19 and ResNet50 architecture frameworks for Image Classification,” in 2021 International conference on disruptive technologies for multi-disciplinary research and applications (CENTCON), IEEE, 2021, pp. 96–99. doi: https://doi.org/10.1109/CENTCON52345.2021.9687944 .
[36] H. Yang, J. Ni, J. Gao, Z. Han, and T. Luan, “A novel method for peanut variety identification and classification by Improved VGG16,” Sci. Rep., vol. 11, no. 1, p. 15756, 2021, doi: https://doi.org/10.1038/s41598-021-95240-y .
[37] D. Albashish, R. Al-Sayyed, A. Abdullah, M. H. Ryalat, and N. A. Almansour, “Deep CNN model based on VGG16 for breast cancer classification,” in 2021 International conference on information technology (ICIT), IEEE, 2021, pp. 805–810. doi: https://doi.org/10.1109/ICIT52682.2021.9491631 .
[38] S. Sharma and K. Guleria, “A deep learning model for early prediction of pneumonia using VGG19 and neural networks,” in Mobile Radio Communications and 5G Networks: Proceedings of Third MRCN 2022, Springer, 2023, pp. 597–612. doi: https://doi.org/10.1007/978-981-19-7982-8_50 .
[39] A. Karacı, “VGGCOV19-NET: automatic detection of COVID-19 cases from X-ray images using modified VGG19 CNN architecture and YOLO algorithm,” Neural Comput. Appl., vol. 34, no. 10, pp. 8253–8274, 2022, doi: https://doi.org/10.1007/s00521-022-06918-x .
[40] O. C. Do, C. M. Luong, P.-H. Dinh, and G. S. Tran, “An efficient approach to medical image fusion based on optimization and transfer learning with VGG19,” Biomed. Signal Process. Control, vol. 87, p. 105370, 2024, doi: https://doi.org/10.1016/j.bspc.2023.105370
[41] S. A. Hasanah, A. A. Pravitasari, A. S. Abdullah, I. N. Yulita, and M. H. Asnawi, “A deep learning review of resnet architecture for lung disease Identification in CXR Image,” Applied Sciences, vol. 13, no. 24, p. 13111, 2023, doi: https://doi.org/10.3390/app132413111.
[42] N. Jain and P. Peddi, “Gender Classification Model based on the Resnet 152 Architecture,” in 2023 IEEE International Carnahan Conference on Security Technology (ICCST), IEEE, 2023, pp. 1–7. doi: https://doi.org/10.1109/ICCST59048.2023.10474266 .
[43] M. K. Panda, B. N. Subudhi, T. Veerakumar, and V. Jakhetiya, “Modified ResNet-152 network with hybrid pyramidal pooling for local change detection,” IEEE Transactions on Artificial Intelligence, vol. 5, no. 4, pp. 1599–1612, 2023, doi: https://doi.org/10.1109/TAI.2023.3299903
[44] M. Mahasin and I. A. Dewi, “Comparison of CSPDarkNet53, CSPResNeXt-50, and EfficientNet-B0 backbones on YOLO v4 as object detector,” International journal of engineering, science and information technology, vol. 2, no. 3, pp. 64–72, 2022, doi: https://doi.org/10.52088/ijesty.v1i4.291 .
[45] J. Frederich, J. Himawan, and M. Rizkinia, “Skin lesion classification using EfficientNet B0 and B1 via transfer learning for computer aided diagnosis,” in AIP Conference Proceedings, AIP Publishing, 2024. doi: https://doi.org/10.1063/5.0200741 .
[46] S. Abd El-Ghany, M. A. Mahmood, and A. A. Abd El-Aziz, “Adaptive Dynamic Learning Rate Optimization Technique for Colorectal Cancer Diagnosis Based on Histopathological Image Using EfficientNet-B0 Deep Learning Model,” Electronics (Basel)., vol. 13, no. 16, p. 3126, 2024, doi: https://doi.org/10.3390/electronics13163126.
[47] H. A. Sanghvi, R. H. Patel, A. Agarwal, S. Gupta, V. Sawhney, and A. S. Pandya, “A deep learning approach for classification of COVID and pneumonia using DenseNet‐201,” Int. J. Imaging Syst. Technol., vol. 33, no. 1, pp. 18–38, 2023, doi: https://doi.org/10.1002/ima.22812.
[48] F. Salim, F. Saeed, S. Basurra, S. N. Qasem, and T. Al-Hadhrami, “DenseNet-201 and Xception pre-trained deep learning models for fruit recognition,” Electronics (Basel)., vol. 12, no. 14, p. 3132, 2023, doi: https://doi.org/10.3390/electronics12143132.
[49] T. Lu, B. Han, L. Chen, F. Yu, and C. Xue, “A generic intelligent tomato classification system for practical applications using DenseNet-201 with transfer learning,” Sci. Rep., vol. 11, no. 1, p. 15824, 2021, doi: https://doi.org/10.1038/s41598-021-95218-w.
[50] S. Liu et al., “Consmax: Hardware-friendly alternative softmax with learnable parameters,” in Proceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design, 2024, pp. 1–9. doi: https://doi.org/10.1145/3676536.3676766.
[51] M. Franke and J. Degen, “The softmax function: Properties, motivation, and interpretation,” 2023, doi: https://doi.org/10.31234/osf.io/vsw47 .
[52] H. Çetiner and İ. Çetiner, “Classification of cataract disease with a DenseNet201 based deep learning model,” Journal of the Institute of Science and Technology, vol. 12, no. 3, pp. 1264–1276, 2022, doi: https://doi.org/10.21597/jist.1098718.
[53] G. Wang, “RL-CWtrans Net: multimodal swimming coaching driven via robot vision,” Front. Neurorobot., vol. 18, p. 1439188, 2024, doi: http://dx.doi.org/10.3389/fnbot.2024.1439188.
[54] H. Chen and X. Yue, “Swimtrans Net: a multimodal robotic system for swimming action recognition driven via Swin-Transformer,” Front. Neurorobot., vol. 18, p. 1452019, 2024, doi: http://dx.doi.org/10.3389/fnbot.2024.1452019.
[55] X. Ke, X. Zhang, T. Zhang, J. Shi, and S. Wei, “Sar ship detection based on swin transformer and feature enhancement feature pyramid network,” in IGARSS 2022-2022 IEEE International Geoscience and Remote Sensing Symposium, IEEE, 2022, pp. 2163–2166. doi: https://doi.org/10.1109/IGARSS46834.2022.9883800.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Al-Khwarizmi Engineering Journal

This work is licensed under a Creative Commons Attribution 4.0 International License.
Copyright: Open Access authors retain the copyrights of their papers, and all open access articles are distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided that the original work is properly cited. The use of general descriptive names, trade names, trademarks, and so forth in this publication, even if not specifically identified, does not imply that these names are not protected by the relevant laws and regulations. While the advice and information in this journal are believed to be true and accurate on the date of its going to press, neither the authors, the editors, nor the publisher can accept any legal responsibility for any errors or omissions that may be made. The publisher makes no warranty, express or implied, with respect to the material contained herein.






