Fine-Tuning Multilingual BERT for Hindi Text Classification: Sentiment Analysis and Topic Categorisation Using the HindiSentiment-6 Corpus
Author(s):Kavita Shukla, Meenakshi Bisht, Saurabh Dewangan
Affiliation: Department of Artificial Intelligence and Machine Learning, Patna Women's College (Autonomous), Patna, Bihar, India
Page No: 150-156
Volume issue & Publishing Year: Volume 3, Issue 5, May 2026
Journal: International Journal of Advanced Engineering Application (IJAEA)
ISSN NO: 3048-6807
DOI: https://doi.org/10.5281/zenodo.20306128
Download PDF Cite this article
Abstract:
Hindi is the most widely spoken language in India with approximately 528 million native speakers and 600 million total speakers, yet natural language processing (NLP) resources and pre-trained language models for Hindi remain substantially less developed than those for English, limiting the deployment of AI-driven text analysis applications in governance, healthcare, education, and digital commerce in Hindi-speaking markets. Transformer-based pre-trained language models — particularly multilingual BERT (mBERT) and its Hindi-specific variant Hindi-BERT — offer a transfer learning pathway for Hindi NLP tasks, but systematic comparison of fine-tuning strategies, domain generalisation, and performance across classification tasks remains limited in the published literature for Indian language NLP. This paper presents a comprehensive evaluation of fine-tuned Hindi-BERT for two text classification tasks: six-class topic categorisation (politics, sports, entertainment, technology, health, business) and three-class sentiment analysis (positive, negative, neutral) using the newly constructed HindiSentiment-6 corpus — a 12,000-document Hindi text dataset scraped from news portals (Dainik Bhaskar, Amar Ujala, Navbharat Times), social media (Twitter/X Hindi accounts), and e-commerce review platforms (Flipkart, Amazon India). Fine-tuned Hindi-BERT achieves 95.6% accuracy and 95.3% macro-F1 on topic classification and 91.2% accuracy on sentiment analysis across five domains — outperforming TF-IDF+SVM (78.4%), fastText (83.1%), character-level CNN (86.2%), BiLSTM (88.7%), and frozen mBERT (91.3%) baselines. Attention weight visualisation confirms the model captures sentiment-bearing words (rohchak: interesting; prernadayak: inspiring; bekar: useless) as high-attention tokens consistent with human linguistic intuition. The HindiSentiment-6 corpus and fine-tuned model weights are released publicly to support the Hindi NLP research community.
Keywords: Hindi NLP, BERT, multilingual transformers, text classification, sentiment analysis, topic categorisation, fine-tuning, HindiSentiment-6, transfer learning, low-resource NLP
Reference:
- [1] Agarwal, B., Poria, S., Mittal, N., Gelbukh, A., & Hussain, A. (2017). Concept-level sentiment analysis with dependency-based semantic parsing: A novel approach. Cognitive Computation, 7(4), 487-499.
- [2] Akhtar, M. S., Kumar, A., Ekbal, A., & Bhatt, C. A. (2016). A hybrid deep learning architecture for sentiment analysis. Proceedings of COLING 2016, 482-493.
- [3] Balamurali, A. R., Joshi, A., & Bhattacharyya, P. (2012). Cross-lingual sentiment analysis for Indian languages using linked WordNets. Proceedings of COLING 2012, 73-82.
- [4] Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., ... & Stoyanov, V. (2020). Unsupervised cross-lingual representation learning at scale. Proceedings of ACL 2020, 8440-8451.
- [5] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL 2019, 4171-4186.
- [6] Jain, S., & Wallace, B. C. (2019). Attention is not explanation. Proceedings of NAACL 2019, 3543-3556.
- [7] Joulin, A., Grave, E., Bojanowski, P., & Mikolov, T. (2017). Bag of tricks for efficient text classification. Proceedings of EACL 2017, 427-431.
- [8] Kakwani, D., Kunchukuttan, A., Golla, S., Gokul, N. C., Bhattacharyya, A., Khapra, M. M., & Kumar, P. (2020). IndicNLPSuite: Monolingual corpora, evaluation benchmarks and pre-trained multilingual language models for Indian languages. Findings of EMNLP 2020, 4948-4961.
- [9] Kim, Y. (2015). Character-aware neural language models. Proceedings of AAAI 2016.
- [10] Kunchukuttan, A., Mehta, P., & Bhattacharyya, P. (2018). The IIT Bombay English-Hindi parallel corpus. Proceedings of LREC 2018.
- [11] Pires, T., Schlinger, E., & Garrette, D. (2019). How multilingual is multilingual BERT? Proceedings of ACL 2019, 4996-5001.
- [12] Veena, G., & Gupta, D. (2021). Sentiment analysis of Hindi text using BERT. International Journal of Advanced Computer Science and Applications, 12(7), 428-435.
📚 Explore Our Related Journals
Looking for the right journal for your next manuscript? Explore our international peer-reviewed journals covering multidisciplinary research, engineering, management, computer science and artificial intelligence.
IJAMA
International Journal of Advanced Multidisciplinary Application
Publishes peer-reviewed research articles in Engineering, Management, Computer Science, Artificial Intelligence, Science, Humanities, Social Sciences and multidisciplinary research.
IJMEM
International Journal of Modern Engineering and Management
Publishes peer-reviewed research articles in Engineering, Management, Computer Science, Artificial Intelligence, Technology and multidisciplinary research.