Tri-lingual Sentiment Detection of Code-Mixed Marathi, Hindi, and English through Hybrid Multilingual Model


Date Published : 8 September 2026

Contributors

Dr. Aniruddha Diliprao Shelotkar

Author

Dr. Sai Kiran Oruganti

Author

Keywords

Code-Mixed Text Multilingual BERT (mBERT) Hybrid Model Script-Aware Attention Tri-lingual Sentiment Analysis

Proceeding

Track

General Track

License

Copyright (c) 2026 Sustainable Global Societies Initiative

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Abstract

Sentiment analysis (SA) is a significant phenomenon in Natural Language Processing (NLP) that categorizes the emotions that are expressed in the text. However, the issue of code-mixed and multi-script data analysis, most of the times in the multilingual environment is also very problematic. Elsewhere in the globe like in India where languages like Marathi, Hindi and English are widely spoken and a change between them is a normal occurrence like during the communication process, this code-switching adds to complexity of sentiment analysis due to language diversity, use of variation in script, and sentiment expressions depending on the situation. In the given paper, the author proposes a new hybrid sentiment analysis model that is specifically applied to code-mixed texts (Marathi, Hindi, and English). The presented model relies on the Multilingual BERT (mBERT) framework and integrates a script-conscious attention that is based on entropy and the possibility to work with multilingual input. This mechanism of attention is dynamic and changes emphasis on different languages in the input during a code-mixing process in such a way that language-specific characteristics can be easily replicated. The model comprises three specialists, one of them is a specialist in the comprehension of the syntactic and semantic specifics of each language, which specializes in Marathi, another in Hindi, and the others in English. The output of these specialists is pooled together in a feature fusion layer where the features of the language are combined to one representation to be applied in sentiment classification. Our method is also more efficient in processing input in multi-script and detecting sentiment in mixed-language input (which is often overlooked by other models) than the conventional sentiment analysis models. The model validation is done with a recently curated tri-lingual dataset and the avoidance of only baseline methods, or the existing multilingual models, including MuRIL and XLM-R, to operate with more complex code-switched data. The results are that our hybrid model is highly sensitive and sound, especially in detecting positive, neutral and negative sentiments in informal communication. The current work provides points of entry into the better sentiment analysis of multilingual sentiments, which can provide more comprehensive solution to the real-life use of multilingual settings, including social media monitoring, analysis of customer reviews and sentiment-based decision making.

References

No References

Downloads

How to Cite

Shelotkar, A., & Oruganti, S. K. . (2026). Tri-lingual Sentiment Detection of Code-Mixed Marathi, Hindi, and English through Hybrid Multilingual Model. Sustainable Global Societies Initiative, 1(9). https://vectmag.com/sgsi/paper/view/1005