Towards Inclusive Digital Accessibility Using Multimodal AI and Braille-Aware Language Models
Contributors
Ajantha Devi Vairamani
Prof. (Dr.) Anand Nayyar
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Digital access is a vital issue for visually impaired people today, especially in a digital world that is based on multimodal communication (text, voice, visual, infographics, statistics, equations, and formatted documents). Existing devices including screen readers, text systems, Optical Character Recognition, and rule-based Braille translation have enhanced ease of reading from text, but these are few when it comes to complex visual and context-rich content. For many high-level users of technology, this has resulted in a complete lack of tools to help them understand text with visual and contextual information. Those constraints are especially severe in science, technology, engineering and mathematics school education because learning resources in these subjects often rely on visual and spatial representations. Recent advancements in artificial intelligence, natural language processing, computer vision, large language models and multimodal learning offer novel possibilities for intelligent accessibility systems with semantic understanding, multimodal interpretation and personalized content generation.
This paper examines and reviews the available literature on accessibility in the digital realm and assistive digital technologies focusing on AI-based accessibility, multimodal AI, and adaptive Braille generation. The review concludes that: Limited contextual awareness; Poor multimodal processing; Poor personalization; Low adaptive Braille generation support. To close this gap, this paper urges the design of a Multimodal Accessibility Engine , an AI-based methodology to bridge the gap of non-intermezzo-language Braille awareness by harnessing multimodes, including multimodal content comprehension, semantic reasoning, and user-adapted Braille generation for inclusive digital participation by visually impaired users.