Prompt Refinement from Text to Image Generation using Multi model Approach
Contributors
MANJULA R
S.Hemalatha
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
The rapid expansion of digital content across textual, visual, and multimodal platforms has significantly increased the demand for intelligent systems. Text to image models, which produce realistic visuals from prompts, have been made possible by a recent discovery in generative artificial intelligence. An iterative prompt refining framework that automatically enhances the prompt for better image production is proposed in this work. The suggested approach combines several AI models, such as Large Language Models for quick refining, BLIP for picture captioning, and Stable Diffusion for image production. The suggested pipeline uses the user's prompt to create an initial image, which is then examined using BLIP to extract a textual description. This process is repeated across multiple iterations to enhance the prompt quality and generated image results. To evaluate the effectiveness of the proposed approach, the refined prompts are compared with the baseline keyword-based prompt refinement method. The images are evaluated using two metrics: CLIP score and aesthetic score, which measures prompt alignment and visual quality. Experimental results show that the LLM-based refinement approach performs better compared to the baseline refinement method as well as improves the alignment of the generated image and prompts