Abstract
Image super-resolution is a technique aimed at enhancing the resolution of images and is widely utilized in fields such as remote sensing and medical imaging. We view image super-resolution as a reconstruction problem for inferring missing structural information from incomplete observations. This perspective highlights the need for methods that do not merely generate pixel-level corrections, but instead analyze and infer latent geometric and semantic patterns embedded in multimodal data. In this work, we present PromptFusionSR (PF-SR), a novel framework for text-assisted image super-resolution. In contrast to conventional approaches that rely solely on visual features, our method leverages the BLIP3 vision language model to extract contextual prompts-short textual descriptions-directly from low-resolution images. A periodic semantic–structural fusion mechanism enables the model to integrate complementary information channels, enhancing its ability to interpret high-frequency patterns under severe degradation. These textual and information cues guide a pretrained stable diffusion model for 4\(\times \) upscaling, yielding high-resolution outputs with improved fidelity and richer visual textures. This work further explores how complementary information channels can be systematically integrated during image reconstruction to improve robustness and perceptual fidelity. On our benchmark datasets, our method surpasses the current state-of-the-art methods across most evaluation metrics, with robust gains in perceptual quality metrics.
Access this chapter
Tax calculation will be finalised at checkout
Purchases are for personal use only
Similar content being viewed by others
References
Freeman, W.T., Pasztor, E.C., Carmichael, O.T.: Learning low-level vision. Int. J. Comput. Vision 40, 25–47 (2000)
Park, S.C., Park, M.K., Kang, M.G.: Super-resolution image reconstruction: a technical overview. IEEE Signal Process. Mag. 20(3), 21–36 (2003)
Mena, F., Arenas, D., Dengel, A.: Missing data as augmentation in the earth observation domain: a multi-view learning approach. Neurocomputing 638, 130175 (2025)
Vexler, J., Vieten, B., Nelke, M., Kramer, S.: Integrating inverse and forward modeling for sparse temporal data from sensor networks. In: International Symposium on Intelligent Data Analysis, pp. 318–329. Springer (2025)
Liu, F., Hu, X., Bu, C., Yu, K.: Fuzzy Bayesian knowledge tracing. IEEE Trans. Fuzzy Syst. 30(7), 2412–2425 (2021)
Wang, Z., Chen, J., Hoi, S.C.: Deep learning for image super-resolution: a survey. IEEE Trans. Pattern Anal. Mach. Intell. 43(10), 3365–3387 (2020)
Persson, D., Wahlberg, W., Vettoruzzo, A., Nowaczyk, S.: Bridging spatial and temporal contexts: sparse transfer learning. In: Krempl, G., Puolamäki, K., Miliou, I. (eds.) Advances in Intelligent Data Analysis XXIII, pp. 330–342. Springer Nature Switzerland, Cham (2025)
Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D.J., Norouzi, M.: Image super-resolution via iterative refinement. IEEE Trans. Pattern Anal. Mach. Intell. 45(4), 4713–4726 (2022)
Wang, J., Yue, Z., Zhou, S., Chan, K.C., Loy, C.C.: Exploiting diffusion prior for real-world image super-resolution. Int. J. Comput. Vision 132(12), 5929–5949 (2024)
Xue, L., et al.: xGen-MM (BLIP-3): a family of open large multimodal models. arXiv preprint arXiv:2408.08872 (2024)
Zhao, Y., Prasad, M., Braytee, A.: DualPrompt-MedCap: a dual-prompt enhanced approach for medical image captioning. In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 205–215. Springer (2025)
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10684–10695 (2022)
Lim, B., Son, S., Kim, H., Nah, S., Mu Lee, K.: Enhanced deep residual networks for single image super-resolution. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 136–144 (2017)
Zhang, Y., Li, K., Li, K., Wang, L., Zhong, B., Fu, Y.: Image super-resolution using very deep residual channel attention networks. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 286–301 (2018)
Chen, C., et al.: Real-world blind super-resolution via feature matching with implicit high-resolution priors. In: Proceedings of the 30th ACM International Conference on Multimedia, pp. 1329–1338 (2022)
Lin, X., et al.: DiffBIR: toward blind image restoration with generative diffusion prior. In: European Conference on Computer Vision, pp. 430–448. Springer (2024)
Yang, T., Wu, R., Ren, P., Xie, X., Zhang, L.: Pixel-aware stable diffusion for realistic image super-resolution and personalized stylization. In: European Conference on Computer Vision, pp. 74–91. Springer (2024)
Wu, R., Yang, T., Sun, L., Zhang, Z., Li, S., Zhang, L.: SeeSR: towards semantics-aware real-world image super-resolution. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 25456–25467 (2024)
Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. In: Advances in Neural Information Processing Systems, vol. 33, pp. 6840–6851 (2020)
Agustsson, E., Timofte, R.: NTIRE 2017 challenge on single image super-resolution: dataset and study. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, July 2017
Wei, P., et al.: Component divide-and-conquer for real-world image super-resolution. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020, Proceedings, Part VIII 16, pp. 101–117. Springer (2020)
Cai, J., Zeng, H., Yong, H., Cao, Z., Zhang, L.: Toward real-world single image super-resolution: a new benchmark and a new model. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3086–3095 (2019)
Hore, A., Ziou, D.: Image quality metrics: PSNR vs. SSIM. In: 2010 20th International Conference on Pattern Recognition, pp. 2366–2369. IEEE (2010)
Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 586–595 (2018)
Ding, K., Ma, K., Wang, S., Simoncelli, E.P.: Image quality assessment: unifying structure and texture similarity. IEEE Trans. Pattern Anal. Mach. Intell., 1 (2020). https://doi.org/10.1109/tpami.2020.3045810
Mittal, A., Soundararajan, R., Bovik, A.C.: Making a “completely blind” image quality analyzer. IEEE Sig. Process. Lett. 20(3), 209–212 (2012)
Author information
Authors and Affiliations
Corresponding author
Editor information
Editors and Affiliations
Ethics declarations
Disclosure of Interests
The authors have no competing interests to declare that are relevant to the content of this article.
Rights and permissions
Copyright information
© 2026 The Author(s), under exclusive license to Springer Nature Switzerland AG
About this paper
Cite this paper
Qu, C., Kwon, I., Thiyagarajan, K., Prasad, M., Braytee, A. (2026). PromptFusionSR: Multimodal Enhancement of Low-Resolution Images with Automatic Prompt-Guided Diffusion. In: Baratchi, M., Nijssen, S., van Rijn, J.N. (eds) Advances in Intelligent Data Analysis XXIV. IDA 2026. Lecture Notes in Computer Science, vol 16513. Springer, Cham. https://doi.org/10.1007/978-3-032-23833-7_7
Download citation
DOI: https://doi.org/10.1007/978-3-032-23833-7_7
Published:
Publisher Name: Springer, Cham
Print ISBN: 978-3-032-23832-0
Online ISBN: 978-3-032-23833-7
eBook Packages: Computer ScienceComputer Science (R0)Springer Nature Proceedings Computer Science


