close
Skip to main content

PromptFusionSR: Multimodal Enhancement of Low-Resolution Images with Automatic Prompt-Guided Diffusion

  • Conference paper
  • First Online:
BERJAYA Advances in Intelligent Data Analysis XXIV (IDA 2026)

Part of the book series: Lecture Notes in Computer Science ((LNCS,volume 16513))

Included in the following conference series:

  • 418 Accesses

Abstract

Image super-resolution is a technique aimed at enhancing the resolution of images and is widely utilized in fields such as remote sensing and medical imaging. We view image super-resolution as a reconstruction problem for inferring missing structural information from incomplete observations. This perspective highlights the need for methods that do not merely generate pixel-level corrections, but instead analyze and infer latent geometric and semantic patterns embedded in multimodal data. In this work, we present PromptFusionSR (PF-SR), a novel framework for text-assisted image super-resolution. In contrast to conventional approaches that rely solely on visual features, our method leverages the BLIP3 vision language model to extract contextual prompts-short textual descriptions-directly from low-resolution images. A periodic semantic–structural fusion mechanism enables the model to integrate complementary information channels, enhancing its ability to interpret high-frequency patterns under severe degradation. These textual and information cues guide a pretrained stable diffusion model for 4\(\times \) upscaling, yielding high-resolution outputs with improved fidelity and richer visual textures. This work further explores how complementary information channels can be systematically integrated during image reconstruction to improve robustness and perceptual fidelity. On our benchmark datasets, our method surpasses the current state-of-the-art methods across most evaluation metrics, with robust gains in perceptual quality metrics.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Subscribe and save

Springer+
from $39.99 /Month
  • Starting from 10 chapters or articles per month
  • Access and download chapters and articles from more than 300k books and 2,500 journals
  • Cancel anytime
View plans

Buy Now

Chapter
USD 29.95
Price excludes VAT (USA)
  • Available as PDF
  • Read on any device
  • Instant download
  • Own it forever
eBook
USD 69.99
Price excludes VAT (USA)
  • Available as EPUB and PDF
  • Read on any device
  • Instant download
  • Own it forever
Softcover Book
USD 89.99
Price excludes VAT (USA)
  • Compact, lightweight edition
  • Free shipping worldwide - view details

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Similar content being viewed by others

References

  1. Freeman, W.T., Pasztor, E.C., Carmichael, O.T.: Learning low-level vision. Int. J. Comput. Vision 40, 25–47 (2000)

    Article  Google Scholar 

  2. Park, S.C., Park, M.K., Kang, M.G.: Super-resolution image reconstruction: a technical overview. IEEE Signal Process. Mag. 20(3), 21–36 (2003)

    Article  Google Scholar 

  3. Mena, F., Arenas, D., Dengel, A.: Missing data as augmentation in the earth observation domain: a multi-view learning approach. Neurocomputing 638, 130175 (2025)

    Article  Google Scholar 

  4. Vexler, J., Vieten, B., Nelke, M., Kramer, S.: Integrating inverse and forward modeling for sparse temporal data from sensor networks. In: International Symposium on Intelligent Data Analysis, pp. 318–329. Springer (2025)

    Google Scholar 

  5. Liu, F., Hu, X., Bu, C., Yu, K.: Fuzzy Bayesian knowledge tracing. IEEE Trans. Fuzzy Syst. 30(7), 2412–2425 (2021)

    Article  Google Scholar 

  6. Wang, Z., Chen, J., Hoi, S.C.: Deep learning for image super-resolution: a survey. IEEE Trans. Pattern Anal. Mach. Intell. 43(10), 3365–3387 (2020)

    Article  Google Scholar 

  7. Persson, D., Wahlberg, W., Vettoruzzo, A., Nowaczyk, S.: Bridging spatial and temporal contexts: sparse transfer learning. In: Krempl, G., Puolamäki, K., Miliou, I. (eds.) Advances in Intelligent Data Analysis XXIII, pp. 330–342. Springer Nature Switzerland, Cham (2025)

    Chapter  Google Scholar 

  8. Saharia, C., Ho, J., Chan, W., Salimans, T., Fleet, D.J., Norouzi, M.: Image super-resolution via iterative refinement. IEEE Trans. Pattern Anal. Mach. Intell. 45(4), 4713–4726 (2022)

    Google Scholar 

  9. Wang, J., Yue, Z., Zhou, S., Chan, K.C., Loy, C.C.: Exploiting diffusion prior for real-world image super-resolution. Int. J. Comput. Vision 132(12), 5929–5949 (2024)

    Article  Google Scholar 

  10. Xue, L., et al.: xGen-MM (BLIP-3): a family of open large multimodal models. arXiv preprint arXiv:2408.08872 (2024)

  11. Zhao, Y., Prasad, M., Braytee, A.: DualPrompt-MedCap: a dual-prompt enhanced approach for medical image captioning. In: International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 205–215. Springer (2025)

    Google Scholar 

  12. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., Ommer, B.: High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 10684–10695 (2022)

    Google Scholar 

  13. Lim, B., Son, S., Kim, H., Nah, S., Mu Lee, K.: Enhanced deep residual networks for single image super-resolution. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, pp. 136–144 (2017)

    Google Scholar 

  14. Zhang, Y., Li, K., Li, K., Wang, L., Zhong, B., Fu, Y.: Image super-resolution using very deep residual channel attention networks. In: Proceedings of the European Conference on Computer Vision (ECCV), pp. 286–301 (2018)

    Google Scholar 

  15. Chen, C., et al.: Real-world blind super-resolution via feature matching with implicit high-resolution priors. In: Proceedings of the 30th ACM International Conference on Multimedia, pp. 1329–1338 (2022)

    Google Scholar 

  16. Lin, X., et al.: DiffBIR: toward blind image restoration with generative diffusion prior. In: European Conference on Computer Vision, pp. 430–448. Springer (2024)

    Google Scholar 

  17. Yang, T., Wu, R., Ren, P., Xie, X., Zhang, L.: Pixel-aware stable diffusion for realistic image super-resolution and personalized stylization. In: European Conference on Computer Vision, pp. 74–91. Springer (2024)

    Google Scholar 

  18. Wu, R., Yang, T., Sun, L., Zhang, Z., Li, S., Zhang, L.: SeeSR: towards semantics-aware real-world image super-resolution. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 25456–25467 (2024)

    Google Scholar 

  19. Ho, J., Jain, A., Abbeel, P.: Denoising diffusion probabilistic models. In: Advances in Neural Information Processing Systems, vol. 33, pp. 6840–6851 (2020)

    Google Scholar 

  20. Agustsson, E., Timofte, R.: NTIRE 2017 challenge on single image super-resolution: dataset and study. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, July 2017

    Google Scholar 

  21. Wei, P., et al.: Component divide-and-conquer for real-world image super-resolution. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, 23–28 August 2020, Proceedings, Part VIII 16, pp. 101–117. Springer (2020)

    Google Scholar 

  22. Cai, J., Zeng, H., Yong, H., Cao, Z., Zhang, L.: Toward real-world single image super-resolution: a new benchmark and a new model. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3086–3095 (2019)

    Google Scholar 

  23. Hore, A., Ziou, D.: Image quality metrics: PSNR vs. SSIM. In: 2010 20th International Conference on Pattern Recognition, pp. 2366–2369. IEEE (2010)

    Google Scholar 

  24. Zhang, R., Isola, P., Efros, A.A., Shechtman, E., Wang, O.: The unreasonable effectiveness of deep features as a perceptual metric. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 586–595 (2018)

    Google Scholar 

  25. Ding, K., Ma, K., Wang, S., Simoncelli, E.P.: Image quality assessment: unifying structure and texture similarity. IEEE Trans. Pattern Anal. Mach. Intell., 1 (2020). https://doi.org/10.1109/tpami.2020.3045810

  26. Mittal, A., Soundararajan, R., Bovik, A.C.: Making a “completely blind” image quality analyzer. IEEE Sig. Process. Lett. 20(3), 209–212 (2012)

    Google Scholar 

Download references

Author information

Authors and Affiliations

Authors

Corresponding author

Correspondence to Ali Braytee.

Editor information

Editors and Affiliations

Ethics declarations

Disclosure of Interests

The authors have no competing interests to declare that are relevant to the content of this article.

Rights and permissions

Reprints and permissions

Copyright information

© 2026 The Author(s), under exclusive license to Springer Nature Switzerland AG

About this paper

Check for updates. Verify currency and authenticity via CrossMark

Cite this paper

Qu, C., Kwon, I., Thiyagarajan, K., Prasad, M., Braytee, A. (2026). PromptFusionSR: Multimodal Enhancement of Low-Resolution Images with Automatic Prompt-Guided Diffusion. In: Baratchi, M., Nijssen, S., van Rijn, J.N. (eds) Advances in Intelligent Data Analysis XXIV. IDA 2026. Lecture Notes in Computer Science, vol 16513. Springer, Cham. https://doi.org/10.1007/978-3-032-23833-7_7

Download citation

Keywords

Publish with us

Policies and ethics