Text-Guided Medical Image Denoising with Vision-Language Fusion

A medical image denoising system that fuses vision and language: a UNet-based denoiser is conditioned on CLIP text embeddings derived from the VQA-RAD dataset, letting textual context guide the reconstruction of noisy scans. Compared to a vision-only baseline, the vision-language fusion improved reconstruction quality by 12.5% in PSNR and 4.6% in SSIM.