Text-Guided Modulation and Boundary-Aware Refinement for Vision-Text COVID-19 CT Lesion Segmentation
DOI:
https://doi.org/10.63313/hmt.9030Keywords:
Medical Image Segmentation, Vision-Language Learning, Lvit, Mosmeddata+, Boundary Refinement, Text-Guided ModulationAbstract
Accurate COVID-19 CT lesion segmentation is challenged by weak boundaries and variable lesion appearance. This paper investigates two enhancement strategies built upon a clean LViT/VT-MFLV baseline: Multi-scale Text-Guided Modulation (MTGM) and Boundary Context Auxiliary Learning (BCA). MTGM converts clinical text embeddings into residual channel gates, while BCA introduces boundary supervision and inference-active boundary-guided feature refinement. On MosMedData+, the clean baseline obtains 0.7318 Dice and 0.5996 IoU. The best evaluated variant, BCA-v2, reaches 0.747385 Dice and 0.613327 IoU. The experiments show that boundary-guided decoder refinement is more effective than training-only boundary supervision, while text-gate placement requires careful scale selection.
References
[1] JIAQILITech. (2023). VT-MFLV: Vision-Text Multimodal Feature Learning V Network for Medical Image Segmentation [Computer software]. https://github.com/JIAQILITech/VT-MFLV
[2] Li, Z., Li, Y., Li, Q., Wang, P., Guo, D., Lu, L., Jin, D., Hong, Q., & Song, M. (2022). LViT: Language meets Vision Transformer in medical image segmentation. arXiv:2206.14718.
[3] Morozov, S. P., Andreychenko, A. E., Pavlov, N. A., Vladzymyrskyy, A. V., Ledikhova, N. V., Gombolevskiy, V. A., Blokhin, I. A., Gelezhe, P. B., Gonchar, A. V., & Chernina, V. Y. (2020). MosMedData: Chest CT scans with COVID-19 related findings dataset. arXiv:2005.06465.
[4] Ma, J., Wang, Y., An, X., Ge, C., Yu, Z., Chen, J., Zhu, Q., Dong, G., He, J., He, Z., Cao, T., Zhu, Y., Nie, Z., & Yang, X. (2021). Toward data-efficient learning: A benchmark for COVID-19 CT lung and infection segmentation. Medical Physics, 48(3), 1197-1210.
[5] Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. Medical Image Computing and Computer-Assisted Intervention, 234-241.
[6] Dosovitskiy, A., Beyer, L., Kolesnikov, A., et al. (2021). An image is worth 16x16 words: Transformers for image recognition at scale. International Conference on Learning Representations.
[7] Oktay, O., Schlemper, J., Folgoc, L. L., et al. (2018). Attention U-Net: Learning where to look for the pancreas. arXiv:1804.03999.
[8] Sudre, C. H., Li, W., Vercauteren, T., Ourselin, S., & Cardoso, M. J. (2017). Generalised Dice overlap as a deep learning loss function for highly unbalanced segmentations. Deep Learning in Medical Image Analysis, 240-248.
[9] Milletari, F., Navab, N., & Ahmadi, S. A. (2016). V-Net: Fully convolutional neural networks for volumetric medical image segmentation. Fourth International Conference on 3D Vision, 565-571.
[10] Chen, J., Lu, Y., Yu, Q., et al. (2021). TransUNet: Transformers make strong encoders for medical image segmentation. arXiv:2102.04306.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 by author(s) and Erytis Publishing Limited

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.







