Grounding DINO-US-SAM: Text-Prompted Multiorgan Segmentation in Ultrasound With LoRA-Tuned Vision–Language Models

Grounding DINO‑US‑SAM: сегментация нескольких органов на УЗИ по текстовому запросу с VLM, тонко настроенными с помощью LoRA
Hamza Rasaee, Taha Koleilat, Hassan Rivaz
2025-09-02

Grounding DINOLoRA (low-rank adaptation)SAM2comparison to SOTA (UniverSeg, MedSAM, MedCLIP-segment anything model, BiomedParse, SAMUS)fine-tuning to ultrasound domaingeneralization to unseen datasetspublic ultrasound datasets (breast, thyroid, liver, prostate, kidney, paraspinal muscle)text-prompted multiorgan segmentationultrasound segmentationvision-language model
Accurate and generalizable object segmentation in ultrasound imaging remains a significant challenge due to anatomical variability, diverse imaging protocols, and limited annotated data. In this study, we propose a prompt-driven vision-language model (VLM) that integrates grounding DINO with SAM2 to enable object segmentation across multiple ultrasound organs. A total of 18 public ultrasound datasets, encompassing the breast, thyroid, liver, prostate, kidney, and paraspinal muscle, were utilized. These datasets were divided into 15 for fine-tuning and validation of grounding DINO using low-rank adaptation (LoRA) to the ultrasound domain, and three were held out entirely for testing to evaluate performance in unseen distributions. Comprehensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) segmentation methods, including UniverSeg, MedSAM, MedCLIP-segment anything model (SAM), BiomedParse, and SAMUS on most seen datasets while maintaining strong performance on unseen datasets without additional fine-tuning. These results underscore the promise of VLMs in scalable and robust ultrasound image analysis, reducing dependence on large, organ-specific annotated datasets. We will publish our code on code.sonography.ai after acceptance.
1
A prompt-driven vision-language model combining grounding DINO with SAM2 enables object segmentation across multiple ultrasound organs.
2
Evaluated on 18 public ultrasound datasets covering breast, thyroid, liver, prostate, kidney, and paraspinal muscle, with three held-out unseen datasets for testing.
3
Low-rank adaptation (LoRA) fine-tuning of grounding DINO on 15 ultrasound datasets adapts the model to the ultrasound domain.
4
The method maintains strong performance on unseen datasets without additional fine-tuning, indicating good generalization across distributions.
5
The proposed approach outperforms SOTA segmentation methods (UniverSeg, MedSAM, MedCLIP-SAM, BiomedParse, SAMUS) on most seen datasets.
6
The study suggests VLMs can reduce dependence on large organ-specific annotated datasets for scalable, robust ultrasound image analysis.

Text-prompted vision–language model system (Grounding DINO integrated with SAM2 and LoRA-tuned) for multiorgan ultrasound image segmentation

Accuracy and generalizability of multiorgan object segmentation in ultrasound images across seen and unseen datasets (performance compared to SOTA methods)

Publication Details
Publication Date
2025-09-02
Journal
Publisher
ISSN
Cited by
7
Access Type
Author Information
Authors
Hamza Rasaee
Taha Koleilat
Hassan Rivaz
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%