Grounding DINO-US-SAM: Text-Prompted Multiorgan Segmentation in Ultrasound With LoRA-Tuned Vision–Language Models
Grounding DINO‑US‑SAM: сегментация нескольких органов на УЗИ по текстовому запросу с VLM, тонко настроенными с помощью LoRA
2025-09-02
SCID: 54.1/my7uhsx5
Discuss with AI
Grounding DINOLoRA (low-rank adaptation)SAM2comparison to SOTA (UniverSeg, MedSAM, MedCLIP-segment anything model, BiomedParse, SAMUS)fine-tuning to ultrasound domaingeneralization to unseen datasetspublic ultrasound datasets (breast, thyroid, liver, prostate, kidney, paraspinal muscle)text-prompted multiorgan segmentationultrasound segmentationvision-language model
Figures from the paper
Abstract (AI)
Accurate and generalizable object segmentation in ultrasound imaging remains a significant challenge due to anatomical variability, diverse imaging protocols, and limited annotated data. In this study, we propose a prompt-driven vision-language model (VLM) that integrates grounding DINO with SAM2 to enable object segmentation across multiple ultrasound organs. A total of 18 public ultrasound datasets, encompassing the breast, thyroid, liver, prostate, kidney, and paraspinal muscle, were utilized. These datasets were divided into 15 for fine-tuning and validation of grounding DINO using low-rank adaptation (LoRA) to the ultrasound domain, and three were held out entirely for testing to evaluate performance in unseen distributions. Comprehensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) segmentation methods, including UniverSeg, MedSAM, MedCLIP-segment anything model (SAM), BiomedParse, and SAMUS on most seen datasets while maintaining strong performance on unseen datasets without additional fine-tuning. These results underscore the promise of VLMs in scalable and robust ultrasound image analysis, reducing dependence on large, organ-specific annotated datasets. We will publish our code on code.sonography.ai after acceptance.
Key Findings
1
A prompt-driven vision-language model combining grounding DINO with SAM2 enables object segmentation across multiple ultrasound organs.
2
Evaluated on 18 public ultrasound datasets covering breast, thyroid, liver, prostate, kidney, and paraspinal muscle, with three held-out unseen datasets for testing.
3
Low-rank adaptation (LoRA) fine-tuning of grounding DINO on 15 ultrasound datasets adapts the model to the ultrasound domain.
4
The method maintains strong performance on unseen datasets without additional fine-tuning, indicating good generalization across distributions.
5
The proposed approach outperforms SOTA segmentation methods (UniverSeg, MedSAM, MedCLIP-SAM, BiomedParse, SAMUS) on most seen datasets.
6
The study suggests VLMs can reduce dependence on large organ-specific annotated datasets for scalable, robust ultrasound image analysis.
Research Object
Text-prompted vision–language model system (Grounding DINO integrated with SAM2 and LoRA-tuned) for multiorgan ultrasound image segmentation
Research Subject
Accuracy and generalizability of multiorgan object segmentation in ultrasound images across seen and unseen datasets (performance compared to SOTA methods)
Publication Details
Publication Date
2025-09-02
Journal
Publisher
ISSN
Cited by
7
Open access PDF
Access Type
Author Information
Download PDF
Subscribe to digest