UNETR: Transformers for 3D Medical Image Segmentation
UNETR: Трансформеры для 3D сегментации медицинских изображений
2022-01-01
SCID: 54.1/gj4brpff
Discuss with AI
3D medical image segmentationMulti Atlas Labeling Beyond The Cranial Vault (BTCV)U-shaped encoder-decoder with skip connectionsUNETRtransformer encoder
Figures from the paper
Abstract (AI)
Fully Convolutional Neural Networks (FCNNs) with contracting and expanding paths have shown prominence for the majority of medical image segmentation applications since the past decade. In FCNNs, the encoder plays an integral role by learning both global and local features and contextual representations which can be utilized for semantic output prediction by the decoder. Despite their success, the locality of convolutional layers in FCNNs, limits the capability of learning long-range spatial dependencies. Inspired by the recent success of transformers for Natural Language Processing (NLP) in long-range sequence learning, we reformulate the task of volumetric (3D) medical image segmentation as a sequence-to-sequence prediction problem. We introduce a novel architecture, dubbed as UNEt TRansformers (UNETR), that utilizes a transformer as the encoder to learn sequence representations of the input volume and effectively capture the global multi-scale information, while also following the successful "U-shaped" network design for the encoder and decoder. The transformer encoder is directly connected to a decoder via skip connections at different resolutions to compute the final semantic segmentation output. We have validated the performance of our method on the Multi Atlas Labeling Beyond The Cranial Vault (BTCV) dataset for multi-organ segmentation and the Medical Segmentation Decathlon (MSD) dataset for brain tumor and spleen segmentation tasks. Our benchmarks demonstrate new state-of-the-art performance on the BTCV leaderboard.
Key Findings
1
Proposed UNETR: a novel architecture that uses a transformer encoder within a U-shaped encoder-decoder with skip connections at multiple resolutions.
2
Reformulated 3D medical image segmentation as a sequence-to-sequence prediction problem using transformers.
3
Transformer encoder in UNETR effectively captures global multi-scale information and long-range spatial dependencies that convolutional layers struggle with.
4
UNETR achieved new state-of-the-art performance on the BTCV leaderboard.
5
Validated UNETR on BTCV (multi-organ) and MSD (brain tumor and spleen) segmentation tasks.
Research Object
Volumetric (3D) medical images for multi-organ and tumor segmentation (input volumes from BTCV and MSD datasets)
Research Subject
Using a transformer-based encoder within a U-shaped network (UNETR) to learn global multi-scale sequence representations and improve semantic segmentation performance by capturing long-range spatial dependencies via encoder–decoder skip connections
Publication Details
Publication Date
2022-01-01
Journal
Publisher
ISSN
Cited by
3162
Access Type
Author Information
Download PDF
Subscribe to digest
References available in scid.ai10
Exploiting Generative AI to Scale up Intelligent Tutoring Systems2023
AI-Assisted Pipeline for Dynamic Generation of Trustworthy Health Supplement Content at Scale2018
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows2021
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale2020
Non-local Neural Networks2018
Emerging Properties in Self-Supervised Vision Transformers2021
Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions2021
TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation2021
Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers2021
Self-Attention Generative Adversarial Networks2018