Alpha-CLIP: A CLIP Model Focusing on Wherever you Want
Alpha-CLIP: модель CLIP, фокусирующаяся там, где нужно
2024-06-16
SCID: 54.1/mfuzyq6q
Discuss with AI
Alpha-CLIPRGBA region-text pairsalpha channel augmentationcontrolled image emphasisregion-focused CLIP
Figures from the paper
Abstract (AI)
Contrastive Language-Image Pre-training (CLIP) plays an essential role in extracting valuable content information from images across diverse tasks. It aligns textual and visual modalities to comprehend the entire image, including all the details, even those irrelevant to specific tasks. How-ever, for a finer understanding and controlled editing of images, it becomes crucial to focus on specific regions of interest, which can be indicated as points, masks, or boxes by humans or perception models. To fulfill the require we introduce Alpha-CLIp, an enhanced version of CLIP with an auxiliary alpha channel to suggest attentive regions and fine-tuned with constructed millions of RGBA region-text pairs. Alpha-CLIP not only preserves the visual recognition ability of CLIP but also enables precise control over the emphasis of image contents. It demonstrates effectiveness in various tasks, including but not limited to open-world recognition, multimodal large language models, and conditional 2D / 3D generation. It has a strong potential to serve as a versatile tool for image-related tasks. Our project is with codes and models available is linked to https://aleafy.github.io/alpha-clip/.
Key Findings
1
Alpha-CLIP adds an auxiliary alpha channel to CLIP to indicate attentive regions (points, masks, or boxes) for focused image understanding.
2
Alpha-CLIP improves applicability in tasks requiring region focus, demonstrated effective for open-world recognition, multimodal LLMs, and conditional 2D/3D generation.
3
Alpha-CLIP is fine-tuned on constructed millions of RGBA region-text pairs to learn region-conditioned visual-language alignment.
4
Alpha-CLIP is released with code and models, enabling use as a versatile tool for image-related tasks.
5
Alpha-CLIP preserves CLIP’s original visual recognition ability while enabling precise control over emphasis of image contents.
Research Object
Alpha-CLIP model (CLIP enhanced with an auxiliary alpha channel trained on RGBA region–text pairs)
Research Subject
Controlled attention to specific image regions (points, masks, or boxes) enabling precise emphasis of image contents while preserving CLIP's visual recognition capabilities
Publication Details
Publication Date
2024-06-16
Journal
Publisher
ISSN
Cited by
90
Access Type
Author Information
Download PDF
Subscribe to digest