Alpha-CLIP: A CLIP Model Focusing on Wherever you Want

Alpha-CLIP: модель CLIP, фокусирующаяся там, где нужно
Dahua Lin, Pan Zhang, Yuhang Zang, Tong Wu, Yuanjun Xiong, Jiaqi Wang, Shu Kong, Fang Ye, Zeyi Sun
2024-06-16

Alpha-CLIPRGBA region-text pairsalpha channel augmentationcontrolled image emphasisregion-focused CLIP
Contrastive Language-Image Pre-training (CLIP) plays an essential role in extracting valuable content information from images across diverse tasks. It aligns textual and visual modalities to comprehend the entire image, including all the details, even those irrelevant to specific tasks. How-ever, for a finer understanding and controlled editing of images, it becomes crucial to focus on specific regions of interest, which can be indicated as points, masks, or boxes by humans or perception models. To fulfill the require we introduce Alpha-CLIp, an enhanced version of CLIP with an auxiliary alpha channel to suggest attentive regions and fine-tuned with constructed millions of RGBA region-text pairs. Alpha-CLIP not only preserves the visual recognition ability of CLIP but also enables precise control over the emphasis of image contents. It demonstrates effectiveness in various tasks, including but not limited to open-world recognition, multimodal large language models, and conditional 2D / 3D generation. It has a strong potential to serve as a versatile tool for image-related tasks. Our project is with codes and models available is linked to https://aleafy.github.io/alpha-clip/.
1
Alpha-CLIP adds an auxiliary alpha channel to CLIP to indicate attentive regions (points, masks, or boxes) for focused image understanding.
2
Alpha-CLIP improves applicability in tasks requiring region focus, demonstrated effective for open-world recognition, multimodal LLMs, and conditional 2D/3D generation.
3
Alpha-CLIP is fine-tuned on constructed millions of RGBA region-text pairs to learn region-conditioned visual-language alignment.
4
Alpha-CLIP is released with code and models, enabling use as a versatile tool for image-related tasks.
5
Alpha-CLIP preserves CLIP’s original visual recognition ability while enabling precise control over emphasis of image contents.

Alpha-CLIP model (CLIP enhanced with an auxiliary alpha channel trained on RGBA region–text pairs)

Controlled attention to specific image regions (points, masks, or boxes) enabling precise emphasis of image contents while preserving CLIP's visual recognition capabilities

Publication Details
Publication Date
2024-06-16
Journal
Publisher
ISSN
Cited by
90
Access Type
Author Information
Authors
Dahua Lin
Pan Zhang
Yuhang Zang
Tong Wu
Yuanjun Xiong
Jiaqi Wang
Shu Kong
Fang Ye
Zeyi Sun
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat
Make a presentation
100%