Stop Randomly Tuning the Image Tower! Zhejiang University IJCAI Paper Reveals the Truth About VLM Asymmetry, Hitting the Brakes on CLIP Fine-Tuning
Published · Aug 17 · Mon Source · 雷峰网 (CN)

Stop Randomly Tuning the Image Tower! Zhejiang University IJCAI Paper Reveals the Truth About VLM Asymmetry, Hitting the Brakes on CLIP Fine-Tuning

Zhejiang University published a paper at IJCAI, pointing out that over-fine-tuning the image encoder of vision-language models (VLMs) like CLIP harms generalization ability, and proposed the A3B2 adaptive asymmetric adapter optimization strategy.

KeywordsStopRandomlyTuningImageTowerZhejiangUniversityIJCAI

Vision-language models often adopt efficient parameter fine-tuning schemes in few-shot transfer tasks. However, there is a prevailing mindset in the industry that increasing the fine-tuning parameters of the image encoder can improve the model's ability to capture visual details.

The latest paper published by the Zhejiang University research team at the IJCAI conference reveals the potential risks of this practice. The study found that overly aggressive fine-tuning on the image encoder actually severely damages the model's generalization ability when facing unknown real-world scenarios, leading to a decline in the original powerful perception ability of the pre-trained model.

To address this issue, the researchers proposed the A3B2 adaptive asymmetric adapter. This method aims to balance the model's performance on specific tasks with its generalization performance on out-of-distribution data through a more restrained fine-tuning strategy, providing new technical reference for the practical application of vision-language models.

This page provides an editorial summary based on publicly available information. It is not a republished article. Use the source link below for the original report.