arxiv preprint - CogVLM: Visual Expert for Pretrained Language Models

In this episode we discuss CogVLM: Visual Expert for Pretrained Language Models
by Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, Jie Tang. CogVLM is an open-source visual language foundation model that significantly improves the integration of vision and language by incorporating a trainable visual expert module within a pre-trained language model’s attention and feed-forward layers. Unlike other models, CogVLM deeply fuses visual and language features without losing any natural language processing capabilities. It delivers state-of-the-art results on several cross-modal benchmarks and is competitive on others, with resources and code accessible publicly.

arxiv preprint – CogVLM: Visual Expert for Pretrained Language Models