Xiaohua Zhai | VitaVision
← Authors

Xiaohua Zhai

2 papers · 9 atlas pages

Papers (2)

Source

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Dosovitskiy, Beyer, Kolesnikov, Weissenborn et al. · ICLR 2021 (arXiv 2020) 2020

arXiv ↗
Source

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Tschannen, Gritsenko, Wang, Naeem et al. · arXiv preprint 2025

arXiv ↗

In the Atlas (9)

Co-authors (23)