Sign in

UNIT: Unifying Image and Text Recognition in One Vision Encoder

By Yi Zhu and others
Currently, vision encoder models like Vision Transformers (ViTs) typically excel at image recognition tasks but cannot simultaneously support text recognition like human visual recognition. To address this limitation, we propose UNIT, a novel training framework aimed at UNifying Image and Text recognition within a single model. Starting with a vision... Show more
September 6, 2024
=
0
Loading PDF…
Loading full text...
Similar articles
Loading recommendations...
=
0
x1
UNIT: Unifying Image and Text Recognition in One Vision Encoder
Click on play to start listening