How It Works
VLM (vision-language model)
An AI model that takes in both images and text and answers in text — it can describe a photo, read a chart, or answer questions about what a camera sees. Most major chatbots that accept image uploads are built on one.
Origin · no single documented coiner
A descriptive label that grew with models such as OpenAI’s CLIP (2021) and DeepMind’s Flamingo (2022), which paired image understanding with language models; no single coiner.
