Multimodal Model
Multimodal Model
A type of model used in machine learning (see also machine learning model) that can process more than one type of input or output data, or "modality," at the same time. For example, a multimodal model can take both an image and text caption as input and then produce a unimodal output in the form of a score indicating how well the text caption describes the image. These models are highly versatile and useful in a variety of tasks, like image captioning and speech recognition.