TRANSFORMERS & LLMs • 61 / 397
Follow text from characters to tokens, attention, transformers and generation.
Multimodal Model
A multimodal model can work with more than one type of input or output, such as text plus images or audio.
Think of it like
Think of Multimodal Model as part of a very fast reader that repeatedly decides what information matters and what should come next.
Real life
Autocomplete, translation and chat assistants depend on concepts such as Multimodal Model.
SRE lens
An SRE assistant might read a dashboard screenshot as well as logs.
Remember thisA multimodal model can work with more than one type of input or output, such as text plus images or audio.