TRANSFORMERS & LLMs • 61 / 397
Follow text from characters to tokens, attention, transformers and generation.

Multimodal Model

A multimodal model can work with more than one type of input or output, such as text plus images or audio.

Think of it like

Think of Multimodal Model as part of a very fast reader that repeatedly decides what information matters and what should come next.

Real life

Autocomplete, translation and chat assistants depend on concepts such as Multimodal Model.

SRE lens

An SRE assistant might read a dashboard screenshot as well as logs.

Remember thisA multimodal model can work with more than one type of input or output, such as text plus images or audio.
AIForSREJump to a concept
Search is optional. The main journey is simply ↓ Next.