ع
Start Topics Teams Reference What's new Saved
AI & models

Multimodal

A model that handles more than just text — it can take in images, screenshots, PDFs, diagrams, and sometimes audio, all in the same conversation. 'Modal' refers to a mode of input; multimodal simply means it understands several at once, so you can mix a picture and a question freely.
Why it matters

It's why you can paste a screenshot of an error, a photo of a whiteboard, or a chart and ask about it directly — no need to retype everything into words first.

Part of How AI actually works
see also

← All terms