Agentic Research

Multimodal

Also: 多模態 · 多模態模型 · 視覺模型 · vision model · VLM · 圖片理解

A model that handles more than text (images, audio, video): non-text input is converted into representations the model can read, sharing one context with the text.

When you will meet it

You meet it the first time you want to "send a screenshot to the AI", or scan a model list for which entries accept images. Without it you conflate two things: assuming every chat model can see (you find out via an error — or worse, via it pretending to see), and assuming a model that reads images can also generate them. Understanding and generation are usually separate machinery.

An analogy

Natively multimodal is like someone born able to see and hear: all senses enter one brain and are thought about together. A bolted-on setup is like a blind person relying on a companion's narration: fine for everyday use when the narrator is good, but spatial relations and small text on a chart are exactly what narration loses.

Minimal example

// 一張圖+一句話的請求(OpenAI 風格 API,示意)
{
  "role": "user",
  "content": [
    { "type": "text",      "text": "這張報表哪裡異常?" },
    { "type": "image_url", "image_url": { "url": "data:image/png;base64,..." } }
  ]
}
// 圖像會被編碼成一串「視覺 token」,
// 和文字一起佔用上下文視窗的額度,也一起計費。

Look at the last two lines: the image is not "attached alongside" — it is converted into tokens and merged into the same context. A single high-resolution image can eat a sizeable slice of the budget, which is why image requests are pricier and hit context limits faster.

What people get wrong

  • Assuming "multimodal" is one master switch. Each modality and each direction (input/output) is an independent capability: reading images is not generating them, hearing speech is not speaking. Check the spec line by line instead of trusting the label.
  • Assuming bolting a vision encoder onto an LLM equals native multimodality. In bolted-on pipelines, visual information passes through an adapter before the language model sees it, and fine-grained cross-modal reasoning — exact small text, complex spatial relations, continuous action in video — is typically weaker.

Related terms

Next