Agentic Research

OCR (Optical Character Recognition)

Also: 光學字元識別 · 文字識別 · 掃描檔轉文字 · OCR 辨識 · PDF 轉文字

Recognising the "picture-of-text" in images, scans and PDFs into real text a computer can process; the first step for any paper document entering a data pipeline.

When you will meet it

If your knowledge base includes scanned annual reports, contracts, invoices or old memos, they must go through OCR first, or the model just sees a page of unreadable image. It is also a common culprit behind "a big chunk of data is mysteriously missing": bad OCR cannot be fixed by stronger RAG downstream.

An analogy

Like hiring a typist to key a stack of paper documents into computer files. Type it accurately and everything downstream is easy; type it wrong (0 read as O, a table scrambled) and the errors ride all the way into retrieval and answers — and nobody notices the mistake came from this step.

Minimal example

一份掃描年報的兩頁:

第 1 頁(印刷體、清晰):
  OCR →「營業額 12,345 千元」        ✓ 準確

第 2 頁(表格+手寫批註+模糊):
  OCR →「營業額 l2,345 干元」        ✗ 錯字連篇
  表格欄位對錯行、手寫幾乎全滅

→ 兩頁都「有輸出」、看起來都成功,但第 2 頁的資料已經壞了

The most dangerous part: OCR failure usually raises no error — it still outputs a string of text, just a wrong one. Downstream embedding, retrieval and answers all build on those typos, and you cannot see from the final answer that OCR was the source. So paper pipelines must spot-check OCR output, especially tables and numbers.

What people get wrong

  • Assuming a PDF that "has text" can be read directly. Many PDFs are scanned images: they look like documents but every page is a picture. Feed one to a model without OCR and it just says it cannot read the content.
  • Applying plain OCR to tables, multi-column layouts or handwriting. Reading order and column mapping go wrong most easily here, and one wrong digit sinks an entire financial analysis. Such documents need OCR that handles tables/layout, plus human review of key figures.

Related terms

Next