Data pipeline
Also: 資料管線 · 資料流程 · data pipeline · ETL · 資料處理流程
A fixed chain of steps that carries data from source to usable state: ingest → clean → transform → store → serve. Almost every AI application sits on such a pipeline.
When you will meet it
However strong the model, if the data fed in is dirty, misformatted or from unreliable sources, the output is unreliable — garbage in, garbage out. Understanding each stage lets you tell, when the AI is wrong, whether the fault is in ingest, clean, transform or retrieval, instead of vaguely blaming the model.
An analogy
Like a water treatment plant: intake (ingest) → filter and settle (clean) → disinfect (transform) → hold in a tower (store) → out of your tap (serve). If any stage fails, you turn the tap and see dirty water — but you only see the tap, not which stage broke.
Minimal example
一條看似正常的管線:
攝: 從 API 抓「發票」→ 存成 {date, amount, vendor}
清洗: 去掉重複、補齊缺的欄位
轉換: amount 統一成整數(分)
儲存: 寫進資料庫
供查詢:RAG/報表來讀
AI 常在這裡悄悄弄壞它:
· 上游某天把 amount 從「分」改成「元」→ schema 漂移,
你的轉換照舊除以 100,金額靜默縮小 100 倍,沒人報錯
· 讓 LLM 抽取欄位卻沒校驗 → 它偶爾把 vendor 填成日期Notice both failures are not crashes with errors but "silently producing wrong data". The most dangerous pipeline failure is not breaking (you would notice) but running on while carrying wrong data all the way through. So every stage needs validation: did the input change, are the fields right, are the values in a sane range.
What people get wrong
- Pouring all effort into the model and ignoring the data pipeline. In practice most problems are in the data: dirty sources, silent schema drift, no validation. A stronger model cannot save bad data fed into it.
- Assuming a pipeline that worked once works forever. Upstream formats change, APIs get revised, sources come and go. A pipeline without monitoring and validation will silently start producing wrong data one day, and you will notice much later.