Agentic Research

Embedding

Also: 向量 · 嵌入 · 向量化 · embedding model

Turning text into a list of numbers so that "similar meaning" becomes "close in distance".

When you will meet it

Semantic search, RAG and similarity recommendations all rest on it. Without it you cannot diagnose why retrieval failed.

An analogy

Like placing every phrase on a huge map so similar meanings land nearby. "database connection failed" and "DB won't connect" barely overlap as words, but they are neighbours on the map.

Minimal example

vec = embed("資料庫連線逾時")
# → [0.021, -0.114, 0.873, ...]  幾百到幾千個浮點數

similarity(vec_a, vec_b)   # 餘弦相似度:1 = 幾乎同義,0 = 無關

No single dimension means anything on its own; what matters is the distance between two texts' vectors.

What people get wrong

  • Using a different embedding model at index time and query time. The coordinate systems differ, so results are effectively random. This is RAG's quietest and most fatal error.
  • Assuming high similarity means correct. Vectors find likeness, not truth.

Related terms

Next