Meaning becomes geometry. Search becomes a prompt. Two ideas, fully interactive — scroll, drag, and run them yourself.
An embedding model is a function: text in, vector out. The magic is that
semantic similarity becomes geometric proximity. Related phrases land near each other;
unrelated ones drift apart. You never train this yourself — you call a model
(text-embedding-3, Cohere, Voyage, bge). Same model for everything, always.
Each chip below is a car-shop phrase embedded into a (toy) 2 D space. Click one to make it the anchor: lines fan out to every other point, colored by how similar they are. Notice the clusters form on their own — sedans, pickups, financing, colours.
Toy 2-D for intuition. Real embeddings live in 768–3072 dimensions — same idea, more room.
Drag the tip of either arrow. Closeness is measured by the angle between vectors, not their length — that's cosine similarity.
Similarity is just the cosine of the angle θ between two vectors:
sim(a,b) = (a·b) / (‖a‖‖b‖) = cos θ
That's the whole primitive. embed("Honda City 1.0 Turbo") → [0.013, −0.92, 0.41, …].
Text in, vector out, meaning preserved as direction. Everything in Part 2 is built on it.
An LLM only knows its training data plus what's in the prompt. Your live inventory, pricing, and dealer policies aren't in the weights — and fine-tuning them in is slow, costly, and stale the moment a car sells. RAG sidesteps all of it: retrieve the relevant facts at query time and paste them into the prompt.
Indexing happens once, offline (top row). Retrieval happens per message (bottom row). Hit Run a query to watch a question flow through.
A toy index of 8 listings. Your question gets embedded, the vector DB returns the nearest matches (the top-k), and only those are handed to the LLM to write a grounded answer.
The model is told: "Answer using only this context." No matching chunk → no answer (instead of a confident hallucination).
Too big = noise drowns the signal; too small = no context. Bad chunks hurt more than
model choice. Tune ~200–500 tokens first, before anything else.
Embeddings miss exact terms — part numbers, model codes, prices. Combine vector search with keyword/BM25. Most vector DBs do both now.
Apply price < 500000 AND transmission = 'auto' in SQL, then rank by
similarity. Don't make the LLM do math a WHERE clause should.
RAG quality = retrieval quality. If the right chunk isn't in the top-k, the model can't use it — and may hallucinate instead. Eval retrieval separately from generation.
| Dimension | RAG | Fine-tuning |
|---|---|---|
| Fresh data | Instant — re-embed the changed row | Stale until you retrain |
| Cost to update | One embedding call (cents) | Full training run |
| Source of truth | Your DB, citable | Baked into weights, opaque |
| Best for | Facts, inventory, docs | Style, format, behavior |