A CTO's interactive primer

Embeddings & RAG

Meaning becomes geometry. Search becomes a prompt. Two ideas, fully interactive — scroll, drag, and run them yourself.

vectors · cosine similarity retrieval-augmented generation built for the AI Concierge
Part 1

Embeddings — turning words into coordinates

An embedding model is a function: text in, vector out. The magic is that semantic similarity becomes geometric proximity. Related phrases land near each other; unrelated ones drift apart. You never train this yourself — you call a model (text-embedding-3, Cohere, Voyage, bge). Same model for everything, always.

① The semantic map — click any point

Each chip below is a car-shop phrase embedded into a (toy) 2 D space. Click one to make it the anchor: lines fan out to every other point, colored by how similar they are. Notice the clusters form on their own — sedans, pickups, financing, colours.

very similar related distant

Toy 2-D for intuition. Real embeddings live in 768–3072 dimensions — same idea, more room.

② Cosine similarity, by hand — drag the arrows

Drag the tip of either arrow. Closeness is measured by the angle between vectors, not their length — that's cosine similarity.

Similarity is just the cosine of the angle θ between two vectors:

sim(a,b) = (a·b) / (‖a‖‖b‖) = cos θ

cosine similarity
θ = 90° 0.00
  • +1 → identical direction (same meaning)
  • 0 → unrelated (perpendicular)
  • −1 → opposite

That's the whole primitive. embed("Honda City 1.0 Turbo") → [0.013, −0.92, 0.41, …]. Text in, vector out, meaning preserved as direction. Everything in Part 2 is built on it.

Part 2

RAG — give the model your facts at query time

An LLM only knows its training data plus what's in the prompt. Your live inventory, pricing, and dealer policies aren't in the weights — and fine-tuning them in is slow, costly, and stale the moment a car sells. RAG sidesteps all of it: retrieve the relevant facts at query time and paste them into the prompt.

The pipeline, animated

Indexing happens once, offline (top row). Retrieval happens per message (bottom row). Hit Run a query to watch a question flow through.

Watch the chunk get embedded, searched, and answered.

Try it: ask the concierge — pick a question

A toy index of 8 listings. Your question gets embedded, the vector DB returns the nearest matches (the top-k), and only those are handed to the LLM to write a grounded answer.

vector DB · ranked by similarity
LLM answer · grounded in top-k only
Pick a question above…

The model is told: "Answer using only this context." No matching chunk → no answer (instead of a confident hallucination).

Part 3

What actually matters in production

🔪 Chunking is the hidden lever

Too big = noise drowns the signal; too small = no context. Bad chunks hurt more than model choice. Tune ~200–500 tokens first, before anything else.

🔎 Hybrid beats pure vector

Embeddings miss exact terms — part numbers, model codes, prices. Combine vector search with keyword/BM25. Most vector DBs do both now.

🧱 Filter before you rank

Apply price < 500000 AND transmission = 'auto' in SQL, then rank by similarity. Don't make the LLM do math a WHERE clause should.

🎯 It's a search engine, not a memory

RAG quality = retrieval quality. If the right chunk isn't in the top-k, the model can't use it — and may hallucinate instead. Eval retrieval separately from generation.

Why not just fine-tune? — a quick contrast

DimensionRAGFine-tuning
Fresh dataInstant — re-embed the changed rowStale until you retrain
Cost to updateOne embedding call (cents)Full training run
Source of truthYour DB, citableBaked into weights, opaque
Best forFacts, inventory, docsStyle, format, behavior
For the AI Concierge specifically: you're already on Postgres → pgvector is the obvious choice, no new infra. Embed each listing + each dealer FAQ, hybrid-search with metadata filters (price, brand, body type, location), feed the top matches to the LLM that talks on LINE / Messenger. Re-embed a listing only when it changes; delete its vector when the car sells.