Lesson 3 · the part that's hard to copy

The Thai-optimized pipeline

Every competitor ships an upload box and an LLM. What they fail at is producing replies that are correct, grounded, and natural when a real Thai buyer types one messy, space-less, multi-intent line. This is how you survive that input.

newmm chunking Gemini / BGE-M3 embeddings hybrid + alias table rerank · confidence gate
The problem, in one message

Seven signals, no spaces, a BE year, slang, and a typo

A naive RAG-over-documents bot reads this as a vague blob and produces a half-hallucinated wall of text. A best-in-class engine parses all seven signals. Hit parse.

The conditions

Why Thai used-car chat breaks naive bots

Each of these silently degrades a generic English pipeline. Click a card to see what it breaks.

Design consequence: you can't bolt Thai onto an English pipeline. Thai handling must be a first-class normalization + hybrid-retrieval layer — and measured on a Thai golden set, because no embedding/reranker vendor publishes Thai benchmarks. Assume nothing; test everything.

Step 1 · query understanding

The Thai normalization layer

A small, fast step that runs before routing and turns messy Thai into clean structured signals. This is where most of the "on-point" feeling comes from. Pick a message and run it through.

→ structured turn_understanding the orchestrator routes on

    

Implement the fuzzy parts with a cheap LLM (Gemini Flash) constrained to JSON; do years, money, and aliases with deterministic regex/dictionaries — don't pay an LLM to parse 1.5 ล้าน.

Step 2 · chunking

Thai-aware chunking — never split mid-word

Thai has no spaces between words and no full stops between sentences. Fixed-size windows slice through the middle of a word and destroy meaning. Toggle to compare.

from pythainlp.tokenize import sent_tokenize, word_tokenize

# 1. split into Thai sentences (there are no full stops)
for sent in sent_tokenize(policy_text):
    # 2. word-safe segmentation before any length fallback
    tokens = word_tokenize(sent, engine="newmm")
    chunk  = pack(tokens, max_tokens=180, overlap=30)
    embed_and_store(chunk, metadata)   # Gemini Embedding 001 / BGE-M3
Step 3 · retrieval

Hybrid is non-negotiable — and so is the alias table

Dense vectors catch meaning; BM25 over newmm-segmented text catches exact terms; an alias table collapses transliteration variants so every spelling of a model name converges. Click a messy spelling.

a buyer types…
the three retrieval legs, fused by RRF
Dense — semantic embedding search (Gemini / BGE-M3). Catches paraphrase, misses exact transliteration.
BM25 — lexical over newmm tokens. The piece generic stacks omit → poor Thai keyword recall.
Alias boost — curated วีออส|วิออส|vios → VIOS table. Same table powers Class A inventory search.

BGE-M3 emits dense + sparse + ColBERT vectors in one pass — hybrid almost for free, and self-hostable (no per-call cost at volume).

Steps 4–5 · rerank & ground

Rerank, then ground — or refuse

Cross-encoder rerank

Score the query–chunk pair jointly over the top ~20 fused candidates; keep the best 3–5. The cheapest big quality win (+20–35%). Cohere Rerank, or self-host BGE reranker.

⚠️ Reranker Thai quality is also unbenchmarked — include rerank in your Thai eval (last section).

Grounding + confidence gate

The LLM answers only from the reranked chunks. If the top rerank score is below threshold, the bot asks one clarifying question or escalates — it does not guess. Filter chunks past expiry_date so dead promos never surface.

Step 6 · used-car guardrails

The claims a customer can sue over

Used-car selling carries liability the generic playbook ignores. The model may phrase grounded facts; it may never originate them. Pick a risky question.

Step 7 · evaluation

The persona / particle leak test

A ค่ะ-persona reply that contains a stray ครับ or unexpected English reads as broken — the exact mixed-locale leak your i18n rule targets. An automated checker runs post-generation. Click a candidate reply to scan it.

Allowlist for English tokens: model names (VIOS, FORTUNER), "PIN", "LINE". Everything else in a Thai reply is a leak. Run this per-locale check on every form and every generated reply — it's the same discipline as the per-locale validation-message test.

What the golden set measures

  • Retrieval hit-rate — right policy chunk in top-k, pre- and post-rerank
  • Routing accuracy — inventory (tools) vs RAG vs persona
  • Groundedness — no price / mileage / installment / availability untraceable to a tool
  • Persona consistency — the leak test above
  • Clarify-vs-guess on low-confidence queries
  • Resolution, not containment — did the buyer's real need get met (booked / answered / handed off)