Every competitor ships an upload box and an LLM. What they fail at is producing replies that are correct, grounded, and natural when a real Thai buyer types one messy, space-less, multi-intent line. This is how you survive that input.
A naive RAG-over-documents bot reads this as a vague blob and produces a half-hallucinated wall of text. A best-in-class engine parses all seven signals. Hit parse.
Each of these silently degrades a generic English pipeline. Click a card to see what it breaks.
Design consequence: you can't bolt Thai onto an English pipeline. Thai handling must be a first-class normalization + hybrid-retrieval layer — and measured on a Thai golden set, because no embedding/reranker vendor publishes Thai benchmarks. Assume nothing; test everything.
A small, fast step that runs before routing and turns messy Thai into clean structured signals. This is where most of the "on-point" feeling comes from. Pick a message and run it through.
turn_understanding the orchestrator routes onImplement the fuzzy parts with a cheap LLM (Gemini Flash) constrained to JSON; do years, money, and aliases with deterministic regex/dictionaries — don't pay an LLM to parse 1.5 ล้าน.
Thai has no spaces between words and no full stops between sentences. Fixed-size windows slice through the middle of a word and destroy meaning. Toggle to compare.
from pythainlp.tokenize import sent_tokenize, word_tokenize
# 1. split into Thai sentences (there are no full stops)
for sent in sent_tokenize(policy_text):
# 2. word-safe segmentation before any length fallback
tokens = word_tokenize(sent, engine="newmm")
chunk = pack(tokens, max_tokens=180, overlap=30)
embed_and_store(chunk, metadata) # Gemini Embedding 001 / BGE-M3
Dense vectors catch meaning; BM25 over newmm-segmented text catches exact terms; an alias table collapses transliteration variants so every spelling of a model name converges. Click a messy spelling.
BGE-M3 emits dense + sparse + ColBERT vectors in one pass — hybrid almost for free, and self-hostable (no per-call cost at volume).
Score the query–chunk pair jointly over the top ~20 fused candidates; keep the best 3–5. The cheapest big quality win (+20–35%). Cohere Rerank, or self-host BGE reranker.
⚠️ Reranker Thai quality is also unbenchmarked — include rerank in your Thai eval (last section).
The LLM answers only from the reranked chunks. If the top rerank score is below
threshold, the bot asks one clarifying question or escalates — it does not guess. Filter chunks
past expiry_date so dead promos never surface.
Used-car selling carries liability the generic playbook ignores. The model may phrase grounded facts; it may never originate them. Pick a risky question.
A ค่ะ-persona reply that contains a stray ครับ or unexpected English reads as broken — the exact mixed-locale leak your i18n rule targets. An automated checker runs post-generation. Click a candidate reply to scan it.
Allowlist for English tokens: model names (VIOS, FORTUNER), "PIN", "LINE". Everything else in a Thai reply is a leak. Run this per-locale check on every form and every generated reply — it's the same discipline as the per-locale validation-message test.