Lessons 1–3 were the theory. This is the working prototype: a real Thai chat engine that runs on mock dealer data, a self-improving data bank built by two AI agents, and an evaluation harness that proves it works — with the real numbers from the build.
It's easy to confuse them, so here's the clean split. Each is real and running today.
A real chat pipeline that turns one messy Thai message into a grounded, persona-correct reply. The logic is production-grade; only the database underneath is mock JSON.
A growing store of Thai vocabulary & test cases, fabricated by two AI agents (Gemini + Codex) and screened before anything touches the engine. Makes the bot understand more slang and typos over time.
A loop that generates customer chats, runs them through the real engine, and scores the replies — so we find bugs before customers do. Every fix below was found this way.
Key idea: real engine, mock data. We can trace and fix the hard parts (Thai understanding, retrieval, grounding) without building a real product database yet — that's the decision we're about to discuss.
Every customer message flows through these six stages (engine/pipeline.py:run_turn).
The expensive, fuzzy stages use an LLM; the parts that must be exact (prices, years, particles) are
deterministic code — we never pay an LLM to do arithmetic or count Thai politeness particles.
Messy Thai → structured signals. Word-segments with pythainlp newmm, converts Buddhist-era years
(ปี 63 → 2020), parses Thai money words (ห้าแสนเจ็ด → 579,000),
maps model aliases (วีออส/ดีแมค → TOYOTA_VIOS), detects colour/transmission, and classifies intent.
Decides which of 7 intents the message carries — it can be several at once
(inventory_search, financing, book_viewing, tradein,
policy_qa, negotiation, smalltalk).
Stock, prices, instalments are structured data — looked up with deterministic functions
(search_inventory, estimate_installment), never retrieved as prose.
Sold/reserved units are excluded; internal cost is stripped before it can ever reach a customer.
Warranty, financing steps, transfer fees are free text → retrieved with hybrid search:
lexical TF-IDF over newmm tokens + dense embeddings (gemini-embedding-2, 3072-dim), fused by
reciprocal-rank fusion, then reranked. Production target: pgvector halfvec(3072).
The LLM (gemini-3.5-flash) writes the answer grounded only in the facts gathered above,
in the tenant's voice — casual น้องกุ้ง (ends ค่ะ) vs formal premium consultant (ends ครับ).
With no API key it falls back to a deterministic Thai safe-template so the pipeline still runs.
A deterministic pass cleans what the LLM gets wrong: collapses doubled particles (นะคะค่ะ → นะคะ), fixes the rising question particle (ไหมค่ะ → ไหมคะ), and flags English leaks / banned phrases.
The single most important decision in the engine. Get this wrong and the bot invents prices.
| Class A — structured | Class B — prose | |
|---|---|---|
| What | stock, price, mileage, year, instalment | warranty, financing steps, transfer/trade-in policy |
| How it's answered | deterministic DB lookup (tool-calling) | hybrid retrieval (RAG) |
| Can the LLM invent it? | No — numbers come from tools only | Only from retrieved chunks |
| Why | a wrong price is a lawsuit; live data must be exact | prose tolerates paraphrase; meaning matters, not digits |
Mock data behind it: 2 tenants — บ้านกุ้งรถสวย (6 cars, 5 policy docs, casual ค่ะ) and Bangkok Auto Premium (5 cars, 4 policy docs, formal ครับ). Each tenant configures its own inventory, knowledge base, persona, and negotiation policy — that's the multi-tenant SaaS shape.
Thai buyers type endless slang, transliterations and typos. We can't hand-write them all, so two different AI agents fabricate them — and nothing is trusted automatically.
Generates realistic customer chat cases (messy messages + what a good reply
should contain). Called over REST from generator.py. Each round uses a fresh nonce so the
cache doesn't replay identical cases.
A file-handoff protocol: we write instructions to codex/CODEX_TASK.md,
run codex exec, it writes Thai vocabulary back to codex_output.json, and a validator
(ingest_codex.py) checks the format before queuing it. Two independent models cross-check each other.
Gemini + Codex propose aliases, intent-trigger fragments, and hard messages → all land in one
review/queue.json, deduped and stamped with where they came from.
auto_review.py annotates every item two ways: deterministic checks (does this alias map to a
real in-stock model? does the same word map two different ways = conflict?) and a Gemini "native-Thai
reviewer" stand-in that approves/rejects with a written reason. Status stays "pending" — this only
pre-sorts the pile.
Auto-approved items are staged to data/provisional_lexicon.json, which the engine loads behind a
flag (USE_PROVISIONAL_LEXICON). So the data bank improves the bot before a human is hired —
but stays clearly separate from human-blessed data.
A native-Thai reviewer confirms items in the "Review" tab; apply_reviews.py promotes them into
the golden set (human-blessed test cases) and data/lexicon.json (which overrides
provisional). The golden set is never used to teach the bot — that's the eval/train separation.
Human-approved so far: 0 — there's no native-Thai reviewer yet. That's exactly why the provisional tier exists: the bot keeps improving on AI-screened data, and the human just confirms a pre-sorted, pre-annotated list when they arrive (minutes, not hours).
A bug-finding loop. The trick is the engine being scored is the real one — so a green score means the actual logic works, not a mock of it.
Deterministic checks = hard guarantees: did it route to the right intent? does every number trace to a real fact (no invented prices)? did it avoid offering a sold car? correct politeness particle?
LLM-judge = the soft stuff: is the Thai natural and helpful? Rated 1–5. The judge only decides the verdict when the reply was really generated by Gemini — never on offline fallbacks, so the score can't be faked.
Engine bug. The code is actually wrong. Fix the engine. (All the fixes in the next section.)
Scorer false-positive. The reply was fine; the check was too strict. Fix the scorer. (e.g. the customer's own number counted as "invented".)
Bad generated expectation. The test itself was unreasonable. (e.g. labelling a car-seeking greeting as pure small-talk.) Not an engine problem.
| Round | บ้านกุ้ง (casual ค่ะ) | Bangkok Auto Premium (formal ครับ) |
|---|---|---|
| 16 | 8/10 | 10/10 |
| 17 | 8/10 → 10/10 after fixes | 9/10 → 10/10 after fixes |
| 18 | 10/10 | 8/10 |
| 19 | 8/10 | 9/10 |
| 20 | 8/10 | 10/10 |
| 21 | 10/10 | 10/10 |
| 22 | 8/10 | 9/10 |
| 23 | 9/10 | 9/10 |
| 24 | 10/10 | 8/10 |
| 25 | 10/12 | 12/12 |
The numbers don't only climb — and that's the point. Each fresh round throws new messy inputs, so dips surface genuinely new edge cases. By round 25 the remaining misses are almost all Bucket C (the evaluator being too strict), not engine defects.
These are real failures from real generated Thai messages, with the real root cause and fix. Click through them.
Input: ค่าโอนรถเท่าไหร่ครับ — a dead-simple question with an exact-match policy doc.
finance, tradein, warranty — the transfer doc with the ฿2,000–3,000 answer was never retrieved. Bot couldn't give the fee.bk-kb-transfer #1; reply states ฿2,000–3,000, free for cash buyers.Input: มีประกันหลังการขายให้ไหมคะ — there is a warranty doc.
Input: ฟอร์จูนเนอร์สีขาวปี 2018 … ลดได้เท่าไหร่ มีงบเก้าแสน — this exactly matches BK-004 (฿949,000).
Input: ดีแมค4ประตูออโต้ปี63มีไหมขอราคาสุดๆสด
The LLM was told "always end with ค่ะ" — and overdid it.
_fix_particles) guarantees correct particles regardless of what the LLM emits — and the style checker now flags both errors so they can't regress silently.Input: ดีแมกซ์สีเทาปี 2019 ยังอยู่ไหมคะ — that unit is reserved, not sold.
sold and reserved together as "gone". A reserved car can come back; telling the customer it's sold loses a live lead.Also fixed this way: multi-car comparison (x3 กับ ux ราคาเท่าไหร่ now answers both), typo-routing (จัดไฟแน้น → financing), small-talk no longer fires alongside real intents, and a scorer that wrongly flagged the customer's own requested discount as an "invented number".
These rules are what stop the system from quietly poisoning itself — important whichever path we pick next.
The golden set (human-blessed tests) is never used to teach the bot. If you train on your test set, your scores become fiction.
AI-screened data lives in a separate tier and is clearly labelled. A human verdict always overrides it. No "AI approved its own homework".
Domain rules (the 7 intents, status meanings) live in one place. The auto-reviewer is even guarded so it can't inject an intent the engine doesn't understand.
Prices, years, particles, groundedness — all handled by code, not LLM guesswork. The LLM only does what only an LLM can: natural language.
What's actually been de-risked: the hard, copy-resistant parts — Thai understanding, hybrid retrieval, grounded generation, persona control, and a measurement loop that catches regressions. The mock data underneath is the easy part to swap.
| Path | What it means | What carries over |
|---|---|---|
| Throwaway | the engine was only ever a learning/eval tool; rebuild the product clean | the knowledge (Class A/B design, the bug fixes as a spec, the golden set + lexicon) |
| Evolve-in-Python | harden this Python engine into the real service (add DB, channels, auth) | everything — code, data bank, eval harness become production assets |
| Rebuild-in-Laravel | reimplement the engine in your usual stack; keep Python only as an NLP sidecar | the data bank, eval harness, and validated logic as a reference implementation |
In all three, the data bank, the golden set, and every documented bug-fix are durable — they're knowledge about Thai car-buyer language, not throwaway code.