Lesson 4 · what we actually built

The engine & the data bank

Lessons 1–3 were the theory. This is the working prototype: a real Thai chat engine that runs on mock dealer data, a self-improving data bank built by two AI agents, and an evaluation harness that proves it works — with the real numbers from the build.

Python engine · no framework Gemini API + Codex CLI pythainlp newmm HITL-gated lexicon 645-item review queue
The big picture

Three things, working together

It's easy to confuse them, so here's the clean split. Each is real and running today.

① The engine

A real chat pipeline that turns one messy Thai message into a grounded, persona-correct reply. The logic is production-grade; only the database underneath is mock JSON.

normalize → route → tools → RAG → reply

② The data bank

A growing store of Thai vocabulary & test cases, fabricated by two AI agents (Gemini + Codex) and screened before anything touches the engine. Makes the bot understand more slang and typos over time.

generate → pre-screen → human-gate → apply

③ The eval harness

A loop that generates customer chats, runs them through the real engine, and scores the replies — so we find bugs before customers do. Every fix below was found this way.

deterministic checks + LLM-judge

Key idea: real engine, mock data. We can trace and fix the hard parts (Thai understanding, retrieval, grounding) without building a real product database yet — that's the decision we're about to discuss.

① the engine · one turn

Six stages from message to reply

Every customer message flows through these six stages (engine/pipeline.py:run_turn). The expensive, fuzzy stages use an LLM; the parts that must be exact (prices, years, particles) are deterministic code — we never pay an LLM to do arithmetic or count Thai politeness particles.

1

Normalize · turn understanding

Messy Thai → structured signals. Word-segments with pythainlp newmm, converts Buddhist-era years (ปี 63 → 2020), parses Thai money words (ห้าแสนเจ็ด → 579,000), maps model aliases (วีออส/ดีแมค → TOYOTA_VIOS), detects colour/transmission, and classifies intent.

tokenizerBE→CE yearmoney parseralias table
↓
2

Route · intent classification

Decides which of 7 intents the message carries — it can be several at once (inventory_search, financing, book_viewing, tradein, policy_qa, negotiation, smalltalk).

multi-intentsubstring triggers
↓
3

Class A tools · live facts, never RAG

Stock, prices, instalments are structured data — looked up with deterministic functions (search_inventory, estimate_installment), never retrieved as prose. Sold/reserved units are excluded; internal cost is stripped before it can ever reach a customer.

tool-callingno hallucinated prices
↓
4

Class B RAG · policy prose

Warranty, financing steps, transfer fees are free text → retrieved with hybrid search: lexical TF-IDF over newmm tokens + dense embeddings (gemini-embedding-2, 3072-dim), fused by reciprocal-rank fusion, then reranked. Production target: pgvector halfvec(3072).

lexical + denseRRF fusionIDF² rerank
↓
5

Persona compose · the reply

The LLM (gemini-3.5-flash) writes the answer grounded only in the facts gathered above, in the tenant's voice — casual น้องกุ้ง (ends ค่ะ) vs formal premium consultant (ends ครับ). With no API key it falls back to a deterministic Thai safe-template so the pipeline still runs.

grounded generationper-tenant persona
↓
6

Style & particle fix · the polish

A deterministic pass cleans what the LLM gets wrong: collapses doubled particles (นะคะค่ะ → นะคะ), fixes the rising question particle (ไหมค่ะ → ไหมคะ), and flags English leaks / banned phrases.

guaranteed-correct Thai
the core design rule

Class A vs Class B — the split that prevents lies

The single most important decision in the engine. Get this wrong and the bot invents prices.

 Class A — structuredClass B — prose
Whatstock, price, mileage, year, instalmentwarranty, financing steps, transfer/trade-in policy
How it's answereddeterministic DB lookup (tool-calling)hybrid retrieval (RAG)
Can the LLM invent it?No — numbers come from tools onlyOnly from retrieved chunks
Whya wrong price is a lawsuit; live data must be exactprose tolerates paraphrase; meaning matters, not digits

Mock data behind it: 2 tenants — บ้านกุ้งรถสวย (6 cars, 5 policy docs, casual ค่ะ) and Bangkok Auto Premium (5 cars, 4 policy docs, formal ครับ). Each tenant configures its own inventory, knowledge base, persona, and negotiation policy — that's the multi-tenant SaaS shape.

② the multi-agent data bank

Two AI agents build the vocabulary; humans get the final say

Thai buyers type endless slang, transliterations and typos. We can't hand-write them all, so two different AI agents fabricate them — and nothing is trusted automatically.

🅰 Gemini · via API

Generates realistic customer chat cases (messy messages + what a good reply should contain). Called over REST from generator.py. Each round uses a fresh nonce so the cache doesn't replay identical cases.

🅱 Codex · via command line

A file-handoff protocol: we write instructions to codex/CODEX_TASK.md, run codex exec, it writes Thai vocabulary back to codex_output.json, and a validator (ingest_codex.py) checks the format before queuing it. Two independent models cross-check each other.

How a Thai word travels from "AI guessed it" to "engine uses it"

1

Collect

Gemini + Codex propose aliases, intent-trigger fragments, and hard messages → all land in one review/queue.json, deduped and stamped with where they came from.

↓
2

Pre-screen · no human needed

auto_review.py annotates every item two ways: deterministic checks (does this alias map to a real in-stock model? does the same word map two different ways = conflict?) and a Gemini "native-Thai reviewer" stand-in that approves/rejects with a written reason. Status stays "pending" — this only pre-sorts the pile.

↓
3

Provisional tier · lifts the engine now

Auto-approved items are staged to data/provisional_lexicon.json, which the engine loads behind a flag (USE_PROVISIONAL_LEXICON). So the data bank improves the bot before a human is hired — but stays clearly separate from human-blessed data.

↓
4

Human gate · the real reviewer, later

A native-Thai reviewer confirms items in the "Review" tab; apply_reviews.py promotes them into the golden set (human-blessed test cases) and data/lexicon.json (which overrides provisional). The golden set is never used to teach the bot — that's the eval/train separation.

The data bank, in real numbers

645
items in review queue (311 cases + 334 lexicon terms)
586
pre-screened in the last auto-review pass
561 / 25
auto-suggested approve / reject · 0 conflicts
491 / 95
screened by Gemini stand-in / by deterministic rules
144
model aliases now live in the engine (e.g. ดีแมค)
167
intent-trigger fragments live across 7 intents

Human-approved so far: 0 — there's no native-Thai reviewer yet. That's exactly why the provisional tier exists: the bot keeps improving on AI-screened data, and the human just confirms a pre-sorted, pre-annotated list when they arrive (minutes, not hours).

③ the eval harness

How we know it works (and find what doesn't)

A bug-finding loop. The trick is the engine being scored is the real one — so a green score means the actual logic works, not a mock of it.

What runs each round

  • Generate — Gemini writes ~10 messy cases per tenant + what each reply should contain.
  • Run — each message goes through the real 6-stage engine.
  • Score — two independent judges (below).
  • Triage — failures sorted into 3 buckets (below).
  • Fix & re-run — patch the engine, replay, confirm green.

Two judges, on purpose

Deterministic checks = hard guarantees: did it route to the right intent? does every number trace to a real fact (no invented prices)? did it avoid offering a sold car? correct politeness particle?

LLM-judge = the soft stuff: is the Thai natural and helpful? Rated 1–5. The judge only decides the verdict when the reply was really generated by Gemini — never on offline fallbacks, so the score can't be faked.

Triage — every failure is one of three things

Bucket A

Engine bug. The code is actually wrong. Fix the engine. (All the fixes in the next section.)

Bucket B

Scorer false-positive. The reply was fine; the check was too strict. Fix the scorer. (e.g. the customer's own number counted as "invented".)

Bucket C

Bad generated expectation. The test itself was unreasonable. (e.g. labelling a car-seeking greeting as pure small-talk.) Not an engine problem.

Real results — adversarial rounds 16–25 (~200 cases)

Roundบ้านกุ้ง (casual ค่ะ)Bangkok Auto Premium (formal ครับ)
168/1010/10
178/10 → 10/10 after fixes9/10 → 10/10 after fixes
1810/108/10
198/109/10
208/1010/10
2110/1010/10
228/109/10
239/109/10
2410/108/10
2510/1212/12

The numbers don't only climb — and that's the point. Each fresh round throws new messy inputs, so dips surface genuinely new edge cases. By round 25 the remaining misses are almost all Bucket C (the evaluator being too strict), not engine defects.

the payoff · real bugs caught & fixed

What the loop actually found

These are real failures from real generated Thai messages, with the real root cause and fix. Click through them.

1 · The bot couldn't answer "how much is the transfer fee?"

Input: ค่าโอนรถเท่าไหร่ครับ — a dead-simple question with an exact-match policy doc.

before
Retrieved finance, tradein, warranty — the transfer doc with the ฿2,000–3,000 answer was never retrieved. Bot couldn't give the fee.
after
Retrieves bk-kb-transfer #1; reply states ฿2,000–3,000, free for cash buyers.
root cause
pythainlp wasn't installed, so the fallback tokenizer (only ~70 words) shattered Thai into single characters. Those lone characters matched every document equally — pure noise drowned the one real match.
fix
Installed pythainlp newmm for proper word segmentation, and added a content-token filter so single-character noise can never dominate retrieval again.

2 · "Do you give a warranty?" → "I don't have that information"

Input: มีประกันหลังการขายให้ไหมคะ — there is a warranty doc.

before
…ตอนนี้น้องกุ้งยังไม่มีข้อมูลในส่วนนี้เลยค่ะ (judge: 1/5)
after
…มีรับประกันเครื่องยนต์และเกียร์ 6 เดือน หรือ 10,000 กม. ค่ะ
root cause
The reranker sorted by raw count of shared words, so a doc sharing many common words (การ/ให้/รถ) beat the warranty doc, which shared only one — the important one. And the doc title wasn't even searched.
fix
Index the title + body, and rerank by IDF²-weighted overlap so one rare on-topic term beats many generic ones. The distinctive word wins.

3 · "Not in stock" — for a car that was literally in stock

Input: ฟอร์จูนเนอร์สีขาวปี 2018 … ลดได้เท่าไหร่ มีงบเก้าแสน — this exactly matches BK-004 (฿949,000).

before
…ในสต๊อกยังไม่มีเข้ามาเลยค่ะ (judge: 2/5)
after
…ราคาอยู่ที่ 949,000 บาท และตอนนี้รถยังอยู่ค่ะ → routes price talk to sales
root cause
The customer's budget (฿900k) was applied as a hard price filter, so the ฿949k car was excluded — even though the customer was negotiating toward that budget, not setting a ceiling.
fix
During a negotiation, budget is a target, not a filter. Plus a near-match fallback: if a named model exists but doesn't match every constraint, show the real unit instead of denying stock.

4 · Asked for a pickup, got recommended a sedan

Input: ดีแมค4ประตูออโต้ปี63มีไหมขอราคาสุดๆสด

before
…ดีแมคยังไม่มีเข้ามาเลยค่ะ แต่มี Honda City … แทนค่ะ (fabricated a sedan)
after
…ดีแมค… ตอนนี้ตัวรถติดจองอยู่ค่ะ (recognises the model, honest status)
root cause
The spelling ดีแมค wasn't in the engine's alias table, so the model wasn't recognised and the search returned junk.
fix
The data bank fixed the engine. The provisional lexicon (built by Codex + Gemini) already had ดีแมค → ISUZU_DMAX plus 5 more spellings. The collect→screen→apply loop paid off.

5 · Malformed Thai politeness particles

The LLM was told "always end with ค่ะ" — and overdid it.

before
…แจ้งวันและเวลาได้เลยนะคะค่ะ  ·  …ดูรูปก่อนไหมค่ะ
after
…ได้เลยนะคะ  ·  …ดูรูปก่อนไหมคะ
root cause
Generative models drift on sentence-final particles: doubling (นะคะ+ค่ะ), and using the falling ค่ะ where a question needs the rising คะ.
fix
A deterministic post-processor (_fix_particles) guarantees correct particles regardless of what the LLM emits — and the style checker now flags both errors so they can't regress silently.

6 · "Reserved" reported as "sold"

Input: ดีแมกซ์สีเทาปี 2019 ยังอยู่ไหมคะ — that unit is reserved, not sold.

before
…ขายไปเรียบร้อยแล้วค่ะ (wrong — it may free up)
after
…ตอนนี้ตัวรถติดจองอยู่ค่ะ แต่ยังมีโอกาสหลุดจองได้ถ้าดีลไม่ผ่านค่ะ
root cause
The engine lumped sold and reserved together as "gone". A reserved car can come back; telling the customer it's sold loses a live lead.
fix
Distinguish the two statuses explicitly: reserved → ติดจองอยู่, never ขายแล้ว.

Also fixed this way: multi-car comparison (x3 กับ ux ราคาเท่าไหร่ now answers both), typo-routing (จัดไฟแน้น → financing), small-talk no longer fires alongside real intents, and a scorer that wrongly flagged the customer's own requested discount as an "invented number".

why it stays trustworthy

The discipline underneath

These rules are what stop the system from quietly poisoning itself — important whichever path we pick next.

eval / train separation

The golden set (human-blessed tests) is never used to teach the bot. If you train on your test set, your scores become fiction.

provisional ≠ human-blessed

AI-screened data lives in a separate tier and is clearly labelled. A human verdict always overrides it. No "AI approved its own homework".

single source of truth

Domain rules (the 7 intents, status meanings) live in one place. The auto-reviewer is even guarded so it can't inject an intent the engine doesn't understand.

deterministic where it counts

Prices, years, particles, groundedness — all handled by code, not LLM guesswork. The LLM only does what only an LLM can: natural language.

where this leaves us

A proven core, ready for the real decision

What's actually been de-risked: the hard, copy-resistant parts — Thai understanding, hybrid retrieval, grounded generation, persona control, and a measurement loop that catches regressions. The mock data underneath is the easy part to swap.

The engine logic is real and tested. The data bank is compounding on AI-screened vocabulary. The only thing the prototype hasn't touched is a production database, real channels (LINE/Messenger), and the dashboard/ROI funnel — which is exactly the next conversation.

The three paths (for our next discussion — not decided here)

PathWhat it meansWhat carries over
Throwawaythe engine was only ever a learning/eval tool; rebuild the product cleanthe knowledge (Class A/B design, the bug fixes as a spec, the golden set + lexicon)
Evolve-in-Pythonharden this Python engine into the real service (add DB, channels, auth)everything — code, data bank, eval harness become production assets
Rebuild-in-Laravelreimplement the engine in your usual stack; keep Python only as an NLP sidecarthe data bank, eval harness, and validated logic as a reference implementation

In all three, the data bank, the golden set, and every documented bug-fix are durable — they're knowledge about Thai car-buyer language, not throwaway code.