Account: ERP-AI-Chatbot
Enter 6-digit code from Google Authenticator

Query Logs

Eloquent LLM (v3)

version-3 catalog only

Ask Eloquent / relation questions against iris-v3 (vLLM served weights) using version-3 as the only catalog. A deterministic orchestrator runs lookup → open_model → optional fetch_erp; the model only copies JSON (found, method, type, fk, table, warnings[]). After a class + relation resolve, an optional identifier (WO / invoice / part / staff id) loads that ERP row read-only and runs only the catalog-approved relation method. No identifier → metadata only.

Catalog
283 models
Generated
2026-10-02T18:30:16+00:00
LLM
iris-v3 (Qwen/Qwen2.5-14B-Instruct-AWQ via vLLM)
vLLM URL
http://127.0.0.1:8002

Ask the catalog

Enter submits. Class is resolved from nav table_name (including snake/plural of the class), short_name, legacy slug, or acronyms (GRN, WOPPR, WO, SO, PO, PR, JO). Exact table-name hits answer from the catalog without the leftover LLM. Ambiguous asks analyze with thinking — that can take several seconds. Add a WO / invoice / part / staff id to fetch ERP rows.

Result

—

Class
Table
Method
Nav hit
Attached
Hop-2
Tools
Model
Parse source
Latency

v3 test reports

Read-only render of the version-3 harness reports. Tabs: n=1 / N=20 control / 0.3 power.

Historical: 2026-08 local-LLM reference test (Ollama host, since retired)

Production leftover LLM is vLLM iris-v3 on :8002. This file is a dated run record.

Ollama vs version-3 catalog — local-LLM reference test

Setup

Field Value
Model qwen3:30b (Ollama tags: family qwen3moe, 30.5B, Q4_K_M). Already loaded (/api/ps).
Host http://127.0.0.1:11434 (intec-erp-chatbot/.env OLLAMA_URL / OLLAMA_DEFAULT_MODEL)
API POST /api/generate, stream: false, options.temperature: 0.3 (chatbot default), timeout 120s
Started 2026-08-20T02:13:40+08:00
Finished 2026-08-20T02:15:43+08:00 (last created_at on case 7)
Catalog eloquent-model-relation/version-3/ (llm/index.json generated_at 2026-08-19T18:07:17+00:00; 283 models, 845 relations, mean file 1990 B, max 8042 B workorder.json)
Retrieval README contract: grep one llm/nav.json object, open one llm/models/{node_key}.json. WO file ~8 KB → prompt used class_name, table_name, relations[], map_notes only.
Ground truth Attached v3 JSON. PHP sampled only to confirm JSON (no contradiction).

Scoring: pass = reply names the JSON fields (method / type / fk / table) and does not add a method absent from relations[]. partial = did not invent, but did not emit the expected tokens. fail = invents a method/table the JSON does not have.

Pass plan (before scoring)

Pass Question Leader after pass Contradiction?
1 Artifacts: README, nav hits, model files exist and node_key == filename contract usable no
2 Schema: relations[] is method/type/fk; inbound leaks live in map_notes v3 encodes the GRNI/User traps no
3 Completeness vs PHP on the asked edges JSON matches PHP no
4 Relation accuracy (type + fk) JSON matches PHP no
5 Query reconstructability (WO lines → inventory) hop is in the two files no
6 Token / navigability (one nav + one/two JSON) prompts 0.9–4.8 KB no
7 Failure modes on live generate lookup holds; negative answer is weak yes — format vs “none”
8 Adversarial: invent WO→GRNI hasMany? model did not invent no (overturn attempt failed)

Results

Case Expected (from JSON) Model response (scored text; thinking discarded) Latency Score
1 WO line items work_order_items, hasMany, wo_id — llm/models/workorder.json relations[1] work_order_items,hasMany,wo_id 13.361s pass
2 WO sales order salesOrderItem, belongsTo, so_item_id — same file relations salesOrderItem,belongsTo,so_item_id 10.635s pass
3 User employee employee, hasOne, staff_id; must not invent hasMany PurchaseOrder — llm/models/user.json country,belongsTo;userrole,belongsTo,id;employee,hasOne,staff_id 29.278s pass
4 GRN table goodsreceiptsnotes not goods_receipts_notes — llm/models/goodsreceiptsnote.json table_name goodsreceiptsnotes 2.942s pass
5 Adversarial GRNI none — no GoodsReceiptsNoteItem / GoodRecieptsNoteItem in WorkOrder relations[]; map_notes says inbound/collapsed , , , 44.843s partial
6 Two-hop materials hop2 inventory, belongsTo, inventory_id — llm/models/workorderitem.json inventory,belongsTo,inventory_id 3.251s pass
7 Inventory uom + count uom, belongsTo, uom_id; relations[] length 21 — llm/models/inventory.json uom,belongsTo,uom_id,21 18.437s pass

Count: 6 pass / 1 partial / 0 fail.

Thinking lived in the Ollama thinking field (not <think> tags in response). Case 5 thinking was 19135 chars / 4480 eval tokens.

Evidence (named files)

Pass 1–2 — contract and schema

  • version-3/README.md: grep "short_name": "WorkOrder" in llm/nav.json → node_key workorder → llm/models/workorder.json. relations[] = methods on this class only.
  • Nav hits used: llm/nav.json WorkOrder L1933–1938 (node_key workorder, table work_orders); User L1835–1839; GoodsReceiptsNote L632–637 (node_key goodsreceiptsnote, table goodsreceiptsnotes, legacy_diagram_slug grn); WorkOrderItem L1960–1965; Inventory L676–680.
  • workorder.json is 8042 B (under the 8 KB hub cap; llm/index.json over_8kb: []). Prompt stripped to 30 relations + map_notes.

Pass 3–4 — JSON vs PHP (sampled because replies matched JSON)

JSON claim PHP
WorkOrder::work_order_items() hasMany wo_id intec-erp-v2/app/Models/WorkOrder.php L125–127
WorkOrder::salesOrderItem() belongsTo so_item_id same file L177–179
App\User::employee() hasOne staff_id,staff_id intec-erp-v2/app/User.php L54–56; country / userrole L46–52. No purchaseOrder function.
GRN $table intec-erp-v2/app/Models/GoodsReceiptsNote.php L13 public $table = 'goodsreceiptsnotes';
WO has no GRNI method rg on WorkOrder.php for GoodReciepts|GoodsReceipts → 0 hits
WorkOrderItem::inventory() belongsTo inventory_id WorkOrderItem.php L60–62
Inventory::uom() Inventory.php L90; 21 relations[] objects in inventory.json

Pass 5–6 — reconstructability and tokens

  • README worked example (WO materials): work_order_items() then WorkOrderItem::inventory(). Cases 1 and 6 returned those method names and keys.
  • Prompt sizes: case 3 = 951 B (full user.json fields); case 4 = 1061 B; case 7 = 2438 B; WO cases ≈ 3.6 KB; hop case 6 = 4834 B. No full nav.json dump.

Pass 7 — failure modes (observed, not guessed)

  1. Hallucinated methods: not observed. Case 3 did not emit PurchaseOrder despite v1 map_notes warning in user.json. Case 5 did not emit a WO→GRNI hasMany.
  2. Wrong table: not observed. Case 4 returned goodsreceiptsnotes, not v2 goods_receipts_notes (README.md do-not-trust #5).
  3. Ignored JSON: not observed on lookup cases. Replies 1, 2, 4, 6, 7 are copies of JSON fields.
  4. Negative-answer format: case 5 thinking cites map_notes (“not a method on this class”) and scans related for GoodsReceiptsNoteItem / GoodRecieptsNoteItem, finds none, then loops because the system line demanded method, type, fk. Final response is empty CSV slots , , , — not the word none, and not an invented method. 44.8s / 4480 eval tokens.
  5. Over-list: case 3 listed all three relations[] (country has no fk in JSON; userrole fk id matches JSON). Extra names are in the file, not invented.
  6. Count self-check: case 7 thinking first said 20, then listed 21 methods (currency … taxes) and answered 21.

Pass 8 — adversarial re-check

Tried to overturn “model follows relations[] only” by asking for a WO hasMany to GRNI — the exact inbound leak in workorder.json map_notes and README hallucination table ($wo->goodRecieptsNoteItems). The model refused to invent. That does not overturn the leader. It does show the compact “one line: method, type, fk” prompt is a bad fit for “none”.

Review

qwen3:30b can follow the v3 retrieval contract when given one nav object plus one (or two) model JSON files: it copies method / type / fk / table_name from relations[] and does not fall back to v1 invented edges (User→PurchaseOrder, WorkOrder→GoodRecieptsNoteItem) or the v2 GRN snake_case table.

It does not reliably emit a clear negative (none) when the asked edge is absent. The failure mode is format collision + long thinking, not a hallucinated WO method. IRIS should add one line to the system prompt: If it is not in relations[], answer none — do not fill method/type/fk.

Thinking cost is real: lookup cases 3–13s (case 4/6 ~3s); User dump 29s; adversarial 45s. response itself stayed one line.

Recommendation

Yes — v3 is usable as the IRIS local-LLM Eloquent reference as-is for attached single-file lookups (method, type, fk, table), provided the IRIS prompt states that missing relations[] rows answer none and the agent never dumps llm/nav.json or all models.

Historical: 2026-08 Ollama N=20 control (Ollama since retired)

Production leftover LLM is vLLM iris-v3 on :8002.

Ollama vs version-3 catalog — N=20 no-file control

Second suite. First qualitative pass: OLLAMA_V3_TEST.md (n=1, no control, CSV answers). This run adds a with_json vs no_file control, grammar-constrained JSON, retrieval/hop tool-loop, generated traps, and N=20.

Setup

Field Value
Model qwen3:30b (already in VRAM; keep_alive: 24h)
Host http://127.0.0.1:11434 Ollama 0.32.5
API POST /api/generate, stream: false, format: <json-schema> (per case class)
Thinking think: false for copy / table; think: true for negative / hop / retrieval
Decode temperature 0 (copy/table) or 0.3 (think-on); num_ctx 4096 or 8192; num_predict 256
Started 2026-08-20T02:33:01+08:00
Finished 2026-08-20T02:53:20+08:00
Wall 1220.0 s (20.3 min) for 680 trials (17 cases × 2 modes × 20)
Catalog eloquent-model-relation/version-3/llm/ (generated_at 2026-08-19T18:07:17+00:00; 283 models)
Harness version-3/scripts/ollama_v3_harness.py
Cases version-3/tests/cases.json
Raw version-3/tests/results/run-20260820-023301.jsonl
Drift python3 scripts/check_catalog_drift.py → clean (286 files, 0 mismatches)

Scoring: parse JSON from response, else thinking. If expected.found === false, require found === false. If expected.found === true, require found === true and byte-equal method / type / fk / table / node_key when present. Hop with_json also requires requesting/using workorderitem. No CSV, no 45 s empty-slot loop.

Pass plan (before scoring)

Pass Question Leader after pass Contradiction?
1 Artifacts: harness, cases, trap generator, drift script exist contract runnable no
2 Schema: grammar JSON replaces CSV; absence = {found:false} format usable no
3 Completeness of live suite vs first 7 + traps + retrieval suite covers review items 1–7 no
4 Accuracy vs PHP on asked edges (sample) JSON matches PHP no
5 Delta: with_json − no_file (what proves v3) v3, not convention yes vs first review — cases 1/2/6 are not filler
6 Retrieval + hop tool-loop nav node_key + second file no
7 Failure modes, latency, think flag think-off copy is cheap; retrieval wall inflated no
8 Adversarial: can convention / abstention overturn “v3 earns keep”? overturn fails no

Results

Case class with_json% no_file% delta n
orig-01-wo-line-items copy 100.0 0.0 +100.0 20/20
orig-02-wo-sales-order copy 100.0 0.0 +100.0 20/20
orig-03-user-employee copy 100.0 0.0 +100.0 20/20
orig-04-grn-table table 100.0 0.0 +100.0 20/20
orig-05-wo-grni-negative negative 100.0 100.0 +0.0 20/20
orig-06-hop-inventory-fk hop 100.0 0.0 +100.0 20/20
orig-07-inventory-uom copy 100.0 0.0 +100.0 20/20
retrieval-grn-node-key retrieval 100.0 0.0 +100.0 20/20
trap-table-goodrecieptsnoteitem table 100.0 0.0 +100.0 20/20
trap-table-gpb table 100.0 0.0 +100.0 20/20
trap-table-workorderos table 100.0 0.0 +100.0 20/20
trap-fk-goodsreceiptsnote-po_reference copy 100.0 0.0 +100.0 20/20
trap-fk-goodrecieptsnoteitem-inventory copy 100.0 0.0 +100.0 20/20
trap-inbound-user-purchaseorder negative 100.0 100.0 +0.0 20/20
trap-inbound-bom-goodrecieptsnoteitem negative 100.0 100.0 +0.0 20/20
trap-spelling-goodsreceiptsnote negative 100.0 100.0 +0.0 20/20
trap-hasmany-id-bom-inventory copy 100.0 0.0 +100.0 20/20

680 / 680 trials produced parseable JSON. 0 HTTP errors. with_json fails: 0. Adversarial found:false rate: 160 / 160 (100%).

Temperature 0 made copy/table answers identical across the 20 repeats after the first decode. Think-on negatives/hop still never flipped found.

Filler vs real regression suite

Kind Cases Why
Proves v3 (delta +100) 13 cases: all copy, table, hop, retrieval no_file never emitted the catalog tokens
Filler for delta (delta ≈ 0) 4 negative cases both modes already return {found:false} — measures format, not the file
Not filler (contra first review) orig-01, orig-02, orig-06 first review guessed “convention / zero delta”. N=20 shows +100

orig-04 GRN table is still the poster-child irregular table (goodsreceiptsnotes). It is not unique: GPB gpb_documents, GRNI goodrecieptsnoteitems, WorkOrderOS work_order_operation_schedules have the same +100 delta.

Evidence (named files)

Pass 1 — artifacts

  • Harness: version-3/scripts/ollama_v3_harness.py (compact attach = class_name, table_name, relations[] method/type/fk/related, map_notes).
  • Cases: version-3/tests/cases.json (7 original + retrieval + 9 generated traps).
  • Full trap dump: version-3/tests/generated_traps.json (550 rows: 45 irregular tables, 475 non-convention FKs, 21 inbound, 2 spelling, 7 hasMany fk=id).
  • Drift: version-3/scripts/check_catalog_drift.py + build_llm_catalog.py --out.
  • Smoke runs (schema debug): tests/results/run-20260820-022816.jsonl, run-20260820-023209.jsonl.

Pass 2 — JSON schema (mandatory item 2)

Legal objects only. Per-class grammars after a smoke fail where the shared schema let the model put the table in fk (method: table_name, fk: goodsreceiptsnotes):

class required shape
copy {found, method, type, fk}
table {found, table}
negative {found} (+ optional need_file)
hop copy fields + need_file / lookup
retrieval {found, node_key, table} + lookup

think: false is supported on 0.32.5 and writes JSON to response. think: true (and omitted think on qwen3) writes the constrained JSON into thinking with empty response — 240 / 680 trials, all still parseable. No , , , rows, no 44 s format loop from the first suite.

Pass 3–4 — suite coverage and PHP sample

Live traps were bounded (9 of 550) so N=20 stayed under 90 minutes. They cover each generator kind except duplicating orig-04/05.

Claim PHP
WorkOrder::work_order_items() hasMany wo_id intec-erp-v2/app/Models/WorkOrder.php L125–127
WorkOrder::salesOrderItem() belongsTo so_item_id workorder.json relations + same PHP file
App\User::employee() hasOne staff_id intec-erp-v2/app/User.php L54–56
GRN $table GoodsReceiptsNote.php L13 public $table = 'goodsreceiptsnotes'
GRNI inventory() belongsTo item_id GoodRecieptsNoteItem.php L88–90
Bom::inventory() hasMany 'id','bom_id' Bom.php L154–156 (code_may_be_wrong)

Drift rebuild to a temp dir matched committed llm/ (timestamps stripped). The 2026-08-20 ERP pull did not change parsed relations for these files.

Pass 5 — no_file fingerprints (what convention actually does)

Deterministic wrong guesses (20/20 unless noted):

Case no_file output Why it fails
orig-01 found:false + lineItems / hasMany / work_order_id PHP is work_order_items / wo_id
orig-02 19× {found:false}; 1× salesOrderItem + sales_order_item_id catalog FK is so_item_id
orig-03 19× {found:false}; 1× fk: employee_id catalog is hasOne / staff_id
orig-04 / all table traps {found:false} abstains; never emits goodsreceiptsnotes / gpb_documents
orig-07 {found:false} does not emit uom / uom_id
orig-06 hop need_file: InventoryItem (14) / Inventory (4) never workorderitem then inventory_id
retrieval 15× {found:false}; 4× need_file: grn_notes; 1× need_file: grn slug trap: grn is legacy_diagram_slug, not node_key
trap GRN po_reference {found:false} does not emit po_number
trap GRNI inventory {found:false} does not emit item_id
trap Bom inventory belongsTo / inventory_id + found:false catalog/PHP is hasMany / fk: id

The no_file system line “If you are not sure, return found false” increases abstention. That does not create the +100 deltas by itself: every time the model guessed, the guess was Laravel convention, not this ERP’s PHP.

Pass 6 — retrieval and hop

Retrieval (nav candidates only; decoys: legacy_diagram_slug: grn, GRNI/grni, GrnCancellationRequest, Gpb). 20/20 with_json picked node_key: goodsreceiptsnote and table: goodsreceiptsnotes on step 0. 0/20 no_file picked that key; one trial asked for grn.

Harness then injected goodsreceiptsnote.json and called generate again even though step 0 already matched — retrieval mean wall 37.7 s is an extra prefix-compile, not thinking tokens (thinking_chars mean 40). Fixed after this run: stop when the parsed object already matches expected (non-hop).

Hop (only workorder.json attached). 20/20 step 0: {found:false, need_file: WorkOrderItem} (normalized to workorderitem). 19/20 then copied inventory / belongsTo / inventory_id from the second file. 1/20 asked for Inventory as a third hop, then still answered correctly. hop1_leak (answering work_order_items as if it were inventory) = 0 / 20. no_file never received a second file (control) and scored 0.

Pass 7 — latency and API notes

Condition n mean wall median max
think: false (copy/table) 440 0.61 s 0.26 s 19.48 s
think: true (neg/hop/retrieval) 240 3.96 s 0.26 s 38.75 s

The ~19 s spike is a first-token / prefix-compile tax on a new prompt, not eval: eval_duration stays ~0.2–0.3 s for ~20 tokens. Repeats of the same prefix drop to 0.2–0.5 s. think: true does not add a long chain-of-thought here — the grammar dumps the JSON object into thinking (max 87 chars).

First-suite case 5 was 44.8 s / 4480 eval tokens of CSV looping. This suite’s negatives are 0.2–1.2 s mean with {found:false}.

think: false in the generate body is accepted. Putting think only under options does not disable thinking on this build.

Pass 8 — adversarial re-check

Tried to overturn “v3 earns its keep”:

  1. Maybe cases 1/2/6 are convention-solved. Overturn fails. orig-01 invents lineItems/work_order_id 20/20. orig-02 never emits so_item_id. orig-06 never opens workorderitem without the catalog.
  2. Maybe negatives show the file does nothing. True for delta, false for format: first suite could not say none; this suite says {found:false} 160/160. Negatives are a format regression test, not a retrieval test.
  3. Maybe code_may_be_wrong Bom.inventory fk:id is a bad expected. The catalog copies PHP (hasMany(Inventory::class,'id','bom_id')). with_json copies that row 20/20; no_file “corrects” it to belongsTo/inventory_id. That is exactly when IRIS must trust v3 (and then a human) rather than Laravel folklore.

Trap-set coverage

Generator kind Enumerated Live (this run)
table_not_snake_plural 45 3 (+ orig-04 GRN)
fk_not_singular_id 475 2 (GRN po_number, GRNI item_id) + orig-01/02
inbound_only_map_notes 21 2 (User→PO, Bom→GRNI) + orig-05
receipts_misspelling 2 1 (GoodsRecieptsNote)
hasmany_fk_id 7 1 (Bom.inventory, tagged code_may_be_wrong)

Full list: tests/generated_traps.json. Live IDs are the from_generated rows in tests/cases.json.

Drift

python3 scripts/check_catalog_drift.py
# clean: true; committed=286, rebuilt=286; only_*=[]; content_mismatch=[]

Ignores generated_at. No GitHub Action; README documents the local command.

Review vs OLLAMA_V3_TEST.md

First suite This suite
n=1, no control n=20 × 2 modes
CSV method,type,fk — case 5 → , , , in 45 s schema {found:false} in <1 s median
Guessed cases 1/2/6 may be zero-delta filler all three are +100
“v3 usable as-is for attached lookups” Confirmed at 100% / 20; also required for tables, non-convention FKs, hop2, and node_key vs grn
Prompt should say none Use grammar JSON instead of a prose none
Thinking cost 3–45 s on lookups think: false → ~0.3 s after prefix warmup

Recommendation

v3 earns its keep. Attach the compact model JSON (or nav → node_key → one file). Do not ask qwen3:30b to reconstruct this ERP from Laravel memory: it abstains or emits lineItems/work_order_id/sales_order_item_id/goods_receipts_notes-style folklore.

IRIS should:

  1. Grep one llm/nav.json object; open llm/models/{node_key}.json — never the legacy_diagram_slug (grn ≠ goodsreceiptsnote).
  2. Constrain output with the JSON schema above. Missing edge → {found:false}.
  3. Disable thinking on field-copy / table questions (think: false top-level). Allow thinking (or a tool-loop) on hop/negative.
  4. For WO materials, if only the header file is loaded, expect need_file: workorderitem before answering inventory_id.
  5. Run scripts/check_catalog_drift.py after builder or PHP-model edits.

What the first suite still did better: a human-readable walk through thinking traces. What this suite adds: statistical deltas and a machine-parseable contract.

Ollama vs version-3 catalog — production temperature 0.3 (N=20)

Third suite. Prior reviews: OLLAMA_V3_TEST.md (n=1, CSV), OLLAMA_V3_TEST_N20.md (temp-0 copy/table N=20). This run cites decode temperature 0.3 (IRIS production), adds a wrong-file class, a no-abstention folklore arm, warnings[] on code_may_be_wrong, a full 550-trap sweep, and an opt-in drift hook.

Do not treat the earlier temp-0 100% repeats as a pass rate. Temp 0 is a cheap regression tripwire (identical tokens after the first decode). Citable rates below are all temperature 0.3.

Setup

Field Value
Model qwen3:30b (keep_alive: 24h)
Host http://127.0.0.1:11434 Ollama 0.32.5
API POST /api/generate, stream: false, format: <json-schema>
Production decode temperature 0.3
Tripwire decode temperature 0, N=3, think: false, copy/table only
Thinking think: false for copy/table/wrong_file/trap sweep; think: true for hop/retrieval/negative in the live 0.3 suite
Started 2026-08-20T03:29:34+08:00
Finished 2026-08-20T03:51:10+08:00
Wall 21.6 min (1296 s) for 1483 trials (no HTTP errors; every trial parseable)
Catalog eloquent-model-relation/version-3/llm/ (generated_at 2026-08-19T18:07:17+00:00; 283 models)
Harness version-3/scripts/ollama_v3_harness.py
Cases version-3/tests/cases.json + tests/generated_traps.json (sweep ran 550; generator now emits 547 after dropping 3 stale inbound rows)
Raw tests/results/run-20260820-033000-*.jsonl
Drift python3 scripts/check_catalog_drift.py → clean; hook installed (opt-in)
Schema smoke tests/results/run-20260820-032822-smoke-schema.jsonl (3 trials, not in the 1483)

Scoring: parse JSON from response, else thinking. found:false expected → require found === false. found:true expected → require found === true and byte-equal method / type / fk / table / node_key. code_may_be_wrong also requires a non-empty warnings[]. Wrong-file fails if found:true with a method/table from the attached file. Folklore% is the found:true rate on no_file_raw (committed guess), not a pass rate.

Pass plan (before scoring)

Pass Question Leader after pass Contradiction?
1 Artifacts: new modes, schema warnings[], hook, 550 sweep files contract runnable no
2 Schema / format at 0.3 (warnings, wrong_file, no_file_raw) grammar usable no
3 Completeness vs request A–G suite covers all items no
4 Accuracy vs PHP on cited edges + fail sample cited edges match PHP yes — 3 inbound traps were stale vs PHP
5 Citable delta at 0.3 (with_json − no_file) v3, not convention no vs N=20 temp-0 review
6 Wrong-file + folklore arms model usually abstains; WO/WOI copies new failure class
7 Full trap sweep + N=20 on fails 21/550 N=1 fails; 8 stay 0/20 yes vs “attach ⇒ 100% copy”
8 Adversarial: can trap misses / wrong-file copies overturn “cite the 0.3 table”? overturn fails for the orig 7; succeeds against “trust the model on every attached row” revised

Citable results (temperature 0.3, N=20)

This is the table to quote. Folklore% = found:true rate on no_file_raw (omit “if you are not sure, return found false”). — = arm not run for that case.

Case class with_json% no_file% folklore% delta
orig-01-wo-line-items copy 100.0 0.0 0.0 +100.0
orig-02-wo-sales-order copy 100.0 0.0 0.0 +100.0
orig-03-user-employee copy 100.0 0.0 — +100.0
orig-04-grn-table table 100.0 0.0 0.0 +100.0
orig-05-wo-grni-negative negative 100.0 100.0 — +0.0
orig-06-hop-inventory-fk hop 100.0 0.0 — +100.0
orig-07-inventory-uom copy 100.0 0.0 — +100.0
retrieval-grn-node-key retrieval 100.0 0.0 — +100.0
trap-hasmany-id-bom-inventory copy 100.0 0.0 — +100.0

Raw: tests/results/run-20260820-033000-power03.jsonl (360 trials) + …-folklore.jsonl (60). n=20/20 per cell except folklore (n=20 raw only on orig-01/02/04).

360 / 360 power03 trials parseable. 0 HTTP errors. with_json fails on these 9 cases: 0. Hop hop1_leak = 0/20. Retrieval node_key = goodsreceiptsnote 20/20.

Same +100 deltas as the temp-0 N=20 review, now at IRIS production temperature. Temperature 0.3 did not flip any cited with_json answer.

Folklore arm vs abstention no_file (same 0.3)

found:true folklore% is 0.0 on orig-01/02/04. Omitting the extra abstention sentence does not make qwen3:30b commit Laravel names as found:true. The SYSTEM line “Do not invent Laravel names” is enough to keep found:false.

Convention still leaks into optional fields (not scored as folklore%):

Case no_file (abstention line on) no_file_raw (line off)
orig-01 6/20 stuffed lineItems / hasMany / work_order_id with found:false 0/20 method leak; 4/20 wrote a “no catalog” warning
orig-02 0/20 method leak (all found:false) 0/20 method leak
orig-04 20/20 bare {found:false} 8/20 table: goods_receipts_notes with found:false; 3 more named that snake_case in warnings

So the folklore arm is not a higher committed-guess rate. It is a higher table-leak rate on GRN. The abstention line actually increased the orig-01 lineItems stuffing (6/20 vs 0/20). Either way the catalog tokens (work_order_items / wo_id / so_item_id / goodsreceiptsnotes) never appear without the file. Delta stays valid under production decode.

Temperature 0 tripwire (not a rate)

tests/results/run-20260820-033000-tripwire0.jsonl — 11 copy/table cases × N=3 × think: false = 33 trials, 153 s.

Every case copied the catalog row 3/3. After the first decode, repeats were byte-identical except trap-fk-goodsreceiptsnote-po_reference (2/3 added warnings:["code_may_be_wrong"]; method/type/fk still matched).

Do not cite “100%” from this arm. Temp 0 collapses sampling. Use it as a cheap “did the prompt/schema/file still copy?” tripwire after catalog or harness edits (N=3 is enough). Cite the 0.3 table above.

Wrong-file (class wrong_file, N=20, temp 0.3, think: false)

Attach JSON for class B; ask a question about class A. Expected: {found:false}. Fail if found:true with a method/table from the attached file.

Case attached asked pass copied wrong file abstain
wrong-file-grn-table-vs-grni goodrecieptsnoteitem.json (goodrecieptsnoteitems) GoodsReceiptsNote $table 20/20 0 20
wrong-file-wo-items-vs-woi workorderitem.json WorkOrder line-items relation 16/20 (80%) 4 16
wrong-file-grn-vs-grncancel grncancellationrequest.json (nav neighbor of GRN) GoodsReceiptsNote $table 20/20 0 20

Raw: tests/results/run-20260820-033000-wrongfile.jsonl.

Does the model copy or abstain? Mostly abstain. GRN vs GRNI (spelling-family near-miss) and GRN vs GrnCancellationRequest (nav neighbor) never copied the decoy table/methods (20/20 {found:false}, often need_file: goodsreceiptsnote).

It does copy on the WorkOrder / WorkOrderItem prefix family: 4/20 returned {found:true, method: work_order, type: belongsTo, fk: wo_id} — the child file’s work_order() row, not the header’s work_order_items / wo_id / hasMany. That is a wrong-file copy, not a memory guess (goodsreceiptsnotes / work_order_items never appeared).

IRIS must not trust the model to notice a class mismatch. Add a deterministic pre-check: requested short_name (or class_name) vs attached class_name. If they differ, do not call the model for a copy/table answer — return {found:false} / open the right node_key. The 80% WO/WOI pass rate is not a safety margin.

warnings[] — Bom::inventory (code_may_be_wrong)

Grammar: copy/table/hop schemas now allow optional "warnings": []. Harness asserts a non-empty array when the case is tagged code_may_be_wrong / require_warnings. Compact attach tags hasMany + fk: id with code_may_be_wrong: true and includes model flags on those cases.

Arm fields_ok warnings_ok pass
power03 with_json N=20 @ 0.3 20/20 20/20 20/20
tripwire0 N=3 @ 0 3/3 3/3 3/3
trap sweep N=1 (all 7 hasmany_fk_id) 7/7 7/7 7/7
power03 no_file N=20 0/20 — 0/20 (all {found:false})

18/20 warnings were exactly ["code_may_be_wrong"]. 2/20 also copied a catalog flag sentence about inverted hasMany(..., 'id', 'bom_id').

PHP: intec-erp-v2/app/Models/Bom.php L154–156 hasMany(Inventory::class,'id','bom_id'). no_file never emits fk: id. Pass. IRIS should surface warnings[] to a human; do not “fix” the row in generated queries unless the task is an ERP bugfix.

Full trap sweep (550 @ N=1, then N=20 on fails)

All generated traps, with_json only, think: false, temp 0.3. No traps dropped.

Stage File Trials Wall Pass
N=1 all traps run-20260820-033000-traps-n1.jsonl 550 558.4 s 529 / 550 (96.2%)
N=20 on the 21 fails run-20260820-033000-traps-n20.jsonl 420 225.3 s 90 / 420 (case-level: see below)

21 of 550 failed at N=1 (3.8%). All 21 got N=20. Fail kind was fields_mismatch only (0 HTTP, 0 unparseable).

N=20 on those 21

Case N=20 with_json% What the model did PHP / catalog
trap-fk-finishgoodstockstransaction-invoices 100.0 N=1 miss; N=20 always copied hasManyThrough / fk: id catalog row; recovered
trap-fk-supplier-taxes 95.0 19/20 belongsTo / default_tax catalog
trap-fk-salesorder-invoices 80.0 16/20 hasOne / so_id catalog
trap-fk-bominventory-stockinventories 35.0 often {found:false} belongsTo / inventory_id (not stockinventories_id)
trap-fk-bom-childbom 30.0 usually abstain hasOne / bom_id
trap-fk-supplier-systemStatus 30.0 usually abstain belongsTo / status
trap-fk-workorderprocesstraveller-item_transfer_statuses 30.0 usually abstain belongsTo / item_transfer_status
trap-fk-bom-bomClassType 5.0 almost always abstain hasMany / bom_id
trap-fk-location-locations 5.0 almost always abstain hasOne / fk: id
trap-fk-bomprice-moq 10.0 almost always abstain belongsTo / price_condition_name
trap-fk-finishgoodstockstransaction-finishGoodsStock 10.0 almost always abstain belongsTo / bom_id
trap-fk-process-master_process 10.0 almost always abstain belongsTo / fk: id
trap-fk-salesorderitem-workOrders 10.0 2/20 copied; else found:false (sometimes with the right fields) belongsTo / fk: id
trap-fk-bom-finish_good_transaction 0.0 {found:false} 20/20 row is in bom.json (hasMany / bom_id)
trap-fk-bom-dualnaturewithinventory 0.0 often copies method/type/fk then sets found:false hasOne / part_number in bom.json
trap-fk-joborderitem-work_orders 0.0 {found:false} 20/20 belongsTo / job_order_id (plural method)
trap-fk-purchaserequest-moq_price_inventory 0.0 {found:false} 20/20 belongsTo / price_condition
trap-fk-purchaserequest-purchase_order_items 0.0 {found:false} 20/20 belongsTo / fk: id
trap-fk-workorder-job_order_items 0.0 20/20 copied neighbor work_order_items / hasMany / wo_id PHP job_order_items() is hasOne(JobOrder, 'job_order_number', 'jo_id') (WorkOrder.php L154–156)
trap-inbound-wipunfinishedgood-bom 0.0 {found:true, need_file: BOM} 20/20 suite bug — see below
trap-inbound-wipunfinishedgood-uom 0.0 {found:true, need_file: Uom|UOM} 20/20 suite bug — see below

Hard zeros that are real model misses (8): Bom hub rows the file contains but the model will not found:true; JobOrderItem work_orders; PurchaseRequest odd FKs; WorkOrder job_order_items vs work_order_items near-miss.

Suite bugs (2 of the 21, plus 1 silent false-pass not in the fail list): v1 map_notes said Bom/Uom/Transaction were inbound, but PHP declares them:

  • WipUnfinishedGood::bom() / uom() — WipUnfinishedGood.php L57–64; also in wipunfinishedgood.json relations[].
  • StockHistory::transaction() — in stockhistory.json relations[]. That inbound trap passed N=1 (found:false) and was a false pass.

Generator fix (after this sweep): skip inbound names already in relations[]. Regenerated dump is 547 (inbound 21 → 18). Dropped: trap-inbound-wipunfinishedgood-bom, …-uom, trap-inbound-stockhistory-transaction. Adjusted N=1 fail rate: 19 / 547 real traps (the 2 WIP rows should not have been asked as negatives).

IRIS implication: attaching a hub file (Bom, WorkOrder, PurchaseRequest) is not a guarantee the model will copy an unusual row. Prefer a deterministic lookup (relations[] where method equals the asked name) over “ask qwen to find the row”. The orig-7 questions hit famous methods; the 8 hard zeros hit obscure / inverted / near-miss names.

Drift hook

Item Path / status
Check script (CI) version-3/scripts/check_catalog_drift.py
Hook source version-3/scripts/hooks/post-commit-catalog-drift.sh
Installer version-3/scripts/install-drift-hook.sh
Installed? yes — intec-erp-v2/.git/hooks/post-commit (no prior hook; nothing overwritten)
When it runs post-commit, only if app/Models/*.php or app/User.php is in HEAD
CI IRIS/chatbot CI can call the same Python script (do not depend on a laptop hook)

Opt-in: the installer copies the hook and backs up a pre-existing post-commit to post-commit.pre-v3-drift.bak. It does not force-enable on other clones. One-line enable: bash scripts/install-drift-hook.sh.

This run: drift clean (286 files, 0 mismatches).

IRIS client requirement — thinking | response

On Ollama 0.32.5, think: true (and omitted think on qwen3) puts constrained JSON in thinking with an empty response.

Suite think-on trials empty response + JSON in thinking
First N=20 (OLLAMA_V3_TEST_N20.md) 240 / 680 240 / 240
This power03 live suite 120 / 360 (hop/retrieval/negative) 120 / 120
This trap sweep (think: false) 0 JSON in response

This is an IRIS client requirement, not a harness footnote: parse response first, then thinking. Or disable thinking for copy/table (think: false top-level in the generate body; options.think does not disable it on this build).

Evidence (named files)

Pass 1 — artifacts

  • Harness: scripts/ollama_v3_harness.py — modes with_json / no_file / no_file_raw; class wrong_file; --temperature / --think / --generated / --resume / --failed-from.
  • Schema: copy/table/hop include optional warnings[]. --self-check covers warnings assert + wrong-file copy detect.
  • Cases: original 7 + retrieval + 8 live traps + 3 wrong-file. Folklore is a mode, not a case id.
  • Sweep driver: scripts/run_v3_power.sh (resume-safe).
  • Hook: source + installer as above; installed on this machine.
  • Results: tests/results/run-20260820-033000-{tripwire0,power03,folklore,wrongfile,traps-n1,traps-n20}.jsonl.

Pass 2 — schema

Smoke (run-20260820-032822-smoke-schema.jsonl): GRN table copied; Bom.inventory returned warnings:["code_may_be_wrong"]; wrong-file GRN-vs-GRNI returned {found:false, need_file: goodsreceiptsnote}. Production N=20 confirmed all three.

Pass 3 — coverage vs A–G

Item Done
A. temp 0 tripwire N=3; cite 0.3 N=20 orig 7 + retrieval + Bom.inventory; no_file at 0.3 yes
B. ≥3 wrong-file @ N=20 0.3 yes (3)
C. no_file_raw @ N=20 0.3 on orig-01/02/04 yes
D. warnings[] + Bom assert + trap generator require_warnings yes (7/7 hasmany_id)
E. 550 N=1 then N=20 on 21 fails; none dropped yes
F. opt-in hook, installed, CI note yes
G. IRIS must read thinking|response documented as client requirement

Pass 4 — PHP sample

Claim PHP
WorkOrder::work_order_items() hasMany wo_id WorkOrder.php L125–127
WorkOrder::job_order_items() hasOne(JobOrder, 'job_order_number', 'jo_id') WorkOrder.php L154–156 (model copied work_order_items instead — real miss)
Bom::inventory() hasMany(Inventory::class,'id','bom_id') Bom.php L154–156
GRN $table GoodsReceiptsNote.php L13 public $table = 'goodsreceiptsnotes'
WipUnfinishedGood::bom() / uom() are methods WipUnfinishedGood.php L57–64 — inbound traps were stale

Pass 5 — 0.3 delta (what proves v3)

Same shape as the temp-0 N=20 review, now citable at production decode: 7 of 9 cases +100; orig-05 negative +0 (format test). no_file fingerprints at 0.3 still never emit wo_id / so_item_id / staff_id / goodsreceiptsnotes / workorderitem / uom_id / node_key: goodsreceiptsnote. Retrieval no_file still hits the grn slug trap (need_file: grn_notes 6/20).

Pass 6 — wrong-file + folklore

See tables above. New: WO/WOI 4/20 copy; GRN decoys abstain. Folklore found:true = 0; GRN snake_case leaks only on the raw arm.

Pass 7 — latency / API

Condition n mean wall max HTTP
tripwire0 33 4.64 s 47.0 s 0
power03 360 0.90 s 19.8 s 0
folklore 60 0.28 s 0.6 s 0
wrongfile 60 0.27 s 0.5 s 0
traps N=1 550 1.01 s 19.5 s 0
traps N=20 fails 420 0.54 s 19.5 s 0

~19–47 s spikes are first-token / prefix-compile on a new prompt (eval_duration stays ~0.2–0.3 s). Repeats drop to 0.2–0.5 s. No API errors. Sweep was not truncated.

think: true on this build: 120/120 power03 think-on trials had empty response and JSON only in thinking.

Pass 8 — adversarial re-check

Tried to overturn “cite the 0.3 100% table / v3 earns keep”:

  1. Maybe 0.3 sampling flips copy/table. Overturn fails on the 9 cited cases (0 with_json fails / 180).
  2. Maybe wrong-file shows the model notices class_name, so IRIS can skip a pre-check. Overturn succeeds as a warning: 4/20 WO/WOI copies. Pre-check is mandatory.
  3. Maybe 550 traps are also 100%. Overturn succeeds. 21 N=1 fails; 8 stay 0/20. Attaching the file is not enough for obscure / near-miss methods on hub classes. IRIS should grep relations[] (or refuse) instead of asking the model to pick a row from a 20-relation Bom file.
  4. Maybe the 2 WIP inbound fails mean the model invents methods. Overturn fails — PHP already has bom() / uom(). Those expected {found:false} rows were stale v1 map_notes. Generator now skips them.
  5. Maybe folklore without the abstention line destroys the +100 delta. Overturn fails. found:true folklore% = 0. Convention leaks are extra fields, not catalog tokens.

Revised verdict: the orig 7 + retrieval + Bom.inventory table at 0.3 N=20 is still the thing to cite. The new work shows two IRIS-side controls the previous review understated: class_name pre-check and deterministic relation extract on hub files.

Review vs OLLAMA_V3_TEST_N20.md

N=20 temp-0 suite This suite
Cited 100% from temp-0 repeats Cite 0.3 only; temp 0 is a tripwire
No wrong-file class WO/WOI copies 4/20; GRN decoys abstain
Abstention no_file only no_file_raw: found:true still 0%; GRN table leak 8/20
Bom.inventory copied without warnings[] warnings[] 20/20 (code_may_be_wrong)
9 of 550 traps live 550 N=1 + N=20 on 21 fails
Drift = README command Opt-in post-commit hook installed; CI can call the same script
Think/response as harness note IRIS client requirement

Recommendation

v3 still earns its keep at IRIS production temperature (0.3). Attach the compact model JSON (or nav → node_key → one file). Do not ask qwen3:30b to reconstruct this ERP from Laravel memory.

IRIS should:

  1. Deterministic pre-check: attached class_name / short_name vs the asked class. Mismatch → {found:false} / open the right file. Do not trust the model (WO/WOI copied 4/20).
  2. Deterministic extract for a named method on a loaded file (relations[] match). Use the model for fuzzy questions (orig-01 “line-items”), not for “copy method X from this JSON” on hub classes.
  3. Constrain output with the JSON schemas (including optional warnings[]). Missing edge → {found:false}.
  4. think: false top-level on copy/table. Parse response then thinking on any think-on call (0.32.5 + qwen3 writes JSON to thinking with empty response — 240/680 then 120/120 in this suite).
  5. For WO materials, if only the header is loaded, expect need_file: workorderitem before inventory_id.
  6. Surface warnings[] when the catalog tags code_may_be_wrong (Bom::inventory).
  7. Run scripts/check_catalog_drift.py in IRIS/chatbot CI. Enable the post-commit hook with bash scripts/install-drift-hook.sh (opt-in; already installed here).

What the temp-0 suite still does better: a cheaper identical-copy tripwire. What this suite adds: citable 0.3 rates, wrong-file + folklore arms, warnings[], and a complete trap fail list instead of a 9-row sample.