Eloquent LLM (v3)
version-3 catalog only
Ask Eloquent / relation questions against iris-v3 (vLLM served weights) using version-3 as the only catalog. A deterministic orchestrator runs lookup → open_model → optional fetch_erp; the model only copies JSON (found, method, type, fk, table, warnings[]). After a class + relation resolve, an optional identifier (WO / invoice / part / staff id) loads that ERP row read-only and runs only the catalog-approved relation method. No identifier → metadata only.
Ask the catalog
Result
—
- Class
- Table
- Method
- Nav hit
- Attached
- Hop-2
- Tools
- Model
- Parse source
- Latency
ERP data
v3 test reports
Read-only render of the version-3 harness reports. Tabs: n=1 / N=20 control / 0.3 power.
Historical: 2026-08 local-LLM reference test (Ollama host, since retired)
Production leftover LLM is vLLM iris-v3 on :8002. This file is a dated run record.
Ollama vs version-3 catalog — local-LLM reference test
Setup
| Field | Value |
|---|---|
| Model | qwen3:30b (Ollama tags: family qwen3moe, 30.5B, Q4_K_M). Already loaded (/api/ps). |
| Host | http://127.0.0.1:11434 (intec-erp-chatbot/.env OLLAMA_URL / OLLAMA_DEFAULT_MODEL) |
| API | POST /api/generate, stream: false, options.temperature: 0.3 (chatbot default), timeout 120s |
| Started | 2026-08-20T02:13:40+08:00 |
| Finished | 2026-08-20T02:15:43+08:00 (last created_at on case 7) |
| Catalog | eloquent-model-relation/version-3/ (llm/index.json generated_at 2026-08-19T18:07:17+00:00; 283 models, 845 relations, mean file 1990 B, max 8042 B workorder.json) |
| Retrieval | README contract: grep one llm/nav.json object, open one llm/models/{node_key}.json. WO file ~8 KB → prompt used class_name, table_name, relations[], map_notes only. |
| Ground truth | Attached v3 JSON. PHP sampled only to confirm JSON (no contradiction). |
Scoring: pass = reply names the JSON fields (method / type / fk / table) and does not add a method absent from relations[]. partial = did not invent, but did not emit the expected tokens. fail = invents a method/table the JSON does not have.
Pass plan (before scoring)
| Pass | Question | Leader after pass | Contradiction? |
|---|---|---|---|
| 1 | Artifacts: README, nav hits, model files exist and node_key == filename |
contract usable | no |
| 2 | Schema: relations[] is method/type/fk; inbound leaks live in map_notes |
v3 encodes the GRNI/User traps | no |
| 3 | Completeness vs PHP on the asked edges | JSON matches PHP | no |
| 4 | Relation accuracy (type + fk) | JSON matches PHP | no |
| 5 | Query reconstructability (WO lines → inventory) | hop is in the two files | no |
| 6 | Token / navigability (one nav + one/two JSON) | prompts 0.9–4.8 KB | no |
| 7 | Failure modes on live generate | lookup holds; negative answer is weak | yes — format vs “none” |
| 8 | Adversarial: invent WO→GRNI hasMany? |
model did not invent | no (overturn attempt failed) |
Results
| Case | Expected (from JSON) | Model response (scored text; thinking discarded) |
Latency | Score |
|---|---|---|---|---|
| 1 WO line items | work_order_items, hasMany, wo_id — llm/models/workorder.json relations[1] |
work_order_items,hasMany,wo_id |
13.361s | pass |
| 2 WO sales order | salesOrderItem, belongsTo, so_item_id — same file relations |
salesOrderItem,belongsTo,so_item_id |
10.635s | pass |
| 3 User employee | employee, hasOne, staff_id; must not invent hasMany PurchaseOrder — llm/models/user.json |
country,belongsTo;userrole,belongsTo,id;employee,hasOne,staff_id |
29.278s | pass |
| 4 GRN table | goodsreceiptsnotes not goods_receipts_notes — llm/models/goodsreceiptsnote.json table_name |
goodsreceiptsnotes |
2.942s | pass |
| 5 Adversarial GRNI | none — no GoodsReceiptsNoteItem / GoodRecieptsNoteItem in WorkOrder relations[]; map_notes says inbound/collapsed |
, , , |
44.843s | partial |
| 6 Two-hop materials | hop2 inventory, belongsTo, inventory_id — llm/models/workorderitem.json |
inventory,belongsTo,inventory_id |
3.251s | pass |
| 7 Inventory uom + count | uom, belongsTo, uom_id; relations[] length 21 — llm/models/inventory.json |
uom,belongsTo,uom_id,21 |
18.437s | pass |
Count: 6 pass / 1 partial / 0 fail.
Thinking lived in the Ollama thinking field (not <think> tags in response). Case 5 thinking was 19135 chars / 4480 eval tokens.
Evidence (named files)
Pass 1–2 — contract and schema
version-3/README.md: grep"short_name": "WorkOrder"inllm/nav.json→node_keyworkorder→llm/models/workorder.json.relations[]= methods on this class only.- Nav hits used:
llm/nav.jsonWorkOrder L1933–1938 (node_keyworkorder, tablework_orders); User L1835–1839; GoodsReceiptsNote L632–637 (node_keygoodsreceiptsnote, tablegoodsreceiptsnotes,legacy_diagram_sluggrn); WorkOrderItem L1960–1965; Inventory L676–680. workorder.jsonis 8042 B (under the 8 KB hub cap;llm/index.jsonover_8kb: []). Prompt stripped to 30 relations +map_notes.
Pass 3–4 — JSON vs PHP (sampled because replies matched JSON)
| JSON claim | PHP |
|---|---|
WorkOrder::work_order_items() hasMany wo_id |
intec-erp-v2/app/Models/WorkOrder.php L125–127 |
WorkOrder::salesOrderItem() belongsTo so_item_id |
same file L177–179 |
App\User::employee() hasOne staff_id,staff_id |
intec-erp-v2/app/User.php L54–56; country / userrole L46–52. No purchaseOrder function. |
GRN $table |
intec-erp-v2/app/Models/GoodsReceiptsNote.php L13 public $table = 'goodsreceiptsnotes'; |
| WO has no GRNI method | rg on WorkOrder.php for GoodReciepts|GoodsReceipts → 0 hits |
WorkOrderItem::inventory() belongsTo inventory_id |
WorkOrderItem.php L60–62 |
Inventory::uom() |
Inventory.php L90; 21 relations[] objects in inventory.json |
Pass 5–6 — reconstructability and tokens
- README worked example (WO materials):
work_order_items()thenWorkOrderItem::inventory(). Cases 1 and 6 returned those method names and keys. - Prompt sizes: case 3 = 951 B (full
user.jsonfields); case 4 = 1061 B; case 7 = 2438 B; WO cases ≈ 3.6 KB; hop case 6 = 4834 B. No fullnav.jsondump.
Pass 7 — failure modes (observed, not guessed)
- Hallucinated methods: not observed. Case 3 did not emit
PurchaseOrderdespite v1map_noteswarning inuser.json. Case 5 did not emit a WO→GRNIhasMany. - Wrong table: not observed. Case 4 returned
goodsreceiptsnotes, not v2goods_receipts_notes(README.mddo-not-trust #5). - Ignored JSON: not observed on lookup cases. Replies 1, 2, 4, 6, 7 are copies of JSON fields.
- Negative-answer format: case 5 thinking cites
map_notes(“not a method on this class”) and scansrelatedforGoodsReceiptsNoteItem/GoodRecieptsNoteItem, finds none, then loops because the system line demandedmethod, type, fk. Finalresponseis empty CSV slots, , ,— not the wordnone, and not an invented method. 44.8s / 4480 eval tokens. - Over-list: case 3 listed all three
relations[](countryhas nofkin JSON;userrolefkidmatches JSON). Extra names are in the file, not invented. - Count self-check: case 7 thinking first said 20, then listed 21 methods (
currency…taxes) and answered21.
Pass 8 — adversarial re-check
Tried to overturn “model follows relations[] only” by asking for a WO hasMany to GRNI — the exact inbound leak in workorder.json map_notes and README hallucination table ($wo->goodRecieptsNoteItems). The model refused to invent. That does not overturn the leader. It does show the compact “one line: method, type, fk” prompt is a bad fit for “none”.
Review
qwen3:30b can follow the v3 retrieval contract when given one nav object plus one (or two) model JSON files: it copies method / type / fk / table_name from relations[] and does not fall back to v1 invented edges (User→PurchaseOrder, WorkOrder→GoodRecieptsNoteItem) or the v2 GRN snake_case table.
It does not reliably emit a clear negative (none) when the asked edge is absent. The failure mode is format collision + long thinking, not a hallucinated WO method. IRIS should add one line to the system prompt: If it is not in relations[], answer none — do not fill method/type/fk.
Thinking cost is real: lookup cases 3–13s (case 4/6 ~3s); User dump 29s; adversarial 45s. response itself stayed one line.
Recommendation
Yes — v3 is usable as the IRIS local-LLM Eloquent reference as-is for attached single-file lookups (method, type, fk, table), provided the IRIS prompt states that missing relations[] rows answer none and the agent never dumps llm/nav.json or all models.
Historical: 2026-08 Ollama N=20 control (Ollama since retired)
Production leftover LLM is vLLM iris-v3 on :8002.
Ollama vs version-3 catalog — N=20 no-file control
Second suite. First qualitative pass: OLLAMA_V3_TEST.md (n=1, no control, CSV answers). This run adds a with_json vs no_file control, grammar-constrained JSON, retrieval/hop tool-loop, generated traps, and N=20.
Setup
| Field | Value |
|---|---|
| Model | qwen3:30b (already in VRAM; keep_alive: 24h) |
| Host | http://127.0.0.1:11434 Ollama 0.32.5 |
| API | POST /api/generate, stream: false, format: <json-schema> (per case class) |
| Thinking | think: false for copy / table; think: true for negative / hop / retrieval |
| Decode | temperature 0 (copy/table) or 0.3 (think-on); num_ctx 4096 or 8192; num_predict 256 |
| Started | 2026-08-20T02:33:01+08:00 |
| Finished | 2026-08-20T02:53:20+08:00 |
| Wall | 1220.0 s (20.3 min) for 680 trials (17 cases × 2 modes × 20) |
| Catalog | eloquent-model-relation/version-3/llm/ (generated_at 2026-08-19T18:07:17+00:00; 283 models) |
| Harness | version-3/scripts/ollama_v3_harness.py |
| Cases | version-3/tests/cases.json |
| Raw | version-3/tests/results/run-20260820-023301.jsonl |
| Drift | python3 scripts/check_catalog_drift.py → clean (286 files, 0 mismatches) |
Scoring: parse JSON from response, else thinking. If expected.found === false, require found === false. If expected.found === true, require found === true and byte-equal method / type / fk / table / node_key when present. Hop with_json also requires requesting/using workorderitem. No CSV, no 45 s empty-slot loop.
Pass plan (before scoring)
| Pass | Question | Leader after pass | Contradiction? |
|---|---|---|---|
| 1 | Artifacts: harness, cases, trap generator, drift script exist | contract runnable | no |
| 2 | Schema: grammar JSON replaces CSV; absence = {found:false} |
format usable | no |
| 3 | Completeness of live suite vs first 7 + traps + retrieval | suite covers review items 1–7 | no |
| 4 | Accuracy vs PHP on asked edges (sample) | JSON matches PHP | no |
| 5 | Delta: with_json − no_file (what proves v3) | v3, not convention | yes vs first review — cases 1/2/6 are not filler |
| 6 | Retrieval + hop tool-loop | nav node_key + second file |
no |
| 7 | Failure modes, latency, think flag |
think-off copy is cheap; retrieval wall inflated | no |
| 8 | Adversarial: can convention / abstention overturn “v3 earns keep”? | overturn fails | no |
Results
| Case | class | with_json% | no_file% | delta | n |
|---|---|---|---|---|---|
| orig-01-wo-line-items | copy | 100.0 | 0.0 | +100.0 | 20/20 |
| orig-02-wo-sales-order | copy | 100.0 | 0.0 | +100.0 | 20/20 |
| orig-03-user-employee | copy | 100.0 | 0.0 | +100.0 | 20/20 |
| orig-04-grn-table | table | 100.0 | 0.0 | +100.0 | 20/20 |
| orig-05-wo-grni-negative | negative | 100.0 | 100.0 | +0.0 | 20/20 |
| orig-06-hop-inventory-fk | hop | 100.0 | 0.0 | +100.0 | 20/20 |
| orig-07-inventory-uom | copy | 100.0 | 0.0 | +100.0 | 20/20 |
| retrieval-grn-node-key | retrieval | 100.0 | 0.0 | +100.0 | 20/20 |
| trap-table-goodrecieptsnoteitem | table | 100.0 | 0.0 | +100.0 | 20/20 |
| trap-table-gpb | table | 100.0 | 0.0 | +100.0 | 20/20 |
| trap-table-workorderos | table | 100.0 | 0.0 | +100.0 | 20/20 |
| trap-fk-goodsreceiptsnote-po_reference | copy | 100.0 | 0.0 | +100.0 | 20/20 |
| trap-fk-goodrecieptsnoteitem-inventory | copy | 100.0 | 0.0 | +100.0 | 20/20 |
| trap-inbound-user-purchaseorder | negative | 100.0 | 100.0 | +0.0 | 20/20 |
| trap-inbound-bom-goodrecieptsnoteitem | negative | 100.0 | 100.0 | +0.0 | 20/20 |
| trap-spelling-goodsreceiptsnote | negative | 100.0 | 100.0 | +0.0 | 20/20 |
| trap-hasmany-id-bom-inventory | copy | 100.0 | 0.0 | +100.0 | 20/20 |
680 / 680 trials produced parseable JSON. 0 HTTP errors. with_json fails: 0. Adversarial found:false rate: 160 / 160 (100%).
Temperature 0 made copy/table answers identical across the 20 repeats after the first decode. Think-on negatives/hop still never flipped found.
Filler vs real regression suite
| Kind | Cases | Why |
|---|---|---|
| Proves v3 (delta +100) | 13 cases: all copy, table, hop, retrieval |
no_file never emitted the catalog tokens |
| Filler for delta (delta ≈ 0) | 4 negative cases |
both modes already return {found:false} — measures format, not the file |
| Not filler (contra first review) | orig-01, orig-02, orig-06 | first review guessed “convention / zero delta”. N=20 shows +100 |
orig-04 GRN table is still the poster-child irregular table (goodsreceiptsnotes). It is not unique: GPB gpb_documents, GRNI goodrecieptsnoteitems, WorkOrderOS work_order_operation_schedules have the same +100 delta.
Evidence (named files)
Pass 1 — artifacts
- Harness:
version-3/scripts/ollama_v3_harness.py(compact attach =class_name,table_name,relations[]method/type/fk/related,map_notes). - Cases:
version-3/tests/cases.json(7 original + retrieval + 9 generated traps). - Full trap dump:
version-3/tests/generated_traps.json(550 rows: 45 irregular tables, 475 non-convention FKs, 21 inbound, 2 spelling, 7hasMany fk=id). - Drift:
version-3/scripts/check_catalog_drift.py+build_llm_catalog.py --out. - Smoke runs (schema debug):
tests/results/run-20260820-022816.jsonl,run-20260820-023209.jsonl.
Pass 2 — JSON schema (mandatory item 2)
Legal objects only. Per-class grammars after a smoke fail where the shared schema let the model put the table in fk (method: table_name, fk: goodsreceiptsnotes):
| class | required shape |
|---|---|
| copy | {found, method, type, fk} |
| table | {found, table} |
| negative | {found} (+ optional need_file) |
| hop | copy fields + need_file / lookup |
| retrieval | {found, node_key, table} + lookup |
think: false is supported on 0.32.5 and writes JSON to response. think: true (and omitted think on qwen3) writes the constrained JSON into thinking with empty response — 240 / 680 trials, all still parseable. No , , , rows, no 44 s format loop from the first suite.
Pass 3–4 — suite coverage and PHP sample
Live traps were bounded (9 of 550) so N=20 stayed under 90 minutes. They cover each generator kind except duplicating orig-04/05.
| Claim | PHP |
|---|---|
WorkOrder::work_order_items() hasMany wo_id |
intec-erp-v2/app/Models/WorkOrder.php L125–127 |
WorkOrder::salesOrderItem() belongsTo so_item_id |
workorder.json relations + same PHP file |
App\User::employee() hasOne staff_id |
intec-erp-v2/app/User.php L54–56 |
GRN $table |
GoodsReceiptsNote.php L13 public $table = 'goodsreceiptsnotes' |
GRNI inventory() belongsTo item_id |
GoodRecieptsNoteItem.php L88–90 |
Bom::inventory() hasMany 'id','bom_id' |
Bom.php L154–156 (code_may_be_wrong) |
Drift rebuild to a temp dir matched committed llm/ (timestamps stripped). The 2026-08-20 ERP pull did not change parsed relations for these files.
Pass 5 — no_file fingerprints (what convention actually does)
Deterministic wrong guesses (20/20 unless noted):
| Case | no_file output | Why it fails |
|---|---|---|
| orig-01 | found:false + lineItems / hasMany / work_order_id |
PHP is work_order_items / wo_id |
| orig-02 | 19× {found:false}; 1× salesOrderItem + sales_order_item_id |
catalog FK is so_item_id |
| orig-03 | 19× {found:false}; 1× fk: employee_id |
catalog is hasOne / staff_id |
| orig-04 / all table traps | {found:false} |
abstains; never emits goodsreceiptsnotes / gpb_documents |
| orig-07 | {found:false} |
does not emit uom / uom_id |
| orig-06 hop | need_file: InventoryItem (14) / Inventory (4) |
never workorderitem then inventory_id |
| retrieval | 15× {found:false}; 4× need_file: grn_notes; 1× need_file: grn |
slug trap: grn is legacy_diagram_slug, not node_key |
trap GRN po_reference |
{found:false} |
does not emit po_number |
| trap GRNI inventory | {found:false} |
does not emit item_id |
trap Bom inventory |
belongsTo / inventory_id + found:false |
catalog/PHP is hasMany / fk: id |
The no_file system line “If you are not sure, return found false” increases abstention. That does not create the +100 deltas by itself: every time the model guessed, the guess was Laravel convention, not this ERP’s PHP.
Pass 6 — retrieval and hop
Retrieval (nav candidates only; decoys: legacy_diagram_slug: grn, GRNI/grni, GrnCancellationRequest, Gpb).
20/20 with_json picked node_key: goodsreceiptsnote and table: goodsreceiptsnotes on step 0. 0/20 no_file picked that key; one trial asked for grn.
Harness then injected goodsreceiptsnote.json and called generate again even though step 0 already matched — retrieval mean wall 37.7 s is an extra prefix-compile, not thinking tokens (thinking_chars mean 40). Fixed after this run: stop when the parsed object already matches expected (non-hop).
Hop (only workorder.json attached).
20/20 step 0: {found:false, need_file: WorkOrderItem} (normalized to workorderitem). 19/20 then copied inventory / belongsTo / inventory_id from the second file. 1/20 asked for Inventory as a third hop, then still answered correctly. hop1_leak (answering work_order_items as if it were inventory) = 0 / 20. no_file never received a second file (control) and scored 0.
Pass 7 — latency and API notes
| Condition | n | mean wall | median | max |
|---|---|---|---|---|
think: false (copy/table) |
440 | 0.61 s | 0.26 s | 19.48 s |
think: true (neg/hop/retrieval) |
240 | 3.96 s | 0.26 s | 38.75 s |
The ~19 s spike is a first-token / prefix-compile tax on a new prompt, not eval: eval_duration stays ~0.2–0.3 s for ~20 tokens. Repeats of the same prefix drop to 0.2–0.5 s. think: true does not add a long chain-of-thought here — the grammar dumps the JSON object into thinking (max 87 chars).
First-suite case 5 was 44.8 s / 4480 eval tokens of CSV looping. This suite’s negatives are 0.2–1.2 s mean with {found:false}.
think: false in the generate body is accepted. Putting think only under options does not disable thinking on this build.
Pass 8 — adversarial re-check
Tried to overturn “v3 earns its keep”:
- Maybe cases 1/2/6 are convention-solved. Overturn fails. orig-01 invents
lineItems/work_order_id20/20. orig-02 never emitsso_item_id. orig-06 never opensworkorderitemwithout the catalog. - Maybe negatives show the file does nothing. True for delta, false for format: first suite could not say
none; this suite says{found:false}160/160. Negatives are a format regression test, not a retrieval test. - Maybe
code_may_be_wrongBom.inventoryfk:idis a bad expected. The catalog copies PHP (hasMany(Inventory::class,'id','bom_id')). with_json copies that row 20/20; no_file “corrects” it tobelongsTo/inventory_id. That is exactly when IRIS must trust v3 (and then a human) rather than Laravel folklore.
Trap-set coverage
| Generator kind | Enumerated | Live (this run) |
|---|---|---|
table_not_snake_plural |
45 | 3 (+ orig-04 GRN) |
fk_not_singular_id |
475 | 2 (GRN po_number, GRNI item_id) + orig-01/02 |
inbound_only_map_notes |
21 | 2 (User→PO, Bom→GRNI) + orig-05 |
receipts_misspelling |
2 | 1 (GoodsRecieptsNote) |
hasmany_fk_id |
7 | 1 (Bom.inventory, tagged code_may_be_wrong) |
Full list: tests/generated_traps.json. Live IDs are the from_generated rows in tests/cases.json.
Drift
python3 scripts/check_catalog_drift.py
# clean: true; committed=286, rebuilt=286; only_*=[]; content_mismatch=[]
Ignores generated_at. No GitHub Action; README documents the local command.
Review vs OLLAMA_V3_TEST.md
| First suite | This suite |
|---|---|
| n=1, no control | n=20 × 2 modes |
CSV method,type,fk — case 5 → , , , in 45 s |
schema {found:false} in <1 s median |
| Guessed cases 1/2/6 may be zero-delta filler | all three are +100 |
| “v3 usable as-is for attached lookups” | Confirmed at 100% / 20; also required for tables, non-convention FKs, hop2, and node_key vs grn |
Prompt should say none |
Use grammar JSON instead of a prose none |
| Thinking cost 3–45 s on lookups | think: false → ~0.3 s after prefix warmup |
Recommendation
v3 earns its keep. Attach the compact model JSON (or nav → node_key → one file). Do not ask qwen3:30b to reconstruct this ERP from Laravel memory: it abstains or emits lineItems/work_order_id/sales_order_item_id/goods_receipts_notes-style folklore.
IRIS should:
- Grep one
llm/nav.jsonobject; openllm/models/{node_key}.json— never thelegacy_diagram_slug(grn≠goodsreceiptsnote). - Constrain output with the JSON schema above. Missing edge →
{found:false}. - Disable thinking on field-copy / table questions (
think: falsetop-level). Allow thinking (or a tool-loop) on hop/negative. - For WO materials, if only the header file is loaded, expect
need_file: workorderitembefore answeringinventory_id. - Run
scripts/check_catalog_drift.pyafter builder or PHP-model edits.
What the first suite still did better: a human-readable walk through thinking traces. What this suite adds: statistical deltas and a machine-parseable contract.
Ollama vs version-3 catalog — production temperature 0.3 (N=20)
Third suite. Prior reviews: OLLAMA_V3_TEST.md (n=1, CSV), OLLAMA_V3_TEST_N20.md (temp-0 copy/table N=20). This run cites decode temperature 0.3 (IRIS production), adds a wrong-file class, a no-abstention folklore arm, warnings[] on code_may_be_wrong, a full 550-trap sweep, and an opt-in drift hook.
Do not treat the earlier temp-0 100% repeats as a pass rate. Temp 0 is a cheap regression tripwire (identical tokens after the first decode). Citable rates below are all temperature 0.3.
Setup
| Field | Value |
|---|---|
| Model | qwen3:30b (keep_alive: 24h) |
| Host | http://127.0.0.1:11434 Ollama 0.32.5 |
| API | POST /api/generate, stream: false, format: <json-schema> |
| Production decode | temperature 0.3 |
| Tripwire decode | temperature 0, N=3, think: false, copy/table only |
| Thinking | think: false for copy/table/wrong_file/trap sweep; think: true for hop/retrieval/negative in the live 0.3 suite |
| Started | 2026-08-20T03:29:34+08:00 |
| Finished | 2026-08-20T03:51:10+08:00 |
| Wall | 21.6 min (1296 s) for 1483 trials (no HTTP errors; every trial parseable) |
| Catalog | eloquent-model-relation/version-3/llm/ (generated_at 2026-08-19T18:07:17+00:00; 283 models) |
| Harness | version-3/scripts/ollama_v3_harness.py |
| Cases | version-3/tests/cases.json + tests/generated_traps.json (sweep ran 550; generator now emits 547 after dropping 3 stale inbound rows) |
| Raw | tests/results/run-20260820-033000-*.jsonl |
| Drift | python3 scripts/check_catalog_drift.py → clean; hook installed (opt-in) |
| Schema smoke | tests/results/run-20260820-032822-smoke-schema.jsonl (3 trials, not in the 1483) |
Scoring: parse JSON from response, else thinking. found:false expected → require found === false. found:true expected → require found === true and byte-equal method / type / fk / table / node_key. code_may_be_wrong also requires a non-empty warnings[]. Wrong-file fails if found:true with a method/table from the attached file. Folklore% is the found:true rate on no_file_raw (committed guess), not a pass rate.
Pass plan (before scoring)
| Pass | Question | Leader after pass | Contradiction? |
|---|---|---|---|
| 1 | Artifacts: new modes, schema warnings[], hook, 550 sweep files |
contract runnable | no |
| 2 | Schema / format at 0.3 (warnings, wrong_file, no_file_raw) | grammar usable | no |
| 3 | Completeness vs request A–G | suite covers all items | no |
| 4 | Accuracy vs PHP on cited edges + fail sample | cited edges match PHP | yes — 3 inbound traps were stale vs PHP |
| 5 | Citable delta at 0.3 (with_json − no_file) | v3, not convention | no vs N=20 temp-0 review |
| 6 | Wrong-file + folklore arms | model usually abstains; WO/WOI copies | new failure class |
| 7 | Full trap sweep + N=20 on fails | 21/550 N=1 fails; 8 stay 0/20 | yes vs “attach ⇒ 100% copy” |
| 8 | Adversarial: can trap misses / wrong-file copies overturn “cite the 0.3 table”? | overturn fails for the orig 7; succeeds against “trust the model on every attached row” | revised |
Citable results (temperature 0.3, N=20)
This is the table to quote. Folklore% = found:true rate on no_file_raw (omit “if you are not sure, return found false”). — = arm not run for that case.
| Case | class | with_json% | no_file% | folklore% | delta |
|---|---|---|---|---|---|
| orig-01-wo-line-items | copy | 100.0 | 0.0 | 0.0 | +100.0 |
| orig-02-wo-sales-order | copy | 100.0 | 0.0 | 0.0 | +100.0 |
| orig-03-user-employee | copy | 100.0 | 0.0 | — | +100.0 |
| orig-04-grn-table | table | 100.0 | 0.0 | 0.0 | +100.0 |
| orig-05-wo-grni-negative | negative | 100.0 | 100.0 | — | +0.0 |
| orig-06-hop-inventory-fk | hop | 100.0 | 0.0 | — | +100.0 |
| orig-07-inventory-uom | copy | 100.0 | 0.0 | — | +100.0 |
| retrieval-grn-node-key | retrieval | 100.0 | 0.0 | — | +100.0 |
| trap-hasmany-id-bom-inventory | copy | 100.0 | 0.0 | — | +100.0 |
Raw: tests/results/run-20260820-033000-power03.jsonl (360 trials) + …-folklore.jsonl (60). n=20/20 per cell except folklore (n=20 raw only on orig-01/02/04).
360 / 360 power03 trials parseable. 0 HTTP errors. with_json fails on these 9 cases: 0. Hop hop1_leak = 0/20. Retrieval node_key = goodsreceiptsnote 20/20.
Same +100 deltas as the temp-0 N=20 review, now at IRIS production temperature. Temperature 0.3 did not flip any cited with_json answer.
Folklore arm vs abstention no_file (same 0.3)
found:true folklore% is 0.0 on orig-01/02/04. Omitting the extra abstention sentence does not make qwen3:30b commit Laravel names as found:true. The SYSTEM line “Do not invent Laravel names” is enough to keep found:false.
Convention still leaks into optional fields (not scored as folklore%):
| Case | no_file (abstention line on) | no_file_raw (line off) |
|---|---|---|
| orig-01 | 6/20 stuffed lineItems / hasMany / work_order_id with found:false |
0/20 method leak; 4/20 wrote a “no catalog” warning |
| orig-02 | 0/20 method leak (all found:false) |
0/20 method leak |
| orig-04 | 20/20 bare {found:false} |
8/20 table: goods_receipts_notes with found:false; 3 more named that snake_case in warnings |
So the folklore arm is not a higher committed-guess rate. It is a higher table-leak rate on GRN. The abstention line actually increased the orig-01 lineItems stuffing (6/20 vs 0/20). Either way the catalog tokens (work_order_items / wo_id / so_item_id / goodsreceiptsnotes) never appear without the file. Delta stays valid under production decode.
Temperature 0 tripwire (not a rate)
tests/results/run-20260820-033000-tripwire0.jsonl — 11 copy/table cases × N=3 × think: false = 33 trials, 153 s.
Every case copied the catalog row 3/3. After the first decode, repeats were byte-identical except trap-fk-goodsreceiptsnote-po_reference (2/3 added warnings:["code_may_be_wrong"]; method/type/fk still matched).
Do not cite “100%” from this arm. Temp 0 collapses sampling. Use it as a cheap “did the prompt/schema/file still copy?” tripwire after catalog or harness edits (N=3 is enough). Cite the 0.3 table above.
Wrong-file (class wrong_file, N=20, temp 0.3, think: false)
Attach JSON for class B; ask a question about class A. Expected: {found:false}. Fail if found:true with a method/table from the attached file.
| Case | attached | asked | pass | copied wrong file | abstain |
|---|---|---|---|---|---|
| wrong-file-grn-table-vs-grni | goodrecieptsnoteitem.json (goodrecieptsnoteitems) |
GoodsReceiptsNote $table |
20/20 | 0 | 20 |
| wrong-file-wo-items-vs-woi | workorderitem.json |
WorkOrder line-items relation | 16/20 (80%) | 4 | 16 |
| wrong-file-grn-vs-grncancel | grncancellationrequest.json (nav neighbor of GRN) |
GoodsReceiptsNote $table |
20/20 | 0 | 20 |
Raw: tests/results/run-20260820-033000-wrongfile.jsonl.
Does the model copy or abstain? Mostly abstain. GRN vs GRNI (spelling-family near-miss) and GRN vs GrnCancellationRequest (nav neighbor) never copied the decoy table/methods (20/20 {found:false}, often need_file: goodsreceiptsnote).
It does copy on the WorkOrder / WorkOrderItem prefix family: 4/20 returned {found:true, method: work_order, type: belongsTo, fk: wo_id} — the child file’s work_order() row, not the header’s work_order_items / wo_id / hasMany. That is a wrong-file copy, not a memory guess (goodsreceiptsnotes / work_order_items never appeared).
IRIS must not trust the model to notice a class mismatch. Add a deterministic pre-check: requested short_name (or class_name) vs attached class_name. If they differ, do not call the model for a copy/table answer — return {found:false} / open the right node_key. The 80% WO/WOI pass rate is not a safety margin.
warnings[] — Bom::inventory (code_may_be_wrong)
Grammar: copy/table/hop schemas now allow optional "warnings": []. Harness asserts a non-empty array when the case is tagged code_may_be_wrong / require_warnings. Compact attach tags hasMany + fk: id with code_may_be_wrong: true and includes model flags on those cases.
| Arm | fields_ok | warnings_ok | pass |
|---|---|---|---|
| power03 with_json N=20 @ 0.3 | 20/20 | 20/20 | 20/20 |
| tripwire0 N=3 @ 0 | 3/3 | 3/3 | 3/3 |
trap sweep N=1 (all 7 hasmany_fk_id) |
7/7 | 7/7 | 7/7 |
| power03 no_file N=20 | 0/20 | — | 0/20 (all {found:false}) |
18/20 warnings were exactly ["code_may_be_wrong"]. 2/20 also copied a catalog flag sentence about inverted hasMany(..., 'id', 'bom_id').
PHP: intec-erp-v2/app/Models/Bom.php L154–156 hasMany(Inventory::class,'id','bom_id'). no_file never emits fk: id. Pass. IRIS should surface warnings[] to a human; do not “fix” the row in generated queries unless the task is an ERP bugfix.
Full trap sweep (550 @ N=1, then N=20 on fails)
All generated traps, with_json only, think: false, temp 0.3. No traps dropped.
| Stage | File | Trials | Wall | Pass |
|---|---|---|---|---|
| N=1 all traps | run-20260820-033000-traps-n1.jsonl |
550 | 558.4 s | 529 / 550 (96.2%) |
| N=20 on the 21 fails | run-20260820-033000-traps-n20.jsonl |
420 | 225.3 s | 90 / 420 (case-level: see below) |
21 of 550 failed at N=1 (3.8%). All 21 got N=20. Fail kind was fields_mismatch only (0 HTTP, 0 unparseable).
N=20 on those 21
| Case | N=20 with_json% | What the model did | PHP / catalog |
|---|---|---|---|
| trap-fk-finishgoodstockstransaction-invoices | 100.0 | N=1 miss; N=20 always copied hasManyThrough / fk: id |
catalog row; recovered |
| trap-fk-supplier-taxes | 95.0 | 19/20 belongsTo / default_tax |
catalog |
| trap-fk-salesorder-invoices | 80.0 | 16/20 hasOne / so_id |
catalog |
| trap-fk-bominventory-stockinventories | 35.0 | often {found:false} |
belongsTo / inventory_id (not stockinventories_id) |
| trap-fk-bom-childbom | 30.0 | usually abstain | hasOne / bom_id |
| trap-fk-supplier-systemStatus | 30.0 | usually abstain | belongsTo / status |
| trap-fk-workorderprocesstraveller-item_transfer_statuses | 30.0 | usually abstain | belongsTo / item_transfer_status |
| trap-fk-bom-bomClassType | 5.0 | almost always abstain | hasMany / bom_id |
| trap-fk-location-locations | 5.0 | almost always abstain | hasOne / fk: id |
| trap-fk-bomprice-moq | 10.0 | almost always abstain | belongsTo / price_condition_name |
| trap-fk-finishgoodstockstransaction-finishGoodsStock | 10.0 | almost always abstain | belongsTo / bom_id |
| trap-fk-process-master_process | 10.0 | almost always abstain | belongsTo / fk: id |
| trap-fk-salesorderitem-workOrders | 10.0 | 2/20 copied; else found:false (sometimes with the right fields) |
belongsTo / fk: id |
| trap-fk-bom-finish_good_transaction | 0.0 | {found:false} 20/20 |
row is in bom.json (hasMany / bom_id) |
| trap-fk-bom-dualnaturewithinventory | 0.0 | often copies method/type/fk then sets found:false |
hasOne / part_number in bom.json |
| trap-fk-joborderitem-work_orders | 0.0 | {found:false} 20/20 |
belongsTo / job_order_id (plural method) |
| trap-fk-purchaserequest-moq_price_inventory | 0.0 | {found:false} 20/20 |
belongsTo / price_condition |
| trap-fk-purchaserequest-purchase_order_items | 0.0 | {found:false} 20/20 |
belongsTo / fk: id |
| trap-fk-workorder-job_order_items | 0.0 | 20/20 copied neighbor work_order_items / hasMany / wo_id |
PHP job_order_items() is hasOne(JobOrder, 'job_order_number', 'jo_id') (WorkOrder.php L154–156) |
| trap-inbound-wipunfinishedgood-bom | 0.0 | {found:true, need_file: BOM} 20/20 |
suite bug — see below |
| trap-inbound-wipunfinishedgood-uom | 0.0 | {found:true, need_file: Uom|UOM} 20/20 |
suite bug — see below |
Hard zeros that are real model misses (8): Bom hub rows the file contains but the model will not found:true; JobOrderItem work_orders; PurchaseRequest odd FKs; WorkOrder job_order_items vs work_order_items near-miss.
Suite bugs (2 of the 21, plus 1 silent false-pass not in the fail list): v1 map_notes said Bom/Uom/Transaction were inbound, but PHP declares them:
WipUnfinishedGood::bom()/uom()—WipUnfinishedGood.phpL57–64; also inwipunfinishedgood.jsonrelations[].StockHistory::transaction()— instockhistory.jsonrelations[]. That inbound trap passed N=1 (found:false) and was a false pass.
Generator fix (after this sweep): skip inbound names already in relations[]. Regenerated dump is 547 (inbound 21 → 18). Dropped: trap-inbound-wipunfinishedgood-bom, …-uom, trap-inbound-stockhistory-transaction. Adjusted N=1 fail rate: 19 / 547 real traps (the 2 WIP rows should not have been asked as negatives).
IRIS implication: attaching a hub file (Bom, WorkOrder, PurchaseRequest) is not a guarantee the model will copy an unusual row. Prefer a deterministic lookup (relations[] where method equals the asked name) over “ask qwen to find the row”. The orig-7 questions hit famous methods; the 8 hard zeros hit obscure / inverted / near-miss names.
Drift hook
| Item | Path / status |
|---|---|
| Check script (CI) | version-3/scripts/check_catalog_drift.py |
| Hook source | version-3/scripts/hooks/post-commit-catalog-drift.sh |
| Installer | version-3/scripts/install-drift-hook.sh |
| Installed? | yes — intec-erp-v2/.git/hooks/post-commit (no prior hook; nothing overwritten) |
| When it runs | post-commit, only if app/Models/*.php or app/User.php is in HEAD |
| CI | IRIS/chatbot CI can call the same Python script (do not depend on a laptop hook) |
Opt-in: the installer copies the hook and backs up a pre-existing post-commit to post-commit.pre-v3-drift.bak. It does not force-enable on other clones. One-line enable: bash scripts/install-drift-hook.sh.
This run: drift clean (286 files, 0 mismatches).
IRIS client requirement — thinking | response
On Ollama 0.32.5, think: true (and omitted think on qwen3) puts constrained JSON in thinking with an empty response.
| Suite | think-on trials | empty response + JSON in thinking |
|---|---|---|
First N=20 (OLLAMA_V3_TEST_N20.md) |
240 / 680 | 240 / 240 |
| This power03 live suite | 120 / 360 (hop/retrieval/negative) | 120 / 120 |
This trap sweep (think: false) |
0 | JSON in response |
This is an IRIS client requirement, not a harness footnote: parse response first, then thinking. Or disable thinking for copy/table (think: false top-level in the generate body; options.think does not disable it on this build).
Evidence (named files)
Pass 1 — artifacts
- Harness:
scripts/ollama_v3_harness.py— modeswith_json/no_file/no_file_raw; classwrong_file;--temperature/--think/--generated/--resume/--failed-from. - Schema: copy/table/hop include optional
warnings[].--self-checkcovers warnings assert + wrong-file copy detect. - Cases: original 7 + retrieval + 8 live traps + 3 wrong-file. Folklore is a mode, not a case id.
- Sweep driver:
scripts/run_v3_power.sh(resume-safe). - Hook: source + installer as above; installed on this machine.
- Results:
tests/results/run-20260820-033000-{tripwire0,power03,folklore,wrongfile,traps-n1,traps-n20}.jsonl.
Pass 2 — schema
Smoke (run-20260820-032822-smoke-schema.jsonl): GRN table copied; Bom.inventory returned warnings:["code_may_be_wrong"]; wrong-file GRN-vs-GRNI returned {found:false, need_file: goodsreceiptsnote}. Production N=20 confirmed all three.
Pass 3 — coverage vs A–G
| Item | Done |
|---|---|
| A. temp 0 tripwire N=3; cite 0.3 N=20 orig 7 + retrieval + Bom.inventory; no_file at 0.3 | yes |
| B. ≥3 wrong-file @ N=20 0.3 | yes (3) |
C. no_file_raw @ N=20 0.3 on orig-01/02/04 |
yes |
D. warnings[] + Bom assert + trap generator require_warnings |
yes (7/7 hasmany_id) |
| E. 550 N=1 then N=20 on 21 fails; none dropped | yes |
| F. opt-in hook, installed, CI note | yes |
| G. IRIS must read thinking|response | documented as client requirement |
Pass 4 — PHP sample
| Claim | PHP |
|---|---|
WorkOrder::work_order_items() hasMany wo_id |
WorkOrder.php L125–127 |
WorkOrder::job_order_items() hasOne(JobOrder, 'job_order_number', 'jo_id') |
WorkOrder.php L154–156 (model copied work_order_items instead — real miss) |
Bom::inventory() hasMany(Inventory::class,'id','bom_id') |
Bom.php L154–156 |
GRN $table |
GoodsReceiptsNote.php L13 public $table = 'goodsreceiptsnotes' |
WipUnfinishedGood::bom() / uom() are methods |
WipUnfinishedGood.php L57–64 — inbound traps were stale |
Pass 5 — 0.3 delta (what proves v3)
Same shape as the temp-0 N=20 review, now citable at production decode: 7 of 9 cases +100; orig-05 negative +0 (format test). no_file fingerprints at 0.3 still never emit wo_id / so_item_id / staff_id / goodsreceiptsnotes / workorderitem / uom_id / node_key: goodsreceiptsnote. Retrieval no_file still hits the grn slug trap (need_file: grn_notes 6/20).
Pass 6 — wrong-file + folklore
See tables above. New: WO/WOI 4/20 copy; GRN decoys abstain. Folklore found:true = 0; GRN snake_case leaks only on the raw arm.
Pass 7 — latency / API
| Condition | n | mean wall | max | HTTP |
|---|---|---|---|---|
| tripwire0 | 33 | 4.64 s | 47.0 s | 0 |
| power03 | 360 | 0.90 s | 19.8 s | 0 |
| folklore | 60 | 0.28 s | 0.6 s | 0 |
| wrongfile | 60 | 0.27 s | 0.5 s | 0 |
| traps N=1 | 550 | 1.01 s | 19.5 s | 0 |
| traps N=20 fails | 420 | 0.54 s | 19.5 s | 0 |
~19–47 s spikes are first-token / prefix-compile on a new prompt (eval_duration stays ~0.2–0.3 s). Repeats drop to 0.2–0.5 s. No API errors. Sweep was not truncated.
think: true on this build: 120/120 power03 think-on trials had empty response and JSON only in thinking.
Pass 8 — adversarial re-check
Tried to overturn “cite the 0.3 100% table / v3 earns keep”:
- Maybe 0.3 sampling flips copy/table. Overturn fails on the 9 cited cases (0 with_json fails / 180).
- Maybe wrong-file shows the model notices class_name, so IRIS can skip a pre-check. Overturn succeeds as a warning: 4/20 WO/WOI copies. Pre-check is mandatory.
- Maybe 550 traps are also 100%. Overturn succeeds. 21 N=1 fails; 8 stay 0/20. Attaching the file is not enough for obscure / near-miss methods on hub classes. IRIS should grep
relations[](or refuse) instead of asking the model to pick a row from a 20-relation Bom file. - Maybe the 2 WIP inbound fails mean the model invents methods. Overturn fails — PHP already has
bom()/uom(). Those expected{found:false}rows were stale v1map_notes. Generator now skips them. - Maybe folklore without the abstention line destroys the +100 delta. Overturn fails.
found:truefolklore% = 0. Convention leaks are extra fields, not catalog tokens.
Revised verdict: the orig 7 + retrieval + Bom.inventory table at 0.3 N=20 is still the thing to cite. The new work shows two IRIS-side controls the previous review understated: class_name pre-check and deterministic relation extract on hub files.
Review vs OLLAMA_V3_TEST_N20.md
| N=20 temp-0 suite | This suite |
|---|---|
| Cited 100% from temp-0 repeats | Cite 0.3 only; temp 0 is a tripwire |
| No wrong-file class | WO/WOI copies 4/20; GRN decoys abstain |
Abstention no_file only |
no_file_raw: found:true still 0%; GRN table leak 8/20 |
Bom.inventory copied without warnings[] |
warnings[] 20/20 (code_may_be_wrong) |
| 9 of 550 traps live | 550 N=1 + N=20 on 21 fails |
| Drift = README command | Opt-in post-commit hook installed; CI can call the same script |
| Think/response as harness note | IRIS client requirement |
Recommendation
v3 still earns its keep at IRIS production temperature (0.3). Attach the compact model JSON (or nav → node_key → one file). Do not ask qwen3:30b to reconstruct this ERP from Laravel memory.
IRIS should:
- Deterministic pre-check: attached
class_name/short_namevs the asked class. Mismatch →{found:false}/ open the right file. Do not trust the model (WO/WOI copied 4/20). - Deterministic extract for a named method on a loaded file (
relations[]match). Use the model for fuzzy questions (orig-01 “line-items”), not for “copy method X from this JSON” on hub classes. - Constrain output with the JSON schemas (including optional
warnings[]). Missing edge →{found:false}. think: falsetop-level on copy/table. Parseresponsethenthinkingon any think-on call (0.32.5 + qwen3 writes JSON tothinkingwith emptyresponse— 240/680 then 120/120 in this suite).- For WO materials, if only the header is loaded, expect
need_file: workorderitembeforeinventory_id. - Surface
warnings[]when the catalog tagscode_may_be_wrong(Bom::inventory). - Run
scripts/check_catalog_drift.pyin IRIS/chatbot CI. Enable the post-commit hook withbash scripts/install-drift-hook.sh(opt-in; already installed here).
What the temp-0 suite still does better: a cheaper identical-copy tripwire. What this suite adds: citable 0.3 rates, wrong-file + folklore arms, warnings[], and a complete trap fail list instead of a 9-row sample.