Evidence for “LLM or JEV? Why not both?”
This is a documentation-only synthesis written on 2026-09-23. No training,
inference benchmark, external API request, prompt change or deployment change
was performed for the essay. It describes the evidence as of protected research
commit 0ed0c8cf8959f468ec9adaeadbe89202dedbce6c. Source studies retain their
own model/binary hashes, raw predictions, settings and measurement protocols.
The shareable article is README.md. Its figure is decision-head.svg, a self-contained 760 × 644 SVG with accessible title/description, explicit white background and no remote assets. It can be included in GitHub Markdown with a normal relative image link.
Table 1: frozen structured-input JEV comparison
- Primary source: JEV report. Detailed candidates: tables.
- JEV 1.13.0, one policy configuration frozen in
c1c3c17before responses; no task-specific fitting or development prompt search. JEV receives explicit priority criteria; local heads learned priority from labels and an action-policy representation. This is a configured-system comparison, not identical prompts. - D means a refit on the existing training pool with v0.5 development selection. The article uses the same frozen D checkpoint throughout the JEV/Laya tables; it does not select A/B/C/D separately for each task after looking at results.
- Anchor allowed-set accuracies, in order start/interruption/priority:
JEV
63/96, 74/106, 99/106; E4B D92/96, 91/106, 75/106; 12B D95/96, 77/106, 86/106. Six start anchors have multiple allowed labels. Followup views are not pooled into these headline denominators. - API latency: 407 scored requests, three excluded warmups, complete JSON body after a reused TLS connection, median 215.946568991 ms. This includes old regression suites as well as v0.5; it is not restricted to the anchor rows. Interruption requests bundle two questions; start requests ask one.
- Earlier representation captures on all 261 v0.5 views: median E4B 54.575695015955716 ms and 12B 201.11658499808982 ms. These are full-prompt GPU forwards, excluding model loading, tokenization, head inference and ASR. They are not simultaneous with the API measurement, matched service latency, or incremental ASR-prefill residual time.
- No exact D head-only or complete D decision latency is supplied in the table. Old NumPy head measurements and current conversational C timings are not silently attached to the structured D accuracy results. No ratio or sum of independently measured medians is reported.
- JEV priority details use all 126 v0.5 interruption views: 33 applicable, 93 non-applicable. JEV has zero false authority calls; E4B D has 15 and 12B D has 14. All three find 33/33 applicable. This is separate from 106 anchors.
- All 261 v0.5 observations passed strict response validation. One older-suite probability vector summed to 0.99 and failed the predeclared strict check; no primary v0.5 score changes. Raw and rounding-sensitivity results remain in the source study, rather than being silently repaired.
Table 2: corrected Laya pilot
- Primary source: corrected report.
Only the
-r2study is cited as a valid fitted-model result. The initial criteria-order/label-mapping bug and invalid run are preserved there. - Laya source
ffaae82e10ca266b30968d4289e91844f6b0fa80, studied version 0.3.7; English ModernBERT-large checkpoint at HF revision5e7b2b1b8ca2ecdd3f2322d94069c9b6ce7e844b. The web project's current version can differ from the pinned experiment; the article does not claim latest. - Corrected protocol
b1da497; selected checkpoint/evaluation code0170164before evaluation. Epoch 2 selected by development NLL; stopped after epoch 5. This was supervised full adaptation except the unused act_head, not upstream RLCD, and not an exhaustive tuning search. A single global temperature fit on calibration data changes probabilities, not the hard class ranking. - Training pool: 648 observations / 206 family IDs, consisting of 511 final and 137 interruption observations. Interruption supplies two separate judgments, giving Laya 785 fitting judgments. Local heads use corresponding task subsets.
- Anchor numerators:
33/96, 42/106, 40/106. E4B/12B D reference predictions are exactly those in Table 1. All 511 training final decisions and all 96 evaluation final anchors were predicted defer; successful training-set fitting was absent. - 407 scored warm loopback HTTP requests: median 15.9252929734 ms, p95 19.1412447 ms; three warmups excluded. In-process SDK median 13.1734139868 ms is preserved in the source but not substituted for the article's HTTP result. Host GPU: RTX PRO 4500. This includes encoder work; it does not reuse Gemma h.
- Local Gemma times are representation-only values from Table 1. “Finishes before the measured input forward” is descriptive of these runs, not a controlled equal-format architectural ablation or a measured live audio advantage.
Table 3: two explicitly separated studies
- Accuracy: v0.2 report
and machine-readable results.
Matching variants are
native-policy-e4bversuspolicy-e4b, andnative-policy-answer-12bversuspolicy-answer-12b. - Native method scores the entire action string plus end-of-turn, with one shared prefill and restored prefix KV per candidate. Argmax is the decision. It is neither greedy live tag generation nor a generated-reasoning baseline. For 12B, the existing empty thought-channel seed puts both readouts on the answer-channel boundary; no reasoning tokens are generated.
- Final dev: E4B native/head
29/43versus39/43; 12B28/43versus34/43. Interruption dev: E4B9/17versus13/17; 12B15/17versus15/17. Ten final and five interruption construction families, with correlated views. Development selected the heads. No native v0.5 comparison is implied. - Timing: later tag study,
tools
203a83a0dd6d6535df7b431921ca7a82dfe6c85e, engineb2e8a99e7357918ccfafa6936dfefafa25c58da9. At that time the conversational head used NumPy in shadow mode, so text decisions remained independent. - Per host, two tagged blocks × six fixed observations = 12 measured tags. One warmup per block excluded. Start: transport read with complete immutable hidden frame; text end: read with closing bracket of the full tag; head end: CPU inference complete, including frame/queue time. These are receiver availability measurements, not exact GPU token timestamps. MTP may coalesce tokens into the same read. Latency is not PCM onset or ASR-to-answer time.
- Union of speak/defer tag ranges and all-tag head median:
nowhere 12B
32.83–50.52 / 0.660 ms; Cortex E4B49.94–101.02 / 0.824 ms; somewhere 12B77.78–80.09 / 0.545 ms. Hardware rows have different MTP settings; do not infer pure hardware ratios. - The timing probes have no new independently established gold labels. Adding a speak tag changed some decisions relative to the native prompt. These probes cannot establish accuracy or behavior neutrality. Live startup outliers remain in the study; the fixed-probe ranges are not universal upper bounds.
- In current control mode, forced output tags express the head's selected action; they cannot serve as an independent text-decision baseline.
Figure and live-runtime description
- Conversational retraining report: 492 training observations / 144 family IDs; three-seed MLP64 ensemble. Input normalization/clipping, tanh, mean logits and temperature softmax. Backbone frozen; E4B hidden dimension 2560, 12B 3840. Training labels do not contain arbitrary authored probability targets. Scores are not established real-room calibrated confidence.
- New conversational heads on 125 eligible previously examined evaluation views:
E4B
98/125 = 78.4%, 12B106/125 = 84.8%. This subset excludes observations outside the live renderer/attribution contract and differs from the 96 anchors. Using the old structured heads on conversational features gave only 35.2% and 47.2%; format transfer must not be assumed. No fresh holdout was created for the integration study. - Resident C deployment:
engine shared-library source
29789ee, tools integration6d52b27plus portability fix736aed8; later acquaintance-prompt revision3a07d32does not change the head implementation. Python owns transport/worker/controller; C runs inference through a resident shared library, without per-call spawning. - Figure is a control/data-flow summary, not a complete Gemma architecture. Capture is the final prompt token after output norm, copied after the final full forward and before generated-token forwards. The vocabulary projection may already have run on that final prompt forward. The later vocabulary/MTP box depicts continued response generation, not a claim that prompt logits were skipped or that the control gate is a transformer layer.
- CPU inference can overlap channel setup. Answer MTP/decoding depends on the returned decision; no fully independent speculative answer path is claimed. Wait is transported distinctly but uses the silent/defer generation path; both may still allow private speaker-name metadata and do not synthesize it. Future input can trigger another observation. The head does not currently rescore every incoming voice during playback.
- Current C warmed readout medians on nowhere: E4B
0.08072663 ms, 12B0.12045247 ms, including Python/C prediction conversion. Excludes loading, hashing, worker scheduling, socket transport, GPU copy, prefill and audio. Fifty repetitions after excluded warmup per vector, interleaved with NumPy. Accuracy parity and full per-host evidence are in the deployment study. - Voice identification is upstream CAM++; head input is a listener-local textual representation containing available transcript, identities and context. The figure's input/output utterances and scores are illustrative, not a logged conversation, a claim of calibrated 70% correctness or an explicit addressee prediction from the head. Acoustic speaker identification is not the classifier.
- Priority and interruption heads are offline experiments; the current live conversational head outputs start-turn action only. Existing acoustic barge-in is retained. The documented Smith/Byers collision demonstrates remaining limits.
Scope and external references
The v0.5 anchor panel contains 202 episode IDs derived from 22 recurring generator prototypes shared across development/calibration/evaluation. Labels are same-agent synthetic annotations, not independent human consensus; uncertainty, limited topics, incomplete real-room coverage and already-examined evaluations constrain all conclusions. Episode bootstrap intervals do not fix shared-prototype leakage. No claim is made about proprietary-model training contamination either way.
- Official TypeSafe introduction: typed decisions from supplied state/questions, rather than our inference about undisclosed JEV internals. Consulted 2026-09-23. JEV is a comparison service, not a component of the classroom pipeline.
- Official Laya repository: open typed judgment implementation; the tested English ModernBERT checkpoint is pinned above. Consulted 2026-09-23. No equivalence to proprietary JEV internals claimed.
Our proof-of-concept (PoC) head failed to demonstrate the potential of this approach for reasoning budget planning. No live reasoning was enabled for this article. A future study must label the benefit/cost of extra reasoning, preserve independent evaluation, calibrate routing, and distinguish missing evidence from computation that can help resolve a problem.