research note

LLM or JEV? Why not both?

A LittleGemma research note · September 23, 2026

An assistant can know the answer and still be wrong to speak. In our experimental voice classroom, Mrs. Byers teaches, Chey asks questions, and Agent Smith joins as a guest. Each runs on a separate computer and listens through a microphone. When Smith asks Chey a question, Byers should usually remain silent. When an announcement requires everyone's attention, even a confident answer may need to stop. These are decisions about participation, not just language generation.

The three-voice classroom, recorded 2026-09-23. Agent Smith (E4B, RTX PRO 4500), Chey (E4B, Jetson Orin NX) and Mrs. Byers (12B, RTX A5000) each run on their own machine; the head-control line on each card shows the decision head's speak / defer call for that turn. Watch on YouTube.

An LLM can make those decisions through text. We previously asked it to emit a private [defer] tag when it should stay silent, and later added [speak] for measurement. That keeps decision and response in one generation, but the controller must wait for enough tokens to recognize the decision. It also ties a small control problem to the model's text-generation behavior.

TypeSafe's JEV offers a different interface: supply state and typed questions, and receive structured decisions, including choices and probability distributions. Its internal architecture is not available to us. Laya provides an open implementation of typed judgments; the English model we tested uses a ModernBERT encoder. Both suggest a useful question: if the desired result is a small decision, must we express it as generated prose?

Our experiment brings a similar decision interface into an LLM that is already processing the conversation. “Both” means combining language generation with a dedicated decision readout. The classroom remains self-contained; JEV is an external benchmark, not a runtime dependency.

A decoder LLM transforms input tokens into contextual hidden states. Its vocabulary head maps those states into next-token scores. We attach another small head to the final prompt token's hidden state after output normalization. In Gemma 4 E4B, that is a vector of 2,560 numbers. The backbone stays frozen. Three small multilayer perceptrons, each with a 64-unit hidden layer, read the vector; their averaged logits are converted into scores for speak, defer, and wait.

The expensive language representation already exists because the LLM needs it to answer. The added head learns to read that representation for a different purpose. It does not generate a rationale, rerun the transformer, or identify voices from audio. Speaker identification and transcription supply the listener's available evidence; the head uses the resulting conversational context to judge whether this listener should answer.

LittleGemma decision head: the final prompt hidden state branches to a CPU classifier, whose action controls response generation.
The figure shows E4B and the current trio control path. Scores are illustrative, not calibrated confidence. The engine gate is control logic, not an extra transformer layer. The immutable vector is copied after the final prompt forward, which may already have computed vocabulary logits, and before generated-token forwards.

The resident C head runs on the CPU and returns its result to the turn controller. The controller checks that it still belongs to the current session and turn, then selects the private response prefix. Speak permits answer decoding with the existing MTP draft-and-verify path. Defer ends this opportunity silently. Wait also stays silent, but leaves the decision open pending more evidence or utterance completion.

This avoids a second decision request and extra GPU kernels for the classifier. Some answer-channel setup can overlap CPU inference, but answer decoding still waits for the decision. We are not claiming that a complete speculative answer runs independently in parallel. The existing acoustic barge-in mechanism remains responsible for collisions after speakers start. The live head currently controls start-turn actions; the separately benchmarked interruption and authority classifiers are not yet a complete deployed replacement for that mechanism.

Authority also needs its own output. “The teacher said it” does not make every sentence an interruption. A routine comment, a quoted announcement, and an applicable live announcement have different implications. Priority should determine precedence under the policy; probability should represent uncertainty about the judgment. A high speak score cannot grant permission to override a required announcement, and authority should not be encoded by reserving a confidence score of 1.0.

What have we measured? The first two tables use the same authored v0.5 evaluation anchors: 96 start-turn observations and 106 interruption observations, with priority scored on the latter. Start actions are speak/defer/wait; interruption actions are continue/pause/yield; priority is ordinary/applicable authority/unresolved. Some observations legitimately permit multiple actions, so accuracy means membership in the allowed set. “D” denotes the previously frozen refitted local checkpoint selected using development data, not a winner selected after seeing the competitor's results.

1. JEV versus our structured-input decision heads

Measurement JEV 1.13.0, zero-shot E4B + head D 12B + head D
Start-turn accuracy 65.6% · 63/96 95.8% · 92/96 99.0% · 95/96
Interruption accuracy 69.8% · 74/106 85.8% · 91/106 72.6% · 77/106
Authority-priority accuracy 93.4% · 99/106 70.8% · 75/106 81.1% · 86/106
Median latency, different boundaries 216 ms · complete API request 54.6 ms · input forward only 201.1 ms · input forward only

JEV did substantially better on authority priority. Across all 126 interruption views, including followups, it recognized all 33 applicable-authority observations and made no false authority calls among the other 93. Both local heads also found those 33, but incorrectly promoted 15 and 14 non-authority observations respectively. This is a useful weakness to learn from.

Our heads performed better on start-turn actions in this panel. JEV tended to defer or pause too often. The small interruption advantage of the 12B head over JEV is inconclusive: its paired interval includes zero. This was one frozen JEV policy configuration without task-specific training or prompt search, not its best possible tuned performance.

The latency row is deliberately labeled: JEV's 407 scored requests include network and service time; local figures are earlier GPU input-forward measurements over 261 v0.5 observations on an RTX PRO 4500. They exclude the head and are not complete decision latencies. We did not record a matched end-to-end runtime for these exact D checkpoints. These figures cannot support a simple speedup ratio. JEV experiment.

2. Laya versus our structured-input decision heads

Measurement Laya, supervised pilot E4B + head D 12B + head D
Start-turn accuracy 34.4% · 33/96 95.8% · 92/96 99.0% · 95/96
Interruption accuracy 39.6% · 42/106 85.8% · 91/106 72.6% · 77/106
Authority-priority accuracy 37.7% · 40/106 70.8% · 75/106 81.1% · 86/106
Median latency, different boundaries 15.93 ms · complete local HTTP request 54.6 ms · input forward only 201.1 ms · input forward only

Laya used the same v0.5 training pool: 648 observations, comprising 511 start-turn and 137 interruption observations. Interruption contributes both an action and a priority label, giving Laya 785 judgments. Our separate heads use the relevant task subsets. This is the same dataset, not the same underlying model: the Laya pilot adapted a 421-million-parameter encoder, while our heads read a frozen Gemma backbone.

The table uses the corrected run after an adapter label-order bug was fixed. Even then, the selected Laya checkpoint deferred on every start-turn anchor and every start-turn training observation. The pilot had not demonstrated successful fitting; it cannot establish that Laya's architecture is incapable of this task. It used one supervised recipe rather than reproducing the project's RLCD training.

Laya's 407 warm local HTTP requests were fast, including its own text encoding, with a 19.14 ms p95. In this standalone setting, it finished before Gemma's measured full input forward. Our architectural advantage is reusing an LLM representation that the application already needs, not claiming that a large LLM is the fastest standalone classifier. Corrected Laya experiment.

3. The LLM's text/vocabulary decision versus an attached head

This comparison has two distinct experiments. Accuracy comes from an earlier v0.2 development study: the policy-instructed LLM scores complete candidate action strings through its normal vocabulary output, without generated reasoning. Timing comes from later fixed probes of generated [speak]/[defer] tags and a refitted conversational head. They must not be presented as one jointly measured accuracy-and-speed result.

Evidence / model LLM text or vocabulary path Attached decision head
v0.2 start-turn dev · E4B 67.4% · 29/43 90.7% · 39/43
v0.2 start-turn dev · 12B 65.1% · 28/43 79.1% · 34/43
v0.2 interruption dev · E4B 52.9% · 9/17 76.5% · 13/17
v0.2 interruption dev · 12B 88.2% · 15/17 88.2% · 15/17
Later timing · RTX PRO 4500 / 12B 32.83–50.52 ms · tag range 0.660 ms · head median
Later timing · Orin NX / E4B 49.94–101.02 ms · tag range 0.824 ms · head median
Later timing · RTX A5000 / 12B 77.78–80.09 ms · tag range 0.545 ms · head median

Both timing paths start at receipt of the complete hidden-state frame. The text path ends when the complete tag arrives; head timing includes frame handling and its CPU queue. Each row covers 12 tagged responses, excluding warmups. This isolates the extra time after the shared representation becomes available. It does not include prefill, measure audible response latency, or establish new accuracy. Live acoustic runs had larger delays, including startup outliers. Accuracy study, tag-timing study.

The accuracy evidence is encouraging but preliminary. The development examples also selected the heads; 43 start-turn observations cover only ten construction families, and the 17 interruption observations cover five. There were no continue targets in that interruption development set. The E4B head's 90.7% headline also hid failure on all three required-wait cases. This is not a broad demonstration of superiority over every well-prompted LLM, nor a matched v0.5 comparison against live greedy text tags.

The timing study used the earlier NumPy implementation. The current resident C implementation separately measured warmed readout medians of 0.081 ms for E4B and 0.120 ms for 12B on the 4500 host's CPU, including Python/C conversion but excluding transport, queues, GPU copies and prefill. Those numbers describe a different boundary from the table. C deployment evidence.

There is another important boundary: benchmark prompts and conversational prompts are different distributions. We preserved the structured benchmark artifacts and trained separate integration heads on 492 conversationally rendered training observations. On 125 eligible, previously examined evaluation views, those heads achieved 78.4% for E4B and 84.8% for 12B. We should not advertise the structured benchmark's 95–99% as the accuracy of the live classroom. Conversational-head study.

All these results belong to a specialized research setting. The larger v0.5 panel still derives its 202 anchor episodes from 22 recurring generator prototypes shared across development, calibration and evaluation. Labels are synthetic, authored by the same agent rather than independently established human judgments. Technical classroom topics dominate. Real accents, similar names, overlapping speech, uncertain identity, unfamiliar social settings and out-of-domain subjects remain incomplete coverage. Neither softmax scores nor success on this panel establish calibrated confidence in a real room.

The architecture, however, need not stop at turn-taking. Another decision head could predict whether this turn merits reasoning, and at what budget: for example, none, brief or extended. The controller could then choose a reasoning path before committing to a long generation, and optionally produce an appropriate conversational cue while processing. This would make reasoning an outcome of decision-making rather than a prerequisite for every decision.

Update: Our proof-of-concept (PoC) head failed to demonstrate the potential of this approach for reasoning budget planning.

Reasoning budget planning remains an unproven capability. It would need examples measuring whether extra reasoning actually improves the answer enough to justify its cost, followed by independent evaluation and calibration. Missing speaker identity may require waiting for evidence; more internal computation cannot supply an unheard fact. The useful design principle is to let one language representation support several trained outputs: words when an answer is needed, and small explicit decisions that determine when and how to produce it.

Metric definitions, pinned experiment identities and source files are collected in evidence.md. No new benchmark was run for this essay.