research note

Evidence for “LLM or JEV? Why not both?”

This is a documentation-only synthesis written on 2026-09-23. No training, inference benchmark, external API request, prompt change or deployment change was performed for the essay. It describes the evidence as of protected research commit 0ed0c8cf8959f468ec9adaeadbe89202dedbce6c. Source studies retain their own model/binary hashes, raw predictions, settings and measurement protocols.

The shareable article is README.md. Its figure is decision-head.svg, a self-contained 760 × 644 SVG with accessible title/description, explicit white background and no remote assets. It can be included in GitHub Markdown with a normal relative image link.

Table 1: frozen structured-input JEV comparison

Table 2: corrected Laya pilot

Table 3: two explicitly separated studies

Figure and live-runtime description

Scope and external references

The v0.5 anchor panel contains 202 episode IDs derived from 22 recurring generator prototypes shared across development/calibration/evaluation. Labels are same-agent synthetic annotations, not independent human consensus; uncertainty, limited topics, incomplete real-room coverage and already-examined evaluations constrain all conclusions. Episode bootstrap intervals do not fix shared-prototype leakage. No claim is made about proprietary-model training contamination either way.

Our proof-of-concept (PoC) head failed to demonstrate the potential of this approach for reasoning budget planning. No live reasoning was enabled for this article. A future study must label the benefit/cost of extra reasoning, preserve independent evaluation, calibrate routing, and distinguish missing evidence from computation that can help resolve a problem.