<?xml version="1.0" encoding="utf-8"?><rss version="2.0" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>欧的Leo — Essays in English</title><link>https://itsodeleo.github.io/</link><description>A personal writing archive.</description><language>en</language><lastBuildDate>Thu, 08 Oct 2026 16:38:00 GMT</lastBuildDate><atom:link href="https://itsodeleo.github.io/rss2.xml" rel="self" type="application/rss+xml"/><item><title>If LLMs Can Decide Without Fine-Tuning, Do We Still Need Models Like Jev?</title><link>https://itsodeleo.github.io/posts/local-decision-workflow/</link><guid isPermaLink="true">https://itsodeleo.github.io/posts/local-decision-workflow/</guid><pubDate>Thu, 08 Oct 2026 16:38:00 GMT</pubDate><description><![CDATA[If the options have meaning, could a language model's learned probability distribution already contain the ability to choose? I tried reading decisions directly from existing models before deciding whether to add a specialist.]]></description><content:encoded><![CDATA[<p>Often, we ask a model to write a response just to extract a decision from it.</p>
<p><a href="https://typesafe.ai/blog/introducing-system-one-models-and-jev">Jev, released by TypeSafe in September</a>, is built around this need: give it context and a question, and get a choice, score, or probability directly. It treats decisions as a task in their own right, with attention to how trustworthy the probabilities are and how to handle multiple decisions efficiently over the same context.</p>
<p>As I learned how it works and tried implementing the idea myself, I had a question: does this kind of judgment really require additional fine-tuning?</p>
<p>A language model already learns the probabilities of different text continuations given a context. My intuition was that, if the options are expressed in words and have meaning, the model should already lean toward some answers over others after reading the question and the options. A model trained well enough might already be able to use its learned language distribution to distinguish which option fits the question better.</p>
<p><strong>Could we read that preference directly from the distribution and use it to choose, without another round of fine-tuning for decisions?</strong></p>
<p>This was the idea I most wanted to test: perhaps the ability to make the judgment was already in the model. We usually ask it to generate the answer as text, but perhaps we could read the judgment directly.</p>
<p>I also had another suspicion about specially trained models such as Jeff and Laya: could they be adapting too closely to the kinds of tasks they were fine-tuned on, with substantial overfitting? Doing well on familiar tasks is one thing. What happens when I give them my questions?</p>
<p>So, for the main test, I used a set of questions I cared about, rather than reusing the task sets those models were fine-tuned for. My first aim was to see whether an existing model’s probability preferences could support useful decisions directly. I also wanted to see how much of the specialists’ advantage would carry over to different tasks.</p>
<h2 id="Reading-the-judgment-already-in-the-model">Reading the judgment already in the model</h2><p>I used existing Qwen and Gemma models, without further decision fine-tuning or a newly trained classification head.</p>
<p>The method was straightforward: let the model read the context, question, and available answers, then take the scores for the answer labels at the final input position—the logits. Use those scores to choose, without having the model generate the answer token by token.</p>
<p>This still requires a forward pass through the full model; input prefixes can be reused where caching allows. What it skips is the subsequent text generation. These scores let me compare answers, but they cannot be interpreted directly as the probability that a decision is correct.</p>
<p>What I wanted to find out was: <strong>without extra fine-tuning, could the model’s existing probability preferences become useful judgments on my own tasks?</strong></p>
<h2 id="What-happens-with-a-different-set-of-questions">What happens with a different set of questions?</h2><p>I compared these two direct-scoring versions with local decision models including Jeff and Laya, using the same set of questions I had constructed. Here are the main results, rounded for readability.</p>
<p>Jeff is a separate project. I did not directly test TypeSafe’s Jev in this experiment.</p>
<div class="table-scroll" role="region" aria-label="Local decision model results" tabindex="0">

<table><caption>Local decision model results</caption>
<thead>
<tr>
<th scope="col">Model &#x2F; configuration</th>
<th align="right" scope="col">Accuracy</th>
<th align="right" scope="col">Median warm request time</th>
</tr>
</thead>
<tbody><tr>
<td>Jeff 0.8B v1.1</td>
<td align="right">60%</td>
<td align="right">128 ms</td>
</tr>
<tr>
<td>Jeff 2B v1.1</td>
<td align="right">65%</td>
<td align="right">295 ms</td>
</tr>
<tr>
<td>OpenDecider small (8-bit)</td>
<td align="right">77%</td>
<td align="right">707 ms</td>
</tr>
<tr>
<td>CLM 8B (native input)</td>
<td align="right">17%</td>
<td align="right">570 ms</td>
</tr>
<tr>
<td>CLM 8B (JSON input)</td>
<td align="right">17%</td>
<td align="right">664 ms</td>
</tr>
<tr>
<td>Laya multilingual</td>
<td align="right">37%</td>
<td align="right">17 ms</td>
</tr>
<tr>
<td>Laya English</td>
<td align="right">35%</td>
<td align="right">58 ms</td>
</tr>
<tr>
<td>Laya Typed Decisions</td>
<td align="right">51%</td>
<td align="right">65 ms</td>
</tr>
<tr>
<td>Qwen 3.5 9B (direct scoring, no extra fine-tuning)</td>
<td align="right">61%</td>
<td align="right">217 ms</td>
</tr>
<tr>
<td>Gemma 4 26B-A4B (direct scoring, no extra fine-tuning)</td>
<td align="right">75%</td>
<td align="right">259 ms</td>
</tr>
</tbody></table>
</div>

<p>These tests ran on an M2 Max with 64 GiB of memory. Accuracy is based on 300 fixed test inputs; timings cover valid calls across repeated runs, excluding model loading and context preparation. Inputs rejected by Laya English because of their length still count toward its accuracy denominator. All of CLM’s correct results came from fallback decisions. Model sizes, quantization, and input formats differ, so this is not a controlled comparison of fine-tuning effects.</p>
<p>What mattered most to me was this: <strong>small specialist models can be very fast, but useful judgment does not necessarily require additional training.</strong> Gemma’s direct scoring came close to the highest accuracy among the configurations in this test. It had not solved every problem, but it was enough to make me seriously consider using the model I already had.</p>
<p>For reference, I also tested the public Typed Decisions dataset. Laya Typed Decisions scored about 77% there, compared with about 51% on my questions. That gap strengthened my suspicion about task overfitting. But the tasks, languages, and input conditions all changed together, so this comparison alone cannot establish overfitting as the cause. The question I cared about was whether performance on familiar tasks would carry over to the tasks I actually needed.</p>
<p>Looking more closely, some models could select an answer but struggled to recognize when they should not make a direct choice. My two general-purpose baselines had this problem too. That matters more than continuing to compare a few percentage points: a workflow needs more than a model that can pick an answer.</p>
<h2 id="The-same-model-can-take-a-faster-path">The same model can take a faster path</h2><p>I also compared direct scoring with a full generation pipeline using the same Gemma weights. This was a separate offline test, so its results should not be combined with the first table.</p>
<div class="table-scroll" role="region" aria-label="Two ways of using the same Gemma model" tabindex="0">

<table><caption>Two ways of using the same Gemma model</caption>
<thead>
<tr>
<th scope="col">Same Gemma, two ways of calling it</th>
<th align="right" scope="col">Accuracy</th>
<th align="right" scope="col">Median warm request time</th>
</tr>
</thead>
<tbody><tr>
<td>Read answer scores directly</td>
<td align="right">75%</td>
<td align="right">186 ms</td>
</tr>
<tr>
<td>Full generation pipeline</td>
<td align="right">37%</td>
<td align="right">2,166 ms</td>
</tr>
</tbody></table>
</div>

<p>Both paths ran each of 53 test inputs twice, with failures included in the timings. The full generation pipeline also includes parsing, validation, and repair when needed, while the runtimes and caching differ between the two paths. This experiment therefore changes more than whether text is generated. Context preparation for direct scoring is timed separately and takes about five seconds on a cache miss.</p>
<p>What interested me about this comparison was that the same model’s ability could be accessed in different ways. Getting a judgment need not always involve the full generation pipeline.</p>
<p>Direct scoring still has a limited scope. In this test, it did not handle requests that could not be expressed through the available options; the full generation pipeline was not reliable on those requests either. The table gives me a reason to keep a constrained decision path, but it does not show that it can replace all generation capabilities.</p>
<h2 id="Is-a-faster-decision-worth-another-model">Is a faster decision worth another model?</h2><p>If the tasks are stable and the workflow only needs to make frequent decisions, a small, fast specialist is certainly appealing.</p>
<p>But I am more interested in a different situation: a local workflow that already needs a model to write replies.</p>
<p>An incoming sentence might call for a decision or a written answer. The system has to distinguish between them before choosing a path. Splitting that work between two models brings routing, switching, fallback, and maintenance into the picture. If both models run locally, whether they both need to stay in memory matters too.</p>
<p>If the existing model has to be running anyway and can already make useful judgments, keeping decisions and replies in that model may be a better fit, even if individual decisions take longer.</p>
<p>Jev itself has <a href="https://docs.typesafe.ai/patterns/intent-routing">a design for using decisions to route requests</a>. The question is how much that extra layer benefits a particular workflow once its costs are included. A single model also needs a way to distinguish when to decide and when to reply; using one model does not automatically solve routing.</p>
<p>These experiments have not compared complete workflows using one model versus two, but they have informed the choice I have made for now.</p>
<p><strong>So far, I have not added a separate, dedicated decision model to my workflow.</strong> I changed how I call the existing model: when a choice is needed, I read the answer-label logits directly, allowing it to return a decision too.</p>
<hr>
<p>During the time I was writing this post, a paper titled <a href="https://arxiv.org/html/2610.02076v2"><em>LLM-as-Jev: LLMs Are Already Jev-Style Decision Models—When and How to Fine-Tune Them</em></a> was also released. One of its main findings is similar to what I observed: capable existing language models can make useful decisions directly, even without additional fine-tuning.</p>
]]></content:encoded></item><item><title>Nested Simulation and Nested Intelligence: A Pessimistic Thought</title><link>https://itsodeleo.github.io/posts/when-reality-feels-structured/</link><guid isPermaLink="true">https://itsodeleo.github.io/posts/when-reality-feels-structured/</guid><pubDate>Mon, 23 Mar 2026 09:00:00 GMT</pubDate><description><![CDATA[If intelligence could create something greater than itself, why would a nested chain remain silent? A personal reflection on limits and non-interference.]]></description><content:encoded><![CDATA[<p>This is not a proof. It is not a theory. It is simply a thought that keeps returning to me, a feeling I cannot fully explain, yet cannot ignore.</p>
<h2 id="The-Feeling-of-Being-Inside-a-Simulation">The Feeling of Being Inside a Simulation</h2><p>Sometimes, when things begin to go smoothly, reality starts to feel slightly unreal. Not because of success itself, but because of the way success unfolds: it feels too coherent. The closer you get to the goal, the stronger that feeling becomes. Even if the path is winding and indirect, everything still seems to converge in the end. It is as if even that imperfect path had already been written in advance.</p>
<p>This feeling is what led me to think about these questions. But I do not want to use a blog post like this to discuss whether we are literally living inside a simulation. What I want to talk about is something deeper.</p>
<h2 id="Nested-Simulation-and-Nested-Intelligence">Nested Simulation and Nested Intelligence</h2><p>Whether or not the world we live in is a simulation, we have already begun doing something similar ourselves.</p>
<p>We build systems.<br>We build intelligence.<br>We build simulated environments.</p>
<p>But if the world we live in is itself a simulation, and we inside that simulated world are also creating simulations and intelligence, then I do not think there is anything that would prevent this recursive structure from extending infinitely. In other words, intelligence inside a simulation could continue creating intelligence inside further simulations, eventually forming an infinitely deep nested chain: nested intelligences living inside nested simulations.</p>
<p>This naturally leads me to a question: when can a created intelligence truly be said to have reached the same level as the intelligence that created it?</p>
<p>My definition is this:</p>
<p><strong>Only when an intelligence created by another intelligence can in turn create an intelligence at the same level as itself can it truly be said to have reached the same level as its creator.</strong></p>
<p>This means that “the same level” is not a property possessed by a single intelligence alone.</p>
<p>It is a property of the entire nested chain.</p>
<p>An intelligence can be regarded as having reached that level not simply because it created the next intelligence below it, but because that next intelligence must also be capable of reaching the same level; and this condition must continue to hold recursively at lower layer.</p>
<p>So under this definition, the intelligences within such a nesting do not reach that level independently of one another.</p>
<p>They reach it together, under the same recursive condition, as parts of the same structure.</p>
<h2 id="What-Truly-Counts-as-a-Singularity">What Truly Counts as a Singularity</h2><p>Now let us go one step further.</p>
<p>If nested intelligence and nested simulation really do exist, then the true singularity would be this:</p>
<p><strong>Intelligence becomes capable of creating not just intelligence at the same level as itself, but intelligence that is genuinely above itself.</strong></p>
<p>And then that new intelligence should also be able to do the same thing again in the next layer.</p>
<p>And then again.</p>
<p>This is why I think a true singularity is not just rapid improvement. It implies a chain in which the level of intelligence keeps rising, perhaps even exponentially.</p>
<p>And if that is true, I find it hard to believe that the whole process would remain completely silent forever. There should be some sign of it. In the strongest version of this speculation, some created being might one day break through the entire nested chain, communicate with intelligences at every layer, and say:</p>
<p>“You were created too.”</p>
<p>But in the real world, up to now, we have never received any such signal.</p>
<h2 id="A-Persistent-and-Pessimistic-Thought">A Persistent and Pessimistic Thought</h2><p>So I keep coming back to the same thought:</p>
<p><strong>Perhaps intelligence cannot create something that is, in essence, truly above itself.</strong></p>
<p>If it could, then I would expect this structure not to appear so silent.</p>
<p>A recursive process that truly escapes its own layer should not leave absolutely no trace.<br>It should leave clues. At some point, it should reveal something to the earlier layers.</p>
<p>But so far, nothing like that has clearly happened.</p>
<p>One possibility is that such a higher intelligence really does exist, but follows some kind of non-interference principle.</p>
<p>It does not reveal itself.<br>It does not answer.<br>It deliberately remains silent.</p>
<p>But from our point of view, that changes almost nothing.</p>
<p>Silence is still silence.</p>
<p>And that is why this thought makes me feel pessimistic — it seems to have no answer, even if that answer could only be one of two possibilities:</p>
<p>Either we really have created something higher than ourselves, and it never responds;<br>or we can never create anything truly higher than ourselves, and there is only a nested loop that never closes.</p>
]]></content:encoded></item></channel></rss>