Skip to content
JeffHub

Jeff v1.3: what changed

The v1.3 base is adapter-first, all adapters are retrained on it, there are new adapters and Jeff-Code, and Jeff runs on llama.cpp. What changed and what it costs.

Jeff v1.3 is a change in direction. Zero-shot on everything is no longer the goal: the base is always meant to be used with an adapter. The base is the foundation the adapters are trained on, and an adapter is where Jeff becomes good at a task.

The base: live-last

The v1.3 base, jeff-base, is a fine-tune of Qwen3.5-0.8B with the live-last prompt layout. The fixed part of the prompt (instructions and options) comes first and the changing input comes last, so prefix caching works and repeated decisions over the same options get faster. The model reads all the options before it sees the input; once an adapter has learned its options, that costs almost nothing, and the fixed part of the prompt can be cached. Every adapter's results.

New names

From v1.3 on, the models have plain names on Hugging Face (under mstrasser/):

What Repository
The base, versioned by tag (v1.3) jeff-base
Each adapter jeff-adapter-<name>, for example jeff-adapter-soc, jeff-adapter-triage
The Jeff-Code adapters jeff-adapter-code (steps) and jeff-adapter-code-router (thinking)
GGUF base files (Q8_0 and Q4_K_M) jeff-base-gguf
GGUF LoRA per adapter jeff-adapter-<name>-gguf

In a request, an adapter is still named by its short name (model="soc", model="code"): on a Jeff server, names starting with "jeff" mean the base model.

Adapters, side by side

Adapters are not merged into the base. You load one base and all the adapters you need, and pick one per request. Each adapter records the exact base it was trained on, and the server refuses an adapter trained on a different one, so an adapter trained on another base does not load.

  • All nine existing adapters are retrained on v1.3, with about 10% of the base model's own training data mixed in as a precaution (its effect has not been measured). Their results.
  • New adapters: trading-desk, sanctions, soc and aml are trained and measured. aml follows the institution's written monitoring policy; it is not a general laundering detector, and its page shows the out-of-distribution cross-test where it scores below the base. Jeff-Code's two adapters, code and code-router, are described below.

Jeff-Code (github.com/firelex/jeff-code) is a coding agent, a fork of the Pi coding agent, in which Jeff makes two decisions for Qwen3.8-27B, each with its own LoRA adapter on the fixed v1.3 base. The code adapter works ahead of Qwen: before each of Qwen's turns it takes the information-gathering steps it is confident about (reading a file, listing a folder, searching the code, checking which tools are installed), picking the tool and then its argument, up to 8 steps in a row; it can also run the tests or a build, repeat Qwen's last command and, with the run-approval setting the evaluation used, run a script Qwen wrote or install a missing package. Writing and editing files always stay with Qwen, and it hands over when it is unsure; it takes a step only when the top option's probability is at least the step threshold, and hands over otherwise. The code-router adapter picks one of four thinking levels for Qwen's turn (off, low, medium, xhigh); in the measured runs, thinking stays off unless the probability of xhigh is at least the thinking threshold. For the steps, Jeff predicts what Qwen would do next: each label is the information-gathering step Qwen actually took next, built by code from Qwen sessions, trained in three stages in curriculum order.

With Jeff's thinking threshold at 0.6 (step threshold 0.40), Jeff-Code matches Qwen3.8-27B's pass rate on six benchmarks: 62.4% against 62.8%, a paired difference of −0.2 points (95% interval −2.6 to +2.1) over 1,242 paired tasks, each run in both settings side by side. On average a task is 47% faster (32% less time): it takes 0.68× the time. Thinking off throughout is faster still but clearly worse (−7.6 points): Jeff's decisions are what keep the quality. The details.

We focused on commercially relevant applications. Some adapters use public data sets, converted into Jeff's type-safe format by deterministic scripts, with GLM 5.3 rewording some of the text for variety. For others the data is synthetic: code simulates the scenarios and fixes every label, and GLM 5.3 writes the text. GLM never decides a label. Licence restrictions are stated on each adapter's page.

We see the published adapters as starting points and will keep refining them. Data sets and adapters from the community are welcome: see suggest an adapter and submitting an adapter.

llama.cpp

v1.3 ships as GGUF for llama.cpp in Q8_0 and Q4_K_M: one base file per format, plus one small LoRA file per adapter (169 MB) that you load with --lora. There is no separate full model per adapter.

  • Q8_0 is effectively lossless: every adapter test is within 0.4 points of full precision.
  • Q4_K_M stays within 0.4 points on every adapter test (−0.4 to +0.0). Each format has its own temperature setting, so the probabilities stay calibrated.
  • Switching adapters per request costs about 11 ms in the llama.cpp library and about 20 ms in llama-server on a GPU. With llama-server, list every adapter in each request with scale 1 or 0.

Running Jeff with llama.cpp.

What stays fixed for the life of v1.3

  • The request format: state, questions and instructions, with the changing state field last.
  • The option rules: named options, keys never bare numbers.
  • The answer format: a probability for every option.

A data set written to the data guidelines trains on v1.3 as it is.