Jeff v1.3: Jeff-Code makes Qwen 3.8-27B finish coding tasks 47% faster (32% less time) on average at the same pass rate; plus 15 adapters & GGUFs
Jeff v1.3 is live, and with it come a number of updates. See jeffhub.ai and github.com/firelex/jeff for full details.
Disclosure: Jeff and Jeff-Code are our project. Everything here is open source and free: weights, adapters and code.
The highlight: Jeff-Code
Jeff-Code is a coding agent with two Jeff v1.3 adapters trained specifically for Qwen 3.8-27B. Jeff-Code is a fork of Pi by Mario Zechner (MIT licence).
We forked Pi because its extension framework doesn't currently let a fast decision model sit deep enough inside the agent loop. Along the way, we made a number of other changes as well (see below).
Aside from hopefully being useful to people who run Qwen 3.8-27B locally as their daily coding model, Jeff-Code is also a conceptually interesting experiment: how far can a System 1 model go inside a coding agent?
The results, run side by side in paired blocks:
- Same quality: with Jeff's thinking threshold at 0.6, Jeff-Code matches Qwen 3.8-27B's pass rate: 62.4% against 62.8%; paired difference −0.2 points, 95% interval −2.6 to +2.1, over 1,242 paired tasks.
- 47% faster (32% less time) per task¹: on average a task takes 0.68× the baseline's time (geometric mean of the per-task time ratios, 95% interval 0.64–0.72; the median task, 0.70×).
- Where it helps most: typical software-engineering work. SWE-bench Verified 0.63×, SWE-rebench 0.66×, Terminal-Bench Pro 0.64× (both over two rounds), Harbor Index 0.71×. On Terminal-Bench 2.0, with its long, hard tasks, there is no clear speed-up (0.96×, interval 0.78–1.16); on SkillsBench neither (0.91×, interval 0.68–1.20).
- The benchmarks: we evaluated only on tasks Jeff never saw in training. SWE-bench Verified ran in full (all 500 tasks; none of its repositories were used for training). For the benchmarks we also trained on, we split the tasks and ran every held-out task; a few pairs hit by repeated infrastructure failures are left out (see below). Terminal-Bench 2.0 (40 of its 89 tasks, 3 attempts each; 45 were used for training, and the other 4 are near-twins of evaluation tasks, so they were used for neither), SWE-rebench (189 held-out tasks, 2 rounds), Terminal-Bench Pro (100 held-out, 2 rounds), SkillsBench (44 held-out) and Harbor Index (41 held-out). Within those splits nothing was sampled. We also ran Terminal-Bench (original) and Terminal-Bench Science, but Qwen solves almost none of those tasks in any setting, so they can't show a difference and are left out of the pooled numbers.
- What it's compared against: Qwen 3.8-27B alone in the same Jeff-Code build with every Jeff feature switched off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. The only remaining differences from original Pi are a rarely triggered runaway cut-off (it stepped in 3 times) and trace logging. Task pairs hit by an infrastructure failure (out of memory, a stalled session, a test environment that wouldn't start) were run again once; the pairs that failed again, and a handful of re-runs still unfinished at launch, are left out for both sides (under 3% of pairs) and listed in the full report. One Terminal-Bench 2.0 task, pytorch-model-recovery, is left out of every comparison: a harness bug stopped its baseline sessions before they began.
- Why not just turn thinking off? We tried: with Qwen's thinking off throughout (and the same safeguards), tasks are faster still but clearly worse: −7.6 points (−10.6 to −4.5), up to −13.5 on Terminal-Bench 2.0. Jeff deciding when Qwen should think is what keeps the quality. That's the case for a small decision model.
¹ Over all tasks combined, total time drops less, by 14% (0.86×, interval 0.80–0.93). Most of that gap comes from a small set of tasks: in about 5% of them, Jeff-Code runs more than 30 minutes longer than Qwen alone, because it keeps going where Qwen alone gives up after a few minutes. The good news is, sometimes that pays off: in those tasks Jeff-Code solved 26 to Qwen's 24. Stopping only the hopeless runs would bring the total to about 0.7× in the best case, but we can't yet recognise them reliably ahead of time (a learned stop rule held up only at about 0.85×). That's the next step.
Per benchmark, against Qwen 3.8-27B alone:
| benchmark | paired tasks | Qwen alone | Jeff-Code | difference, points (95% interval) | time per task |
|---|---|---|---|---|---|
| SWE-bench Verified | 486 | 70.6% | 70.8% | +0.4 (−3.5 to +4.3) | 0.63× |
| SWE-rebench (2 rounds) | 370 | 58.9% | 58.1% | −0.8 (−5.7 to +3.5) | 0.66× |
| Terminal-Bench Pro (2 rounds) | 195 | 61.2% | 62.8% | +1.5 (−4.6 to +7.7) | 0.64× |
| Terminal-Bench 2.0 (3 attempts) | 108 | 75.9% | 70.0% | −4.6 (−12.1 to +3.7) | 0.96× |
| SkillsBench | 42 | 28.6% | 31.0% | +2.4 (−11.9 to +16.7) | 0.91× |
| Harbor Index | 41 | 12.2% | 9.8% | −2.4 (−12.2 to +7.3) | 0.71× |
| all six, pooled | 1,242 | 62.8% | 62.4% | −0.2 (−2.6 to +2.1) | 0.68× |
No benchmark shows a clear pass-rate difference: every interval includes zero. Benchmarks we ran twice are combined, with intervals computed over both rounds. The pass rates count every finished session, and the difference counts only tasks finished in both settings, so it isn't exactly the gap between the two pass rates.
How it works: Jeff makes two kinds of decisions around every Qwen turn, each handled by a small adapter on the same Jeff base.
First, Jeff works ahead of Qwen. If it can take the next information-gathering step itself (read a file, list a folder, search the code, check which tools are installed), it does. It picks the tool first and then its argument, and can take several steps in a row if it's confident. Qwen then starts its turn with those results already in front of it, rather than spending a slow turn fetching them.
It can also run the tests or a build, repeat Qwen's last command and, with the run-approval setting the evaluation used, run a script Qwen wrote or install a missing package. Writing and editing files always stay with Qwen, and whenever Jeff is unsure, it hands over to Qwen. This is interesting because it's an example of a System 1 model that not only picks a tool, but also parameterizes it.
Second, Jeff decides whether Qwen needs to think hard on its next turn. Thinking stays off unless Jeff is confident the turn needs it (the 0.6 above is that threshold). Because Jeff returns a calibrated probability for every choice, each of these behaviours is controlled by a single setting.
How we trained it: Jeff predicts what Qwen would do next.
For information-gathering, each training label is simply the step Qwen actually took next. If Jeff can take that step early, Qwen gets the result without spending a turn on it. These labels are built directly from Qwen sessions by code, with no other model in the loop.
The thinking decision is labelled differently. For each recorded Qwen turn at full thinking, we re-asked the same request with thinking off, then low, then medium. The label is the cheapest level whose action was as good as the original, or full thinking if none was.
"As good" is decided by code wherever possible: for example, the same kind of step on the same target. Otherwise, Qwen3.8-Max with thinking off judges whether the cheaper step would serve the task just as well at that moment.
About 27% of turns didn't need thinking at all.
These labels are deliberately strict, which made the router cautious, and that's what preserved the pass rate. We also tried the looser question, "Is the full-thinking step materially better?" The problem is that, on an individual turn, the judge can't reliably distinguish thinking-off from a second full-thinking answer either. A router trained that way would therefore switch thinking off almost everywhere. As the full thinking-off run shows, that costs quality over the course of an entire task.
Everything above was measured using adapters: small LoRA files on the fixed Jeff v1.3 base, which is the point of the whole design. We also trained a single full fine-tune to make both decisions and compared it offline. Its accuracy was practically identical. So we're releasing the adapters, which preserve the multi-adapter architecture without giving up accuracy.
Here are some of the main differences between Pi and Jeff-Code:
- Thinking per turn: Pi only switches thinking on or off, so Qwen always thought at its highest level (in our tests, Qwen's low and medium levels thought about as long as the highest one, so they saved no time). Jeff-Code sets Qwen's thinking level for each turn, fixed or decided by Jeff.
- Safeguards for thinking off: a loop guard catches repeated or near-identical actions up to six steps back (a file write only counts as progress if it changes the file); a caught repeat is thrown away and that turn is asked again with full thinking, and near-identical outputs or two failed commands in a row send the next turn to full thinking. The thinking-off comparison run had exactly the same safeguards, the thinking limit below and the same escalations to full thinking (about 3% of its turns ended up thinking); the only differences from Jeff-Code are Jeff's decisions. So its quality gap is down to Jeff.
- Runaway cut-off: if Qwen's thinking or text keeps repeating itself, the reply is stopped and re-asked with full thinking. This was on in every run, including the full-thinking baseline (it stepped in 3 times there, across the six pooled benchmarks).
- Thinking limit: at 8,000 thinking tokens, Qwen answers from what it has thought so far. That also rescues replies that would otherwise hit the 32K output limit, which ends a Pi session. The full-thinking baseline ran without this limit, as plain Pi does; the thinking-off run and Jeff-Code had it.
- Jeff steps: before each Qwen turn, Jeff-Code builds a menu of concrete next steps from what's already known, and Jeff takes them itself when it's confident.
- A pool of Jeff servers, one per GPU, keeps each decision at about 0.2 s, and every Qwen request and Jeff decision is logged, which is what made the training data and the evaluation possible.
We'll keep it maintained, and we'd happily upstream whatever Pi wants to take.
A new base model
The main differences between Jeff v1.3 and v1.2:
Live-last prompt structure. The fixed part of the prompt (instructions and options) now comes first and the live data comes last, so prefix caching works much better and repeated decisions over the same options get faster.
A deliberate benchmark trade-off. Because of this, the zero-shot performance of Jeff-base over long option lists has dropped significantly. With the new prompt order, the model reads all the options before it sees the input, and a 0.8B model is just much worse at going back over a long option list than at reading the options with the input already in mind. On our general panel the v1.3 base is roughly level with v1.2 (78.6% vs 78.8%, calibration error 0.028 vs 0.024), but on unfamiliar tasks, especially with long option lists, the base on its own falls apart:
| task (no adapter) | options | v1.2 base | v1.3 base |
|---|---|---|---|
| legal-clause classification | 100 | 66.0% | 7.4% |
| support intents | 7–64 | 85.1% | 24.2% |
| triage | 2–36 | 67.1% | 47.6% |
| tool choice | 4–136 | 57.8% | 30.2% |
Here's the good news: once you add the right adapter, almost all of the accuracy lost in the base comes back: v1.3 + adapter is within −0.7 to +0.3 points of v1.2 + adapter on every task except legal-clauses (83.6% vs 85.7%). Once the adapter has learned the options, putting them first costs almost nothing, and the fixed part of the prompt can be cached. We decided the speed-up was worth it.
A new philosophy: Zero-shot on everything is no longer the goal
With v1.3, we made a big change in direction: Jeff-base is no longer trying to compete on zero-shot benchmarks. Instead, it's always meant to be used with an adapter. The base is the foundation the adapters are trained on, and an adapter is where Jeff becomes good at a given task. Adapters don't get merged into the base: you load one base and all the adapters you need, and pick one per request. We figured that a big step up in speed, with adapter accuracy the same as v1.2, was worth going all in on adapters.
Switching adapters per request costs about 11 ms in the llama.cpp library and about 20 ms in llama-server on a GPU. That's mostly down to how today's servers are built: llama.cpp rebuilds part of its compute setup when the active adapter changes, and a cached prompt prefix computed under one adapter can't be reused by another. Both could largely be avoided with a separate prefix cache per adapter and a server designed for fast switching. Since total inference time stays low, we haven't optimised it yet; meanwhile, grouping requests by adapter or giving a busy adapter its own server avoids most of the cost.
So why a Jeff base at all?
If the adapters do all the work, why keep a Jeff base? To find out, we trained the same adapters on plain Qwen3.5-0.8B (never trained by us) and on the Jeff v1.3 base, with identical recipes.
Triage (classifying support messages; accuracy on a fixed 2,000-row test sample):
| training examples | adapter on Jeff v1.3 base | adapter on plain Qwen3.5-0.8B |
|---|---|---|
| 0 (no adapter) | 47.2% | 33.8% |
| 250 | 76.4% | 76.0% |
| 500 | 79.1% | 76.5% |
| 1,000 | 80.9% | 79.6% |
| 72,000 (all) | 91.5% | 91.2% |
- With plenty of data, the base hardly matters. LoRA adapters work on plain Qwen too.
- With little data, the Jeff base gives a head start, here 0.4 to 2.6 points, because it already knows how to make a decision from a list of options and give a calibrated answer; the adapter only has to learn the task.
That matters in practice, because most real tasks don't come with tens of thousands of labelled decisions. For now, v1.3 is a single, general Jeff base. But it raises a new research question: should there be domain-specific bases (finance, security, legal and so on), instead of or alongside the general one? Perhaps. We'll look into it and report back in a separate post.
New adapters
All nine existing adapters have been retrained on v1.3.
We've also added six new ones, bringing the total to 15: two for Jeff-Code (above) and four for real-world decision tasks.
For the latter, some data comes from public data sets, converted into Jeff's type-safe schema by deterministic scripts, with GLM 5.3 rewording some of the text for variety. Some is synthetic: code simulates the scenarios and fixes every label, and GLM 5.3 writes the text. GLM never decides a label. Licence restrictions are stated per adapter.
We view the published adapters as starting points and will work on refining them over time. Any community contributions in terms of data sets and/or other adapters are always welcome.
The four real-world adapters are trading-desk, aml (anti-money laundering), sanctions and soc (security-operations alert triage). Their results, licences and limits are on jeffhub.ai; note that some build on data with non-commercial licences.
We also received one community suggestion, which we'll train next. We haven't forgotten!
GGUFs
We're not Unsloth, but we wanted to create some basic GGUFs that can be used with llama.cpp.
v1.3 ships as GGUF for llama.cpp in Q8_0 and Q4_K_M: one base file per format, plus one small LoRA file per adapter (about 169 MB) that you load with --lora. No separate full model per adapter. Every adapter, including the two Jeff-Code ones, is available this way.
- Q8_0 is effectively lossless: every adapter is within 0.4 points of full precision.
- Q4_K_M stays within about two-thirds of a point on every adapter (−0.6 to +0.7). Each quantization level has its own temperature setting so the probabilities stay calibrated. One exception: run Jeff-Code's thinking router on the Q8_0 base. Its decisions are close calls, and at Q4_K_M about 6% of its on/off decisions differ from full precision.
- To switch adapters per request in llama-server, list every adapter in each request with scale 1 or 0 (see the README).
Links
- Jeff-Code: github.com/firelex/jeff-code, results on jeffhub.ai/jeff-code
- Jeff (server, client, adapter kit): github.com/firelex/jeff
- Models on Hugging Face: mstrasser/jeff-base (revision v1.3) and the adapters as
mstrasser/jeff-adapter-<name>, all listed on huggingface.co/mstrasser - GGUF: mstrasser/jeff-base-gguf and
mstrasser/jeff-adapter-<name>-gguf - Scores, data cards and licences for every adapter: jeffhub.ai/results and each adapter's page on jeffhub.ai
