Skip to content
JeffHub

Jeff-Code

A coding agent, forked from Pi, in which Jeff makes two decisions around every turn of Qwen3.8-27B: it takes the routine information-gathering steps itself, and it decides when Qwen3.8-27B needs to think hard. Each decision is a small adapter on the fixed Jeff v1.3 base. The same pass rate, 47% faster per task (32% less time).

Results

47% faster¹¹ 32% less time per task, on average: a task takes 0.68× of Qwen3.8-27B's time alone (95% interval 0.64× to 0.72×).
62.4% vs 62.8%The same pass rate: paired difference −0.2 points (95% interval −2.6 to +2.1) over 1,242 paired tasks.
−7.6 pointsFaster, but clearly worse (95% interval −10.6 to −4.5): Jeff's decisions are what keep the quality.
92% / 68%code (steps) / code-router (thinking), on held-out scoring tasks.
By benchmark
Per benchmark: paired tasks, Jeff-Code's time as a share of Qwen3.8-27B alone's, and the pass rates
BenchmarkTasksTime vs Qwen (95%)Speed-upPass rate, Qwen → Jeff-CodeDifference (95%)
SWE-bench Verified4860.63× (0.57× to 0.69×)clear70.6% → 70.8%+0.4 (−3.5 to +4.3)
SWE-rebench · two rounds3700.66× (0.61× to 0.73×)clear58.9% → 58.1%−0.8 (−5.7 to +3.5)
Terminal-Bench Pro · two rounds1950.64× (0.55× to 0.76×)clear61.2% → 62.8%+1.5 (−4.6 to +7.7)
Terminal-Bench 2.01080.96× (0.78× to 1.16×)none clear75.9% → 70.0%−4.6 (−12.1 to +3.7)
SkillsBench420.91× (0.68× to 1.20×)none clear28.6% → 31.0%+2.4 (−11.9 to +16.7)
Harbor Index410.71× (0.51× to 0.99×)clear12.2% → 9.8%−2.4 (−12.2 to +7.3)

Benchmarks run twice have both rounds in one row, the intervals resampling whole tasks over both. Every pass-rate interval includes no difference. The median task takes 0.70× the time; the biggest gains are on typical software-engineering work.

How it was measured

  • Against: Qwen3.8-27B alone in the same Jeff-Code build, every Jeff feature off, full thinking on every turn and no thinking limit.
  • Paired: each task run in both settings side by side, at the same time on the same Qwen server, and compared pair by pair.
  • Six benchmarks: SWE-bench Verified, SWE-rebench, Terminal-Bench Pro, Terminal-Bench 2.0, SkillsBench, Harbor Index; held-out tasks only.
  • Settings: step threshold 0.40; thinking off unless P(xhigh) ≥ 0.6.
  • Thinking off: down to −13.5 points on Terminal-Bench 2.0.
  • Left out: an immaterial share of task pairs, for both sides, where the system killed a session (for example, out of memory), plus the few sessions still running when the report was made.

Total time over all tasks: 0.86× (95% interval 0.80× to 0.93×), a smaller drop than per task, because a few long sessions dominate the sum: on some hard tasks Jeff-Code keeps going until the time limit where Qwen alone gives up after a few minutes. About 5% of tasks take more than 30 minutes longer; on those, Jeff-Code solved 26 against 24.

The Jeff-Code repository. Offline accuracy reported by the Jeff-Code session: practically the same as the full fine-tune compared offline (step 92%, router 68% for both). Numbers supplied by the maintainers on 2026-10-05. Source: results/sources/v1.3/jeff-code-owner-supplied.json (owner-supplied)

How it works

Jeff works ahead

Before each of Qwen3.8-27B's turns, code takes the information-gathering steps it is confident about: read a file, list a folder, search the code, check which tools are installed. It picks the tool, then its argument, and can take several steps in a row. Qwen3.8-27B starts its turn with the results already in front of it. It can also run the tests or a build, repeat Qwen3.8-27B's last command and, with the run-approval setting the evaluation used, run a script Qwen3.8-27B wrote or install a missing package. Writing and editing files always stay with Qwen3.8-27B, and whenever Jeff is unsure, it hands over.

Jeff decides when to think

code-router decides whether Qwen3.8-27B needs to think hard on its next turn. Thinking stays off unless Jeff is at least 60% sure the turn needs it. Because Jeff returns a calibrated probability, that is a single setting: a higher threshold is faster and less careful.

The two adapters

Both are LoRA adapters on the fixed Jeff v1.3 base, loaded beside any other adapters on one server. They are built for Jeff-Code with Qwen3.8-27B and were measured only there; Jeff-Code builds their requests itself.

Decision 1 · steps

code: the steps Jeff takes itself

In the Jeff-Code agent, takes the information-gathering steps ahead of Qwen3.8-27B (reads, listings, searches) and hands over when unsure.

How it was made

In short: Imitates Qwen3.8-27B: each label is the information-gathering step Qwen actually took next, built by code from Qwen sessions, with no other model in the loop.

  • Test set: Whole tasks are held out, and the evaluation uses only held-out tasks of six benchmarks: SWE-bench Verified in full (500 tasks; none of its repositories is used for training), Terminal-Bench 2.0 40 of its 89 tasks (45 train; the other 4 are near-twins of evaluation tasks and are used for neither), SWE-rebench 189 tasks (671 train; split by repository), Terminal-Bench Pro 100 (90 train), SkillsBench 44 (42 train) and Harbor Index 41 (30 train). A leak check compares every training task with the evaluation tasks.
  • Training data: not published.
  • How it works: Jeff works ahead of the large model. Before each of Qwen's turns, Jeff decides whether it can already take the next information-gathering step itself: read a file, list a folder, search the code, check which tools are installed. It picks the tool first, then its argument, and can take up to 8 steps in a row. Qwen then starts its turn with those results already in front of it. It can also run the tests or a build, repeat Qwen's last command and, with the run-approval setting the evaluation used, run a script Qwen wrote or install a missing package. Writing and editing files always stay with Qwen, and whenever Jeff is unsure, it hands over. Because Jeff gives a calibrated probability for every choice, its step threshold is a single setting: Jeff takes a page's top option only when its probability is at least the threshold (0.40 in the measured runs), and hands over otherwise. A second adapter, code-router, decides on every turn whether Qwen should think hard.
  • How we trained it: Jeff predicts what Qwen would do next. Each training label is the information-gathering step Qwen actually took next, built by code from Qwen sessions; if Jeff takes that step early, Qwen gets the result without spending a turn on it.
  • The data comes in three stages: (1) about 14,400 public Qwen3.8-27B agent sessions (ukisai/Qwen3.8-27B-multi-turn-agent-sft), with Jeff's option lists rebuilt from each transcript; (2) public Terminal-Bench sessions of Qwen3.8-27B, replayed in each task's container; (3) our own sessions of Qwen3.8-27B in Jeff-Code, where Jeff logs its full option list before every turn without acting.
  • Training order: the adapter was trained in curriculum order, stage 1, then stage 2, then stage 3, with stage 1 sampled down to half the size of stage 3.
  • Questions longer than 8,192 tokens are cut to fit, keeping the head and tail of the state; the question and its options are never cut. The same rule applies at run time.
  • Jeff only ever offers files, folders, programs or packages that have already appeared in the task or in earlier output, the same things Qwen could act on.
  • To keep training and real use identical, Qwen works with bash only in Jeff-Code.
  • Jeff-Code is a fork of the Pi coding agent (MIT), with its own changes such as a loop guard and Jeff acting first.
  • Jeff-Code's measured results come entirely from LoRA adapters on the fixed Jeff v1.3 base - this one and code-router - both loaded beside the other adapters on one jeff-base.

Data and licence

Adapter: Apache-2.0

  • ukisai/Qwen3.8-27B-multi-turn-agent-sft (about 14,400 public Qwen3.8-27B agent sessions)Generated by Qwen3.8-27B (the sessions); labels built by codeJeff's option lists are rebuilt from each transcript; the label is the step Qwen took next.
  • openguardrails Terminal-Bench sessions of Qwen3.8-27BGenerated by Qwen3.8-27B (the sessions); labels built by codePublic sessions, replayed in each task's container.
  • Terminal-Bench 2.0 tasks (outside the 40 evaluation tasks and their four near-twins)Not generated by a modelThe tasks our own Jeff-Code sessions and the replays run on.
  • SWE-rebench tasks (the 671 training tasks; whole repositories held out, 189 tasks)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub swe-rebench/swe-rebench-leaderboard).
  • Terminal-Bench Pro tasks (the 90 training tasks; 100 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub terminal-bench-pro/terminal-bench-pro).
  • SkillsBench tasks (the 42 training tasks; 44 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub benchflow/skillsbench).
  • Harbor Index tasks (the 30 training tasks; 41 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub harbor-index/harbor-index-1.0).
  • Terminal-Bench tasks (the 27 training tasks of the 66-task set on the Harbor hub; 33 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub terminal-bench/terminal-bench).
  • Terminal-Bench Science tasks (the 35 training tasks; 35 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub terminal-bench-science/terminal-bench-science).
  • Our own Jeff-Code sessions with Qwen3.8-27BBuilt for this adapterGenerated by Qwen3.8-27B (local), the sessions; labels built by codeQwen3.8-27B in Jeff-Code (bash only); Jeff logs its full option list before every turn without acting.

Download

Through llama.cpp: the GGUF takes the same action (act on the top option at 0.40, otherwise hand over) as full precision on 99.9% of the 991 development rows at Q8_0, 97.9% at Q4_K_M. Running Jeff with llama.cpp

Decision 2 · thinking

code-router: when Qwen thinks hard

In the Jeff-Code agent, picks how hard Qwen3.8-27B should think on each turn (off, low, medium, xhigh), so easy turns run fast.

Serve code-router on the Q8_0 base only. Its choices are close calls, so at Q4_K_M it changes about 6% of its thinking on/off decisions (Q8_0: 0.8%).

How it was made

In short: Each label is the lowest thinking level at which Qwen3.8-27B still took a good step, found by re-asking recorded turns; judged by a code rule, else by Qwen3.8-Max (hosted, Alibaba Cloud DashScope).

  • Test set: Whole tasks are held out, as for the code adapter - the benchmark tasks used to measure Jeff-Code are never used for training.
  • Training data: not published.
  • How it works: before each of Qwen's turns, the router gives a probability for each of four thinking levels - off, low, medium and xhigh. In the measured runs, thinking stays off unless the probability of xhigh, out of all four, is at least the thinking threshold (0.6). The code adapter takes the routine information-gathering steps; together they are Jeff-Code's two decisions.
  • Why it matters: thinking off throughout is faster still but clearly worse (7.6 points lower pass rate; down to 13.5 on Terminal-Bench 2.0); the router's decisions are what keep the quality.
  • How the labels are made: each recorded Qwen3.8-27B turn (recorded at xhigh) was asked again with exactly the same request at thinking off (temperature 0), then low, then medium (normal sampling), stopping at the first level whose action was good. The label is that level, or xhigh if none was.
  • What counts as good: first a code rule - the same intent as the recorded xhigh action, meaning the same kind of step and the same target per command. Otherwise Qwen3.8-Max (hosted, Alibaba Cloud DashScope; thinking off, temperature 0) answered yes or no to: 'Would Step 2 serve the task as well as Step 1 at this moment?', seeing the task and the last 3 steps. Inline scripts, program runs, file writes and edits, and final answers always went to that judge.
  • Sessions recorded at medium thinking were left out.
  • Label shares in training: off about 27%, low about 8%, medium about 5%, xhigh about 60%.
  • In the evaluated runs, low and medium were never used: thinking was off unless the probability of xhigh was at least 0.6.
  • Questions longer than 8,192 tokens are cut to fit, keeping the head and tail of the state; the question and its options are never cut. The same rule applies at run time.
  • Jeff-Code's measured results come entirely from LoRA adapters on the fixed Jeff v1.3 base - this one and code - loaded beside the other adapters on one jeff-base.

Data and licence

Adapter: Apache-2.0

  • Recorded Qwen3.8-27B sessions (at xhigh), re-asked at lower thinking levelsGenerated by Qwen3.8-27B (the re-asked answers); Qwen3.8-Max (hosted, Alibaba Cloud DashScope), for the yes-or-no judgements the code rule could not makeSessions recorded at medium thinking were left out. The re-asked answers come from Qwen3.8-27B; where the code rule could not decide, Qwen3.8-Max judged whether the lower-level step was as good.

Download

Through llama.cpp: the GGUF takes the same thinking on/off decision (on when P(xhigh) ≥ 0.6) as full precision on 99.2% of the 987 development rows at Q8_0, 94.3% at Q4_K_M. Running Jeff with llama.cpp