code: the steps Jeff takes itself
In the Jeff-Code agent, takes the information-gathering steps ahead of Qwen3.8-27B (reads, listings, searches) and hands over when unsure.
How it was made
In short: Imitates Qwen3.8-27B: each label is the information-gathering step Qwen actually took next, built by code from Qwen sessions, with no other model in the loop.
- Test set: Whole tasks are held out, and the evaluation uses only held-out tasks of six benchmarks: SWE-bench Verified in full (500 tasks; none of its repositories is used for training), Terminal-Bench 2.0 40 of its 89 tasks (45 train; the other 4 are near-twins of evaluation tasks and are used for neither), SWE-rebench 189 tasks (671 train; split by repository), Terminal-Bench Pro 100 (90 train), SkillsBench 44 (42 train) and Harbor Index 41 (30 train). A leak check compares every training task with the evaluation tasks.
- Training data: not published.
- How it works: Jeff works ahead of the large model. Before each of Qwen's turns, Jeff decides whether it can already take the next information-gathering step itself: read a file, list a folder, search the code, check which tools are installed. It picks the tool first, then its argument, and can take up to 8 steps in a row. Qwen then starts its turn with those results already in front of it. It can also run the tests or a build, repeat Qwen's last command and, with the run-approval setting the evaluation used, run a script Qwen wrote or install a missing package. Writing and editing files always stay with Qwen, and whenever Jeff is unsure, it hands over. Because Jeff gives a calibrated probability for every choice, its step threshold is a single setting: Jeff takes a page's top option only when its probability is at least the threshold (0.40 in the measured runs), and hands over otherwise. A second adapter, code-router, decides on every turn whether Qwen should think hard.
- How we trained it: Jeff predicts what Qwen would do next. Each training label is the information-gathering step Qwen actually took next, built by code from Qwen sessions; if Jeff takes that step early, Qwen gets the result without spending a turn on it.
- The data comes in three stages: (1) about 14,400 public Qwen3.8-27B agent sessions (ukisai/Qwen3.8-27B-multi-turn-agent-sft), with Jeff's option lists rebuilt from each transcript; (2) public Terminal-Bench sessions of Qwen3.8-27B, replayed in each task's container; (3) our own sessions of Qwen3.8-27B in Jeff-Code, where Jeff logs its full option list before every turn without acting.
- Training order: the adapter was trained in curriculum order, stage 1, then stage 2, then stage 3, with stage 1 sampled down to half the size of stage 3.
- Questions longer than 8,192 tokens are cut to fit, keeping the head and tail of the state; the question and its options are never cut. The same rule applies at run time.
- Jeff only ever offers files, folders, programs or packages that have already appeared in the task or in earlier output, the same things Qwen could act on.
- To keep training and real use identical, Qwen works with bash only in Jeff-Code.
- Jeff-Code is a fork of the Pi coding agent (MIT), with its own changes such as a loop guard and Jeff acting first.
- Jeff-Code's measured results come entirely from LoRA adapters on the fixed Jeff v1.3 base - this one and code-router - both loaded beside the other adapters on one jeff-base.
Data and licence
Adapter: Apache-2.0
- ukisai/Qwen3.8-27B-multi-turn-agent-sft (about 14,400 public Qwen3.8-27B agent sessions)Generated by Qwen3.8-27B (the sessions); labels built by codeJeff's option lists are rebuilt from each transcript; the label is the step Qwen took next.
- openguardrails Terminal-Bench sessions of Qwen3.8-27BGenerated by Qwen3.8-27B (the sessions); labels built by codePublic sessions, replayed in each task's container.
- Terminal-Bench 2.0 tasks (outside the 40 evaluation tasks and their four near-twins)Not generated by a modelThe tasks our own Jeff-Code sessions and the replays run on.
- SWE-rebench tasks (the 671 training tasks; whole repositories held out, 189 tasks)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub swe-rebench/swe-rebench-leaderboard).
- Terminal-Bench Pro tasks (the 90 training tasks; 100 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub terminal-bench-pro/terminal-bench-pro).
- SkillsBench tasks (the 42 training tasks; 44 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub benchflow/skillsbench).
- Harbor Index tasks (the 30 training tasks; 41 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub harbor-index/harbor-index-1.0).
- Terminal-Bench tasks (the 27 training tasks of the 66-task set on the Harbor hub; 33 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub terminal-bench/terminal-bench).
- Terminal-Bench Science tasks (the 35 training tasks; 35 held out)Not generated by a modelThe tasks our own Jeff-Code sessions ran on (harbor hub terminal-bench-science/terminal-bench-science).
- Our own Jeff-Code sessions with Qwen3.8-27BBuilt for this adapterGenerated by Qwen3.8-27B (local), the sessions; labels built by codeQwen3.8-27B in Jeff-Code (bash only); Jeff logs its full option list before every turn without acting.
Download
- Weights: mstrasser/jeff-adapter-code
- LoRA GGUF: mstrasser/jeff-adapter-code-gguf
Through llama.cpp: the GGUF takes the same action (act on the top option at 0.40, otherwise hand over) as full precision on 99.9% of the 991 development rows at Q8_0, 97.9% at Q4_K_M. Running Jeff with llama.cpp
