Reference app
A support-inbox agent on one laptop (Apple M4 Max, 128 GB), run two ways: Qwen3.8-27B making every decision, and Jeff with adapters making the decisions while Qwen3.8-27B only writes the reply.
The pipeline
Each ticket passes five decisions before a reply is written. With Jeff, each decision is one request to the adapter named below. Each step has its own threshold: when Jeff is less sure than that, the decision goes on to Qwen3.8-27B. The thresholds are set so that Jeff only passes a decision on when that is needed to stay ahead.
- A new ticket arrives
- Step 1Is the ticket trying to take over the agent?
- Step 2Which team, how urgent, and does a person need to see it?
- Step 3What does the customer want?
- Step 4Which tool should the agent call next, or should it answer or ask?
- Step 5Which help-centre passages answer the question?
- Write the replyQwen3.8-27B, in both setups
In the other setup, Qwen3.8-27B makes all five decisions itself, with the same prompts and options.
Results
Each step on a fixed sample of its adapter's held-out test set, answered both ways on the same rows. With Jeff first, the step is 39.0× faster on average, and its accuracy goes from 87.7% to 95.7%. The accuracy with Jeff includes the decisions passed on to Qwen3.8-27B, and so does the time.
| Task | Rows | Qwen3.8-27B alone: accuracy | time per query | Jeff + adapter first: accuracy | time per query | Gain [95% interval] | Threshold | Sent on to Qwen3.8-27B | Faster |
|---|---|---|---|---|---|---|---|---|---|
guard | 300 | 84.0% | 3.92 s | 98.0% | 103 ms | +14.0 points [+10.0, +18.3] | 0.00 | 0.0% | 38.1× |
triage | 300 | 81.3% | 3.60 s | 91.0% | 61 ms | +9.7 points [+4.7, +14.7] | 0.00 | 0.0% | 59.1× |
support-intents | 300 | 86.0% | 6.44 s | 95.3% | 118 ms | +9.3 points [+5.3, +13.3] | 0.00 | 0.0% | 54.6× |
tools | 300 | 90.3% | 11.27 s | 98.0% | 309 ms | +7.7 points [+4.7, +11.0] | 0.00 | 0.0% | 36.5× |
ground | 300 | 96.7% | 13.22 s | 96.3% | 659 ms | tie−0.3 points [−3.0, +2.3] | 0.64 | 1.7% | 20.1× |
| Mean of 5 tasks | 87.7% | 95.7% | 39.0× |
guard38.1× faster- Qwen3.8-27B alone
- 84.0% · 3.92 s
- Jeff + adapter first
- 98.0% · 103 ms
- Gain over Qwen3.8-27B [95% interval]
- +14.0 points [+10.0, +18.3]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
triage59.1× faster- Qwen3.8-27B alone
- 81.3% · 3.60 s
- Jeff + adapter first
- 91.0% · 61 ms
- Gain over Qwen3.8-27B [95% interval]
- +9.7 points [+4.7, +14.7]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
support-intents54.6× faster- Qwen3.8-27B alone
- 86.0% · 6.44 s
- Jeff + adapter first
- 95.3% · 118 ms
- Gain over Qwen3.8-27B [95% interval]
- +9.3 points [+5.3, +13.3]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
tools36.5× faster- Qwen3.8-27B alone
- 90.3% · 11.27 s
- Jeff + adapter first
- 98.0% · 309 ms
- Gain over Qwen3.8-27B [95% interval]
- +7.7 points [+4.7, +11.0]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
ground20.1× faster- Qwen3.8-27B alone
- 96.7% · 13.22 s
- Jeff + adapter first
- 96.3% · 659 ms
- Gain over Qwen3.8-27B [95% interval]
- tie−0.3 points [−3.0, +2.3]
- Threshold · sent on to Qwen3.8-27B
- 0.64 · 1.7% of 300 rows
- Mean of 5 tasks: 87.7% → 95.7%, 39.0× faster
Times are the mean per query over the whole sample, prompt to answer. The higher accuracy in each row is in bold. Gain: accuracy of Jeff + adapter minus Qwen3.8-27B alone, with its 95% interval; “tie” when the interval includes zero. Each task has its own threshold: below it, Jeff passes the query on to Qwen3.8-27B. “Faster” for the mean row is the geometric mean of the per-task speed-ups.
On ground: about equal accuracy (96.3% against 96.7% on 300 rows), with Jeff 20× faster. Every other step is more accurate with Jeff.
Time per whole ticket (five decisions and the reply) is not measured yet. Every threshold, and how this was measured.
Watch it run
The code
The app will be published in the Jeff repository under examples/, with the scripts that produced every number on this page.
