Skip to content
JeffHub

Reference app

A support-inbox agent on one laptop (Apple M4 Max, 128 GB), run two ways: Qwen3.8-27B making every decision, and Jeff with adapters making the decisions while Qwen3.8-27B only writes the reply.

The pipeline

Each ticket passes five decisions before a reply is written. With Jeff, each decision is one request to the adapter named below. Each step has its own threshold: when Jeff is less sure than that, the decision goes on to Qwen3.8-27B. The thresholds are set so that Jeff only passes a decision on when that is needed to stay ahead.

  1. A new ticket arrives
  2. Step 1Is the ticket trying to take over the agent?
    Jeff + guardnever needs to pass on toQwen3.8-27B
  3. Step 2Which team, how urgent, and does a person need to see it?
    Jeff + triagenever needs to pass on toQwen3.8-27B
  4. Step 3What does the customer want?
    Jeff + support-intentsnever needs to pass on toQwen3.8-27B
  5. Step 4Which tool should the agent call next, or should it answer or ask?
    Jeff + toolsnever needs to pass on toQwen3.8-27B
  6. Step 5Which help-centre passages answer the question?
    Jeff + groundunder 64% sure →Qwen3.8-27B
  7. Write the replyQwen3.8-27B, in both setups

In the other setup, Qwen3.8-27B makes all five decisions itself, with the same prompts and options.

Results

Each step on a fixed sample of its adapter's held-out test set, answered both ways on the same rows. With Jeff first, the step is 39.0× faster on average, and its accuracy goes from 87.7% to 95.7%. The accuracy with Jeff includes the decisions passed on to Qwen3.8-27B, and so does the time.

  • guard38.1× faster
    Qwen3.8-27B alone
    84.0% · 3.92 s
    98.0% · 103 ms
    Gain over Qwen3.8-27B [95% interval]
    +14.0 points [+10.0, +18.3]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • triage59.1× faster
    Qwen3.8-27B alone
    81.3% · 3.60 s
    91.0% · 61 ms
    Gain over Qwen3.8-27B [95% interval]
    +9.7 points [+4.7, +14.7]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • support-intents54.6× faster
    Qwen3.8-27B alone
    86.0% · 6.44 s
    95.3% · 118 ms
    Gain over Qwen3.8-27B [95% interval]
    +9.3 points [+5.3, +13.3]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • tools36.5× faster
    Qwen3.8-27B alone
    90.3% · 11.27 s
    98.0% · 309 ms
    Gain over Qwen3.8-27B [95% interval]
    +7.7 points [+4.7, +11.0]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • ground20.1× faster
    Qwen3.8-27B alone
    96.7% · 13.22 s
    96.3% · 659 ms
    Gain over Qwen3.8-27B [95% interval]
    tie−0.3 points [−3.0, +2.3]
    Threshold · sent on to Qwen3.8-27B
    0.64 · 1.7% of 300 rows
  • Mean of 5 tasks: 87.7% → 95.7%, 39.0× faster

Times are the mean per query over the whole sample, prompt to answer. The higher accuracy in each row is in bold. Gain: accuracy of Jeff + adapter minus Qwen3.8-27B alone, with its 95% interval; “tie” when the interval includes zero. Each task has its own threshold: below it, Jeff passes the query on to Qwen3.8-27B. “Faster” for the mean row is the geometric mean of the per-task speed-ups.

On ground: about equal accuracy (96.3% against 96.7% on 300 rows), with Jeff 20× faster. Every other step is more accurate with Jeff.

Time per whole ticket (five decisions and the reply) is not measured yet. Every threshold, and how this was measured.

Watch it run

A short screen recording of both setups answering the same tickets will be here.

The code

The app will be published in the Jeff repository under examples/, with the scripts that produced every number on this page.