Skip to content
JeffHub

Results

What Jeff and its adapters do, measured: every adapter on its own held-out test set, Jeff + adapters in front of Qwen3.8-27B on the same test rows, and how fast and small the server is. Numbers as of 2026-10-01; every one names where it came from.

55.6% → 91.5%
mean accuracy of Jeff alone → Jeff + adapter, over 9 adapters
86.6% → 95.3%
accuracy: Qwen3.8-27B alone → Jeff + adapters; gain +8.7 points [+7.4, +10.1]
37.7× faster
than Qwen3.8-27B alone, 8 adapters
about 30 ms
per decision with an adapter (median, NVIDIA RTX PRO 6000)

What each adapter adds

Each adapter is scored on its full held-out test set, which it never trained on, three ways: the untrained Qwen3.5-0.8B that Jeff is built from, Jeff v1.2 0.8B alone, and Jeff v1.2 0.8B with the adapter. Across the 9 adapters, mean accuracy goes from 29.5% untrained to 55.6% for Jeff alone and 91.5% with the adapter.

Accuracy is the share of test questions answered correctly. Calibration error (ECE) measures how far the stated confidence is from the real hit rate: 0.01 means the stated confidence is, on average, about 1 percentage point away from how often those answers are right. What calibration means, and why it matters.

  • guard6,552 test rows
    Qwen3.5-0.8B untrained
    43.7% · 0.065
    Jeff v1.2 0.8B alone
    46.9% · 0.280
    98.4% · 0.004
  • triage7,256 test rows
    Qwen3.5-0.8B untrained
    44.4% · 0.107
    Jeff v1.2 0.8B alone
    67.1% · 0.026
    91.8% · 0.009
  • support-intents5,577 test rows
    Qwen3.5-0.8B untrained
    33.9% · 0.162
    Jeff v1.2 0.8B alone
    85.1% · 0.080
    96.8% · 0.006
  • tools5,157 test rows
    Qwen3.5-0.8B untrained
    18.0% · 0.064
    Jeff v1.2 0.8B alone
    57.8% · 0.091
    97.9% · 0.004
  • ground4,160 test rows
    Qwen3.5-0.8B untrained
    28.7% · 0.061
    Jeff v1.2 0.8B alone
    49.0% · 0.237
    97.0% · 0.012
  • nav3,300 test rows
    Qwen3.5-0.8B untrained
    12.6% · 0.038
    Jeff v1.2 0.8B alone
    23.8% · 0.195
    97.0% · 0.005
  • emotion5,408 test rows
    Qwen3.5-0.8B untrained
    12.5% · 0.044
    Jeff v1.2 0.8B alone
    32.2% · 0.038
    60.6% · 0.020
  • spam3,603 test rows
    Qwen3.5-0.8B untrained
    59.4% · 0.060
    Jeff v1.2 0.8B alone
    72.3% · 0.109
    98.4% · 0.008
  • legal-clauses9,895 test rows
    Qwen3.5-0.8B untrained
    12.5% · 0.094
    Jeff v1.2 0.8B alone
    66.0% · 0.032
    85.7% · 0.011

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

Accuracy, step by stepQwen3.5-0.8B untrainedAdded by Jeff's trainingAdded by the adapter
  • guard43.7% → 46.9% → 98.4%
  • triage44.4% → 67.1% → 91.8%
  • support-intents33.9% → 85.1% → 96.8%
  • tools18.0% → 57.8% → 97.9%
  • ground28.7% → 49.0% → 97.0%
  • nav12.6% → 23.8% → 97.0%
  • emotion12.5% → 32.2% → 60.6%
  • spam59.4% → 72.3% → 98.4%
  • legal-clauses12.5% → 66.0% → 85.7%

Scale 0 to 100% accuracy on each adapter's own held-out test set.

Jeff + adapters against Qwen3.8-27B

On 8 of 9 tasks, Jeff with its adapter beats the 27B on its own; the 27B stays as a fallback you can tune per task, and for anything no adapter covers.

The adapters do the work: with the thresholds below, Jeff sends no query at all to Qwen3.8-27B on 8 of the 9 tasks. Both setups answered the same fixed random sample of each task's held-out rows on one Apple M4 Max, 128 GB; both models with MLX, one at a time.

  • Qwen3.8-27B alone: the 27B answers every query. The model: Qwen3.8-27B, 8-bit MLX weights, prompted, step-by-step reasoning off; it uses 28.6 GB.
  • Jeff + adapter first: Jeff + the task's adapter answers first; when its confidence is below the threshold, the 27B answers too and its answer is used (time = Jeff + 27B).
  • The threshold: each task has its own, chosen on each task's own calibration rows: the lowest threshold (fastest) whose accuracy beats the 27B alone by at least 1 point, and then measured on the test rows. In plain words: Jeff only passes a query on to Qwen3.8-27B when that is needed to stay ahead.
  • Timing: each task's fixed random sample of held-out rows; the same rows for both routes; times are the mean per query over the whole sample, prompt to answer.
  • How certain: each gain has a 95% interval from a paired bootstrap over the test rows, 10,000 resamples per adapter (2,000 for the mean). “Tie” means the interval includes zero.
  • Sample size: Qwen3.8-27B takes seconds per decision, so the comparison uses a fixed random sample of 300 rows per task (500 for emotion and legal-clauses). Jeff + adapter was also scored on each full test set (3,300 to 9,895 rows), and agrees with the sample to within 2.1 points on every adapter.

8 adapters

The headline is the mean over these 8 adapters, each at its own threshold; emotion is shown separately below.

  • guard38.1× faster
    Qwen3.8-27B alone
    84.0% · 3.92 s
    98.0% · 103 ms
    Gain over Qwen3.8-27B [95% interval]
    +14.0 points [+10.0, +18.3]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • triage59.1× faster
    Qwen3.8-27B alone
    81.3% · 3.60 s
    91.0% · 61 ms
    Gain over Qwen3.8-27B [95% interval]
    +9.7 points [+4.7, +14.7]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • support-intents54.6× faster
    Qwen3.8-27B alone
    86.0% · 6.44 s
    95.3% · 118 ms
    Gain over Qwen3.8-27B [95% interval]
    +9.3 points [+5.3, +13.3]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • tools36.5× faster
    Qwen3.8-27B alone
    90.3% · 11.27 s
    98.0% · 309 ms
    Gain over Qwen3.8-27B [95% interval]
    +7.7 points [+4.7, +11.0]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • ground20.1× faster
    Qwen3.8-27B alone
    96.7% · 13.22 s
    96.3% · 659 ms
    Gain over Qwen3.8-27B [95% interval]
    tie−0.3 points [−3.0, +2.3]
    Threshold · sent on to Qwen3.8-27B
    0.64 · 1.7% of 300 rows
  • nav34.6× faster
    Qwen3.8-27B alone
    91.3% · 7.21 s
    97.0% · 208 ms
    Gain over Qwen3.8-27B [95% interval]
    +5.7 points [+2.7, +9.0]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • spam36.6× faster
    Qwen3.8-27B alone
    88.0% · 2.67 s
    98.7% · 73 ms
    Gain over Qwen3.8-27B [95% interval]
    +10.7 points [+7.0, +14.7]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 300 rows
  • legal-clauses35.5× faster
    Qwen3.8-27B alone
    75.0% · 16.53 s
    87.8% · 466 ms
    Gain over Qwen3.8-27B [95% interval]
    +12.8 points [+9.2, +16.4]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 500 rows
  • Mean of 8 tasks: 86.6% → 95.3%, 37.7× faster; gain +8.7 points [+7.4, +10.1]

Times are the mean per query over the whole sample, prompt to answer. The higher accuracy in each row is in bold. Gain: accuracy of Jeff + adapter minus Qwen3.8-27B alone, with its 95% interval; “tie” when the interval includes zero. Each task has its own threshold: below it, Jeff passes the query on to Qwen3.8-27B. “Faster” for the mean row is the geometric mean of the per-task speed-ups.

Where the gain comes from

The same 8 adapters, the same rows, four setups. Against Qwen3.8-27B alone, the adapters on their own add 8.6 points of accuracy; routing on its own adds 0.5.

Accuracy and time per decision for each setup, mean over 8 adapters
SetupAccuracyTime
Qwen3.8-27B aloneanswers every query86.6%8.11 s
Adapters onlyJeff + adapter; never asks Qwen3.8-27B95.2%212 ms
Routing onlyJeff without adapters; unsure queries go to Qwen3.8-27B87.1%6.99 s
Both (the headline)Jeff + adapter; Qwen3.8-27B as a fallback95.3%250 ms

“Routing only” had its threshold picked on the test rows, because plain Jeff has no calibration rows, so its result is flattered.

ground is the only task that passes queries on to Qwen3.8-27B: when Jeff is less than 64% sure (threshold 0.64), which happened for 1.7% of its queries. On the test rows it scores 96.3% against Qwen3.8-27B's 96.7%: one question in 300, within noise, and the rule was fixed in advance on the calibration rows, not tuned on the test. So: about equal accuracy, 20× faster.

Every other task sends nothing to Qwen3.8-27B and is more accurate.

Left out of the averages

  • emotion42.5× faster
    Qwen3.8-27B alone
    35.6% · 4.79 s
    60.6% · 113 ms
    Gain over Qwen3.8-27B [95% interval]
    +25.0 points [+20.0, +30.0]
    Threshold · sent on to Qwen3.8-27B
    0.00 · 0.0% of 500 rows

Times are the mean per query over the whole sample, prompt to answer. The higher accuracy in each row is in bold. Gain: accuracy of Jeff + adapter minus Qwen3.8-27B alone, with its 95% interval; “tie” when the interval includes zero. Each task has its own threshold: below it, Jeff passes the query on to Qwen3.8-27B.

Emotion is left out of the averages: picking the single strongest of 27 emotions (or neutral) in short Reddit comments is hard even for people, and the human labels often disagree. Jeff + adapter scores 60.6% there against the 27B's 35.6%, at 42× the speed. Including it, the average across all adapters is 91.4% against 80.9%, so leaving it out makes the gain smaller, not larger.

One shared threshold, for comparison

The threshold trades speed for care. At 0 Jeff answers everything; at 1 Qwen3.8-27B answers everything. Here the same threshold is applied to every task at once, to show the whole dial; the dot marks 0.51, the single threshold used before each task got its own.

Shared threshold 0.51: accuracy 95.8% on the test rows (96.9% on the calibration rows), 30.5× faster, 0.9% sent on to Qwen3.8-27B.

Accuracy95.8%
85.0%92.5%100.0%0.000.250.500.751.00
Speed-up30.5×
0.0×25.0×50.0×0.000.250.500.751.00
Sent on to Qwen3.8-27B0.9%
0.0%50.0%100.0%0.000.250.500.751.00

Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the shared threshold. Accuracy is the mean over the tasks on their test rows; speed-up the geometric mean. The full table is below. Point at a chart to read any threshold.

Show the thresholds as a table
Accuracy, speed-up and share sent to Qwen3.8-27B at each threshold
ThresholdCalibration accuracyTest accuracySent to Qwen3.8-27BFaster
0.0096.8%95.6%0.0%44.1×
0.0596.8%95.6%0.0%44.1×
0.1096.8%95.6%0.0%44.1×
0.1596.8%95.6%0.0%44.1×
0.2096.8%95.6%0.0%44.1×
0.2596.8%95.6%0.0%44.1×
0.3096.8%95.6%0.0%44.1×
0.3596.8%95.6%0.0%44.1×
0.4096.8%95.6%0.0%44.1×
0.4596.8%95.7%0.1%42.2×
0.5096.8%95.7%0.6%33.0×
0.51 (shared)96.9%95.8%0.9%30.5×
0.5596.8%95.9%1.6%25.6×
0.6096.7%95.9%2.4%20.5×
0.6596.7%95.7%3.1%18.3×
0.7096.5%95.9%4.5%16.1×
0.7596.1%95.9%5.7%14.0×
0.8095.7%95.5%7.6%11.4×
0.8595.7%95.3%9.3%9.9×
0.9095.5%95.2%11.7%8.2×
0.9595.2%94.7%15.1%6.3×
1.0088.7%87.7%100.0%1.0×

Every fifth threshold of the 101 measured, plus the chosen one.

The base models

Jeff itself, before any adapter, on its general benchmark of 4,599 questions. The adapters on this site are trained on the 0.8B model.

Base models on the general benchmark
ModelQuestionsAccuracyCalibration errorRun
Jeff v1.2 0.8B4,59978.7%0.0280.8b-20260929-2258
Jeff v1.2 2B4,59981.7%0.0212b-20260930-2347

Serving speed and memory

One Jeff server on an NVIDIA RTX PRO 6000: the base model loaded once, adapters beside it, one decision per request. With the default serving (LoRA in bfloat16), a decision with an adapter takes a median of 31.2 ms, against 25.9 ms for the base alone. With all nine adapters loaded and a different one on every request, it is 30.0 ms: switching costs nothing measurable. Merged mode folds one adapter into the weights and runs at the base's speed (25.7 ms), but then serves only that adapter.

Time per decision and GPU memory by serving setting, NVIDIA RTX PRO 6000
SettingAdapters loadedMedian95th percentileGPU memory loadedPeak GPU memory
base alone025.9 ms35.5 ms1.74 GB2.30 GB
base + one adapter (LoRA in bfloat16, the default)131.2 ms40.3 ms1.79 GB2.32 GB
base + all nine adapters, switching adapter on every request930.0 ms39.4 ms1.96 GB2.49 GB
one adapter merged into the weights (JEFF_ADAPTER_MODE=merged)125.7 ms35.7 ms1.77 GB2.30 GB
base + one adapter (LoRA in float32)134.1 ms47.4 ms1.81 GB2.34 GB
base + all nine adapters in float32, switching every request934.4 ms47.0 ms2.14 GB2.67 GB

675 requests per setting. For comparison, Qwen3.8-27B uses 28.6 GB on the Mac; Jeff's memory on the Mac is not measured yet.

Not measured yet

Nothing on this site is estimated. These are still to come:

  • Jeff's memory on the Mac (measured on the RTX only)
  • end-to-end ticket times (not run yet)

Where the numbers came from

  • Adapters and serving: jeff-finetunes/adapters/BASELINE.md
  • Against Qwen3.8-27B: jeff-reference-app results/cascade.json
  • Jeff v1.2 0.8B: jev/runs/eval/0.8b-20260929-2258-final-calibrated.json
  • Jeff v1.2 2B: jev/runs/eval/2b-20260930-2347-final-calibrated.json

Collected into one file on 2026-10-01T11:56:57+00:00.