Results
What Jeff and its adapters do, measured: every adapter on its own held-out test set, Jeff + adapters in front of Qwen3.8-27B on the same test rows, and how fast and small the server is. Numbers as of 2026-10-01; every one names where it came from.
- 55.6% → 91.5%
- mean accuracy of Jeff alone → Jeff + adapter, over 9 adapters
- 86.6% → 95.3%
- accuracy: Qwen3.8-27B alone → Jeff + adapters; gain +8.7 points [+7.4, +10.1]
- 37.7× faster
- than Qwen3.8-27B alone, 8 adapters
- about 30 ms
- per decision with an adapter (median, NVIDIA RTX PRO 6000)
What each adapter adds
Each adapter is scored on its full held-out test set, which it never trained on, three ways: the untrained Qwen3.5-0.8B that Jeff is built from, Jeff v1.2 0.8B alone, and Jeff v1.2 0.8B with the adapter. Across the 9 adapters, mean accuracy goes from 29.5% untrained to 55.6% for Jeff alone and 91.5% with the adapter.
Accuracy is the share of test questions answered correctly. Calibration error (ECE) measures how far the stated confidence is from the real hit rate: 0.01 means the stated confidence is, on average, about 1 percentage point away from how often those answers are right. What calibration means, and why it matters.
| Adapter test set | Test rows | Qwen3.5-0.8B untrained | Jeff v1.2 0.8B alone | Jeff v1.2 0.8B + adapter |
|---|---|---|---|---|
guard | 6,552 | 43.7% · 0.065 | 46.9% · 0.280 | 98.4% · 0.004 |
triage | 7,256 | 44.4% · 0.107 | 67.1% · 0.026 | 91.8% · 0.009 |
support-intents | 5,577 | 33.9% · 0.162 | 85.1% · 0.080 | 96.8% · 0.006 |
tools | 5,157 | 18.0% · 0.064 | 57.8% · 0.091 | 97.9% · 0.004 |
ground | 4,160 | 28.7% · 0.061 | 49.0% · 0.237 | 97.0% · 0.012 |
nav | 3,300 | 12.6% · 0.038 | 23.8% · 0.195 | 97.0% · 0.005 |
emotion | 5,408 | 12.5% · 0.044 | 32.2% · 0.038 | 60.6% · 0.020 |
spam | 3,603 | 59.4% · 0.060 | 72.3% · 0.109 | 98.4% · 0.008 |
legal-clauses | 9,895 | 12.5% · 0.094 | 66.0% · 0.032 | 85.7% · 0.011 |
guard6,552 test rows- Qwen3.5-0.8B untrained
- 43.7% · 0.065
- Jeff v1.2 0.8B alone
- 46.9% · 0.280
- Jeff v1.2 0.8B + adapter
- 98.4% · 0.004
triage7,256 test rows- Qwen3.5-0.8B untrained
- 44.4% · 0.107
- Jeff v1.2 0.8B alone
- 67.1% · 0.026
- Jeff v1.2 0.8B + adapter
- 91.8% · 0.009
support-intents5,577 test rows- Qwen3.5-0.8B untrained
- 33.9% · 0.162
- Jeff v1.2 0.8B alone
- 85.1% · 0.080
- Jeff v1.2 0.8B + adapter
- 96.8% · 0.006
tools5,157 test rows- Qwen3.5-0.8B untrained
- 18.0% · 0.064
- Jeff v1.2 0.8B alone
- 57.8% · 0.091
- Jeff v1.2 0.8B + adapter
- 97.9% · 0.004
ground4,160 test rows- Qwen3.5-0.8B untrained
- 28.7% · 0.061
- Jeff v1.2 0.8B alone
- 49.0% · 0.237
- Jeff v1.2 0.8B + adapter
- 97.0% · 0.012
nav3,300 test rows- Qwen3.5-0.8B untrained
- 12.6% · 0.038
- Jeff v1.2 0.8B alone
- 23.8% · 0.195
- Jeff v1.2 0.8B + adapter
- 97.0% · 0.005
emotion5,408 test rows- Qwen3.5-0.8B untrained
- 12.5% · 0.044
- Jeff v1.2 0.8B alone
- 32.2% · 0.038
- Jeff v1.2 0.8B + adapter
- 60.6% · 0.020
spam3,603 test rows- Qwen3.5-0.8B untrained
- 59.4% · 0.060
- Jeff v1.2 0.8B alone
- 72.3% · 0.109
- Jeff v1.2 0.8B + adapter
- 98.4% · 0.008
legal-clauses9,895 test rows- Qwen3.5-0.8B untrained
- 12.5% · 0.094
- Jeff v1.2 0.8B alone
- 66.0% · 0.032
- Jeff v1.2 0.8B + adapter
- 85.7% · 0.011
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
guard43.7% → 46.9% → 98.4%triage44.4% → 67.1% → 91.8%support-intents33.9% → 85.1% → 96.8%tools18.0% → 57.8% → 97.9%ground28.7% → 49.0% → 97.0%nav12.6% → 23.8% → 97.0%emotion12.5% → 32.2% → 60.6%spam59.4% → 72.3% → 98.4%legal-clauses12.5% → 66.0% → 85.7%
Scale 0 to 100% accuracy on each adapter's own held-out test set.
Jeff + adapters against Qwen3.8-27B
On 8 of 9 tasks, Jeff with its adapter beats the 27B on its own; the 27B stays as a fallback you can tune per task, and for anything no adapter covers.
The adapters do the work: with the thresholds below, Jeff sends no query at all to Qwen3.8-27B on 8 of the 9 tasks. Both setups answered the same fixed random sample of each task's held-out rows on one Apple M4 Max, 128 GB; both models with MLX, one at a time.
- Qwen3.8-27B alone: the 27B answers every query. The model: Qwen3.8-27B, 8-bit MLX weights, prompted, step-by-step reasoning off; it uses 28.6 GB.
- Jeff + adapter first: Jeff + the task's adapter answers first; when its confidence is below the threshold, the 27B answers too and its answer is used (time = Jeff + 27B).
- The threshold: each task has its own, chosen on each task's own calibration rows: the lowest threshold (fastest) whose accuracy beats the 27B alone by at least 1 point, and then measured on the test rows. In plain words: Jeff only passes a query on to Qwen3.8-27B when that is needed to stay ahead.
- Timing: each task's fixed random sample of held-out rows; the same rows for both routes; times are the mean per query over the whole sample, prompt to answer.
- How certain: each gain has a 95% interval from a paired bootstrap over the test rows, 10,000 resamples per adapter (2,000 for the mean). “Tie” means the interval includes zero.
- Sample size: Qwen3.8-27B takes seconds per decision, so the comparison uses a fixed random sample of 300 rows per task (500 for emotion and legal-clauses). Jeff + adapter was also scored on each full test set (3,300 to 9,895 rows), and agrees with the sample to within 2.1 points on every adapter.
8 adapters
The headline is the mean over these 8 adapters, each at its own threshold; emotion is shown separately below.
| Task | Rows | Qwen3.8-27B alone: accuracy | time per query | Jeff + adapter first: accuracy | time per query | Gain [95% interval] | Threshold | Sent on to Qwen3.8-27B | Faster |
|---|---|---|---|---|---|---|---|---|---|
guard | 300 | 84.0% | 3.92 s | 98.0% | 103 ms | +14.0 points [+10.0, +18.3] | 0.00 | 0.0% | 38.1× |
triage | 300 | 81.3% | 3.60 s | 91.0% | 61 ms | +9.7 points [+4.7, +14.7] | 0.00 | 0.0% | 59.1× |
support-intents | 300 | 86.0% | 6.44 s | 95.3% | 118 ms | +9.3 points [+5.3, +13.3] | 0.00 | 0.0% | 54.6× |
tools | 300 | 90.3% | 11.27 s | 98.0% | 309 ms | +7.7 points [+4.7, +11.0] | 0.00 | 0.0% | 36.5× |
ground | 300 | 96.7% | 13.22 s | 96.3% | 659 ms | tie−0.3 points [−3.0, +2.3] | 0.64 | 1.7% | 20.1× |
nav | 300 | 91.3% | 7.21 s | 97.0% | 208 ms | +5.7 points [+2.7, +9.0] | 0.00 | 0.0% | 34.6× |
spam | 300 | 88.0% | 2.67 s | 98.7% | 73 ms | +10.7 points [+7.0, +14.7] | 0.00 | 0.0% | 36.6× |
legal-clauses | 500 | 75.0% | 16.53 s | 87.8% | 466 ms | +12.8 points [+9.2, +16.4] | 0.00 | 0.0% | 35.5× |
| Mean of 8 tasks | 86.6% | 95.3% | +8.7 points [+7.4, +10.1] | 37.7× |
guard38.1× faster- Qwen3.8-27B alone
- 84.0% · 3.92 s
- Jeff + adapter first
- 98.0% · 103 ms
- Gain over Qwen3.8-27B [95% interval]
- +14.0 points [+10.0, +18.3]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
triage59.1× faster- Qwen3.8-27B alone
- 81.3% · 3.60 s
- Jeff + adapter first
- 91.0% · 61 ms
- Gain over Qwen3.8-27B [95% interval]
- +9.7 points [+4.7, +14.7]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
support-intents54.6× faster- Qwen3.8-27B alone
- 86.0% · 6.44 s
- Jeff + adapter first
- 95.3% · 118 ms
- Gain over Qwen3.8-27B [95% interval]
- +9.3 points [+5.3, +13.3]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
tools36.5× faster- Qwen3.8-27B alone
- 90.3% · 11.27 s
- Jeff + adapter first
- 98.0% · 309 ms
- Gain over Qwen3.8-27B [95% interval]
- +7.7 points [+4.7, +11.0]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
ground20.1× faster- Qwen3.8-27B alone
- 96.7% · 13.22 s
- Jeff + adapter first
- 96.3% · 659 ms
- Gain over Qwen3.8-27B [95% interval]
- tie−0.3 points [−3.0, +2.3]
- Threshold · sent on to Qwen3.8-27B
- 0.64 · 1.7% of 300 rows
nav34.6× faster- Qwen3.8-27B alone
- 91.3% · 7.21 s
- Jeff + adapter first
- 97.0% · 208 ms
- Gain over Qwen3.8-27B [95% interval]
- +5.7 points [+2.7, +9.0]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
spam36.6× faster- Qwen3.8-27B alone
- 88.0% · 2.67 s
- Jeff + adapter first
- 98.7% · 73 ms
- Gain over Qwen3.8-27B [95% interval]
- +10.7 points [+7.0, +14.7]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 300 rows
legal-clauses35.5× faster- Qwen3.8-27B alone
- 75.0% · 16.53 s
- Jeff + adapter first
- 87.8% · 466 ms
- Gain over Qwen3.8-27B [95% interval]
- +12.8 points [+9.2, +16.4]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 500 rows
- Mean of 8 tasks: 86.6% → 95.3%, 37.7× faster; gain +8.7 points [+7.4, +10.1]
Times are the mean per query over the whole sample, prompt to answer. The higher accuracy in each row is in bold. Gain: accuracy of Jeff + adapter minus Qwen3.8-27B alone, with its 95% interval; “tie” when the interval includes zero. Each task has its own threshold: below it, Jeff passes the query on to Qwen3.8-27B. “Faster” for the mean row is the geometric mean of the per-task speed-ups.
Where the gain comes from
The same 8 adapters, the same rows, four setups. Against Qwen3.8-27B alone, the adapters on their own add 8.6 points of accuracy; routing on its own adds 0.5.
| Setup | Accuracy | TimeTime per decision |
|---|---|---|
| Qwen3.8-27B aloneanswers every query | 86.6% | 8.11 s |
| Adapters onlyJeff + adapter; never asks Qwen3.8-27B | 95.2% | 212 ms |
| Routing onlyJeff without adapters; unsure queries go to Qwen3.8-27B | 87.1% | 6.99 s |
| Both (the headline)Jeff + adapter; Qwen3.8-27B as a fallback | 95.3% | 250 ms |
“Routing only” had its threshold picked on the test rows, because plain Jeff has no calibration rows, so its result is flattered.
ground is the only task that passes queries on to Qwen3.8-27B: when Jeff is less than 64% sure (threshold 0.64), which happened for 1.7% of its queries. On the test rows it scores 96.3% against Qwen3.8-27B's 96.7%: one question in 300, within noise, and the rule was fixed in advance on the calibration rows, not tuned on the test. So: about equal accuracy, 20× faster.
Every other task sends nothing to Qwen3.8-27B and is more accurate.
Left out of the averages
| Task | Rows | Qwen3.8-27B alone: accuracy | time per query | Jeff + adapter first: accuracy | time per query | Gain [95% interval] | Threshold | Sent on to Qwen3.8-27B | Faster |
|---|---|---|---|---|---|---|---|---|---|
emotion | 500 | 35.6% | 4.79 s | 60.6% | 113 ms | +25.0 points [+20.0, +30.0] | 0.00 | 0.0% | 42.5× |
emotion42.5× faster- Qwen3.8-27B alone
- 35.6% · 4.79 s
- Jeff + adapter first
- 60.6% · 113 ms
- Gain over Qwen3.8-27B [95% interval]
- +25.0 points [+20.0, +30.0]
- Threshold · sent on to Qwen3.8-27B
- 0.00 · 0.0% of 500 rows
Times are the mean per query over the whole sample, prompt to answer. The higher accuracy in each row is in bold. Gain: accuracy of Jeff + adapter minus Qwen3.8-27B alone, with its 95% interval; “tie” when the interval includes zero. Each task has its own threshold: below it, Jeff passes the query on to Qwen3.8-27B.
Emotion is left out of the averages: picking the single strongest of 27 emotions (or neutral) in short Reddit comments is hard even for people, and the human labels often disagree. Jeff + adapter scores 60.6% there against the 27B's 35.6%, at 42× the speed. Including it, the average across all adapters is 91.4% against 80.9%, so leaving it out makes the gain smaller, not larger.
One shared threshold, for comparison
The threshold trades speed for care. At 0 Jeff answers everything; at 1 Qwen3.8-27B answers everything. Here the same threshold is applied to every task at once, to show the whole dial; the dot marks 0.51, the single threshold used before each task got its own.
Shared threshold 0.51: accuracy 95.8% on the test rows (96.9% on the calibration rows), 30.5× faster, 0.9% sent on to Qwen3.8-27B.
Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the shared threshold. Accuracy is the mean over the tasks on their test rows; speed-up the geometric mean. The full table is below. Point at a chart to read any threshold.
Show the thresholds as a table
| Threshold | Calibration accuracy | Test accuracy | Sent to Qwen3.8-27B | Faster |
|---|---|---|---|---|
| 0.00 | 96.8% | 95.6% | 0.0% | 44.1× |
| 0.05 | 96.8% | 95.6% | 0.0% | 44.1× |
| 0.10 | 96.8% | 95.6% | 0.0% | 44.1× |
| 0.15 | 96.8% | 95.6% | 0.0% | 44.1× |
| 0.20 | 96.8% | 95.6% | 0.0% | 44.1× |
| 0.25 | 96.8% | 95.6% | 0.0% | 44.1× |
| 0.30 | 96.8% | 95.6% | 0.0% | 44.1× |
| 0.35 | 96.8% | 95.6% | 0.0% | 44.1× |
| 0.40 | 96.8% | 95.6% | 0.0% | 44.1× |
| 0.45 | 96.8% | 95.7% | 0.1% | 42.2× |
| 0.50 | 96.8% | 95.7% | 0.6% | 33.0× |
| 0.51 (shared) | 96.9% | 95.8% | 0.9% | 30.5× |
| 0.55 | 96.8% | 95.9% | 1.6% | 25.6× |
| 0.60 | 96.7% | 95.9% | 2.4% | 20.5× |
| 0.65 | 96.7% | 95.7% | 3.1% | 18.3× |
| 0.70 | 96.5% | 95.9% | 4.5% | 16.1× |
| 0.75 | 96.1% | 95.9% | 5.7% | 14.0× |
| 0.80 | 95.7% | 95.5% | 7.6% | 11.4× |
| 0.85 | 95.7% | 95.3% | 9.3% | 9.9× |
| 0.90 | 95.5% | 95.2% | 11.7% | 8.2× |
| 0.95 | 95.2% | 94.7% | 15.1% | 6.3× |
| 1.00 | 88.7% | 87.7% | 100.0% | 1.0× |
Every fifth threshold of the 101 measured, plus the chosen one.
The base models
Jeff itself, before any adapter, on its general benchmark of 4,599 questions. The adapters on this site are trained on the 0.8B model.
| Model | Questions | Accuracy | Calibration error | Run |
|---|---|---|---|---|
| Jeff v1.2 0.8B | 4,599 | 78.7% | 0.028 | 0.8b-20260929-2258 |
| Jeff v1.2 2B | 4,599 | 81.7% | 0.021 | 2b-20260930-2347 |
Serving speed and memory
One Jeff server on an NVIDIA RTX PRO 6000: the base model loaded once, adapters beside it, one decision per request. With the default serving (LoRA in bfloat16), a decision with an adapter takes a median of 31.2 ms, against 25.9 ms for the base alone. With all nine adapters loaded and a different one on every request, it is 30.0 ms: switching costs nothing measurable. Merged mode folds one adapter into the weights and runs at the base's speed (25.7 ms), but then serves only that adapter.
| Setting | Adapters loaded | Median | 95th percentile | GPU memory loaded | Peak GPU memory |
|---|---|---|---|---|---|
| base alone | 0 | 25.9 ms | 35.5 ms | 1.74 GB | 2.30 GB |
| base + one adapter (LoRA in bfloat16, the default) | 1 | 31.2 ms | 40.3 ms | 1.79 GB | 2.32 GB |
| base + all nine adapters, switching adapter on every request | 9 | 30.0 ms | 39.4 ms | 1.96 GB | 2.49 GB |
| one adapter merged into the weights (JEFF_ADAPTER_MODE=merged) | 1 | 25.7 ms | 35.7 ms | 1.77 GB | 2.30 GB |
| base + one adapter (LoRA in float32) | 1 | 34.1 ms | 47.4 ms | 1.81 GB | 2.34 GB |
| base + all nine adapters in float32, switching every request | 9 | 34.4 ms | 47.0 ms | 2.14 GB | 2.67 GB |
675 requests per setting. For comparison, Qwen3.8-27B uses 28.6 GB on the Mac; Jeff's memory on the Mac is not measured yet.
Not measured yet
Nothing on this site is estimated. These are still to come:
- Jeff's memory on the Mac (measured on the RTX only)
- end-to-end ticket times (not run yet)
Where the numbers came from
- Adapters and serving:
jeff-finetunes/adapters/BASELINE.md - Against Qwen3.8-27B:
jeff-reference-app results/cascade.json - Jeff v1.2 0.8B:
jev/runs/eval/0.8b-20260929-2258-final-calibrated.json - Jeff v1.2 2B:
jev/runs/eval/2b-20260930-2347-final-calibrated.json
Collected into one file on 2026-10-01T11:56:57+00:00.
