Skip to content
JeffHub

Why Jeff

An agent makes many small decisions for every one that needs real thought. Jeff makes the small ones on your own hardware, in milliseconds, and says how sure it is. A large model does the thinking.

System 1 and System 2

Psychologists describe two kinds of thinking: System 1 is fast and automatic, System 2 slow and careful. An agent needs both. Before it writes a single reply, it has to decide whether a message is safe, where it should go, what the person wants and which tool to call. Each of those is a choice between options you can name in advance.

Sending each of them to a large model such as Qwen3.8-27B costs time, because the model generates text to answer; it costs memory, because the large model must stay loaded for every step; and it costs money or hardware when many requests arrive at once. Jeff answers each of those decisions in one forward pass of a 0.8B model, and the large model is left to reason and write.

On 8 of 9 tasks, Jeff with its adapter beats the 27B on its own; the 27B stays as a fallback you can tune per task, and for anything no adapter covers. Measured on an Apple M4 Max, 128 GB over 8 adapters: with Jeff and its adapters making the decisions, a query took 250 ms on average instead of 8.11 s (37.7× faster), and accuracy rose from 86.6% to 95.3%, a gain of +8.7 points [+7.4, +10.1] with its 95% interval. Where the gain comes from.

Speed and memory, measured

Qwen3.8-27B alone against Qwen3.8-27B with Jeff and adapters
Measure27B alone27B + Jeff
Time per query (average over the 8 adapters)8.11 s250 ms
Accuracy (mean of the 8 adapters)86.6%95.3%
Accuracy gain, with its 95% interval—+8.7 points [+7.4, +10.1]
Queries Qwen3.8-27B has to answer100%0.2%
Memory28.6 GB30.6 GB

The emotion adapter is left out of the averages (why). Time, accuracy and share answered on the same test rows on an Apple M4 Max, 128 GB; each task has its own threshold. Memory: Qwen3.8-27B measured on the Mac; Jeff's 1.96 GB with all nine adapters loaded was measured on an NVIDIA RTX PRO 6000 and is not yet measured on the Mac. Every decision, in detail.

about 25 ms
per decision, Jeff alone (median)
about 30 ms
per decision with an adapter (median)
37.7× faster
than Qwen3.8-27B alone, same decisions
1.96 GB
Jeff with all nine adapters loaded

Speed and memory measured on an NVIDIA RTX PRO 6000, one decision per request, median over the same prompts in every setting. Switching to a different adapter on every request costs nothing measurable (30.0 ms median, against 31.2 ms with one adapter). For comparison, Qwen3.8-27B uses 28.6 GB with 8-bit weights on the Mac: Jeff with all nine adapters uses 6.9% of that. Full speed and memory table.

What calibrated means

Jeff does not just pick an option; it gives each option a probability. Calibrated means those probabilities can be taken at face value: of all the answers Jeff gives at 70% confidence, about 70% are right. Of those it gives at 99%, about 99% are right.

That makes a simple rule possible. When Jeff is sure, act on its answer. When it is not, hand the decision to a large model or a person. Most decisions go the fast way, and the hard few get more care. The reference app gives each decision its own threshold, and hands a decision over only when that is needed to stay ahead of the large model.

Calibration error (ECE) measures how far the stated confidence is from the real hit rate, on average. Lower is better; 0 is perfect. On each adapter's own test set:

Calibration error per adapter
Adapter test setJeff v1.2 0.8B aloneJeff v1.2 0.8B + adapter
guard0.2800.004
triage0.0260.009
support-intents0.0800.006
tools0.0910.004
ground0.2370.012
nav0.1950.005
emotion0.0380.020
spam0.1090.008
legal-clauses0.0320.011

ECE with 15 confidence bins, after each model's fitted temperature. All results.

What Jeff cannot do

  • It does not write text. It chooses between options and gives probabilities. Replies, summaries and plans need a language model.
  • It needs the options listed. You name the choices (teams, tools, labels); Jeff picks among them. It cannot invent an option you did not give.
  • Hard, fine-grained tasks stay hard. Picking one of 28 emotions for a short comment, the emotion adapter is right 60.6% of the time on its 5,408-row test set: far better than without it, but only moderate. Its calibration tells you which answers to trust.
  • It does not reason at length. Decisions that need several steps of thought, long documents, or a stated reason are better left to a large model.

When to use Jeff, and when a larger model is the better choice.