When to use Jeff
For decisions you make often and fast, a small model with an adapter beats a big one. For the hard few, hand off.
Jeff is built for System 1 decisions: quick judgement calls between options you can name in advance. For most of those, one small model with an adapter per job is the right tool. For some it is not. This page says which is which.
Where Jeff with an adapter wins
Use Jeff when the decision is:
- Frequent. Thousands or millions of decisions, where cost per decision and the number of GPUs matter.
- Latency-bound. A user or a system is waiting: voice commands, choosing an agent's next tool, a guard in front of every request, routing a ticket as it arrives.
- Short-input. A message, a ticket, a screen, a clause. Jeff reads up to about 8,000 tokens.
- Learnable. You have examples, or can generate them, and the right answer follows from the input rather than from long reasoning.
- Private. The data must stay on your own hardware or on the device.
- Reasonably stable. The options and the kinds of input do not change every week.
That covers routing, triage, intent, moderation, spam, clause typing, grounding checks and interface commands. For these, a large model is slower and more expensive, and once an adapter is trained it is often no more accurate.
Where a larger model is the better choice
- The decision needs reasoning. Knowing the decision in advance does not make it easy. Cross-references in a contract, arithmetic in a grounding check or weighing several facts at once need a model that reasons. On the reasoning-heavy benchmarks Jeff stays well below large models (Jeff v1.2 0.8B on its public benchmark panel: BBH 63.2 against Jev's published 94.3; JudgeBench 60.9 against 78.6). It does well where the task is recognition rather than reasoning: Financial PhraseBank 96.3, RAGTruth 85.5, WinoGrande 68.7, and 78.7 over the whole panel (the 2B scores 81.7). An adapter narrows the gap for one task; it does not add reasoning a 0.8B model lacks.
- Long inputs. Whole contracts, long email threads or large documents are beyond what Jeff was trained on.
- Inputs that drift, or attack. An adapter is good on inputs like its training data and improves only when it is retrained. New scam patterns, new products or deliberate attacks on a prompt-injection guard wear it down, and its confidence becomes less reliable at the same time.
- Few decisions, costly mistakes. At a few hundred decisions a day, speed and cost barely matter and a few points of accuracy do.
- No training data. Rare categories or a brand-new decision with a handful of examples: a large model with those examples in its prompt does better than an adapter trained on them.
- You need a stated reason. Jeff returns probabilities, not explanations. Some regulated decisions need a rationale on record.
- Many decisions that change often. Every adapter needs data, training, testing and retraining when the base model changes. Changing a large model's prompt takes minutes.
Often best: both, in a cascade
Jeff answers every request first. When it is confident, its answer stands. When it is not, the request goes to the large model.
from jeff import Client
jeff = Client("http://localhost:8765", model="support-intents")
picked = jeff.choose(state, options, instructions)
if picked.probability >= threshold:
intent = picked.key # most requests: milliseconds, on your hardware
else:
intent = ask_large_model(state, options, instructions) # the hard fewThis works because Jeff's probabilities are calibrated: fitted so that, on inputs like its test data, answers given at 90% are right about 90% of the time. Check this on your own data before relying on a threshold. You keep most of the speed and cost saving, and the large model covers the cases above.
Choosing the threshold
Run Jeff and the large model on the adapter's held-out test set. For each confidence threshold, note how many requests Jeff answers and how accurate those answers are. If Jeff alone matches the large model, you do not need the large model for this decision. If not, the threshold where Jeff's accuracy meets your bar tells you how much traffic the large model will see.
Next: Request format
