Training an adapter
Build clean data in Jeff's row format, train a LoRA adapter on one GPU, and measure it honestly. With the lessons from the nine official adapters.
An adapter is worth training when Jeff alone is not accurate enough for your job. Across the nine official adapters, mean accuracy went from 55.6% with Jeff alone to 91.5% with the adapter (all results).
Building those adapters taught us that the hard part is the data, not the training. Every one of our generated data sets failed our own quality review at least once. Read the lessons below before you write a single row, and follow the data guidelines: a data set that meets them carries over to v1.3 unchanged, while v1.2 weights have to be retrained (roadmap).
Lessons from the official adapters
Shortcuts are the main danger. A shortcut is a surface feature that gives the answer away without understanding the input. A model trained on shortcuts scores well on test data made the same way, then fails on real users. These are shortcuts we found in our own data and fixed:
- The correct option was more often the longest one.
- Questions were much longer than commands: length alone predicted the answer 49% of the time, against 33% by chance.
- One generation round wrote all its "none of these" requests in a recognisable style.
- Generated short questions dropped "the" and "a", so missing articles meant "question".
- A phishing corpus leaked its recipient address, so the address alone meant "phishing".
- Unfilled template slots such as
{{Order Number}}were left in support messages.
Check what a model that ignores meaning can do. Train a tiny model on surface features only: length, word count, punctuation, the position of the right option. If it clearly beats chance, fix the data before training. The shortcut report does this for you.
Fixing one shortcut can create another. When we added short questions to break the length shortcut, filler words such as "um" and "uh" became the new giveaway, because only the generated commands had them. Re-run every check after every fix.
Hold out whole groups for the test, not random rows: whole apps, companies or sources. Then the test measures new situations, not memory. For example, the triage test set is the messages to 10% of its generated organisations, none of which appear in training. split does this.
Check for leaks. No test row, and no near-copy of one (the same message with one word changed), may appear in training. leak-check fails if one does.
Mix in about 10% replay data, so the adapter does not forget general skills. The official adapters used a sample of the base model's own training data. That data is not available to you today: use your own replay rows or none, until the official replay set ships with v1.3. replay-mix mixes them in.
Train one pass over the data (one epoch), as every official adapter was trained.
Have someone who did not build the data review it before training. Our independent reviews caught problems the builders had missed.
Record which model wrote each generated row, and the licence of every source. You will need both for the data card.
Measure three ways on the same test set: the untrained base model, Jeff without the adapter, and Jeff with it. evaluate-three-ways runs all three.
1. Write the data
Training rows are JSON Lines, one question per row. Where your application asks several questions about the same
state, write one row per question with the same state and family. The family is the group you hold out together.
{"id": "r1", "suite": "mytask", "family": "company-17",
"state": {"company": "…", "channel": "email", "message": "…"},
"question": {"type": "choice", "instructions": "Which team should handle this message?",
"criteria": {"other": "Not for any of these teams", "k1": "…", "k2": "…"}},
"label": "k1", "target": "k1"}- The changing field last in
state, as in the request format. - Option keys never bare numbers; shuffle options per row unless their order means something (score scales).
- Split by family, not by row, as above. Keep public benchmarks you report as external tests out of training entirely.
- Instructions: mostly one canonical wording per question, with some rewordings, so the adapter is not brittle.
- Only data you may use: public data sets must allow training and redistribution.
check-rows checks the format.
2. Train
Adapter training needs an NVIDIA GPU (training on a Mac is not supported) and the lora extra.
uv sync --extra cuda --extra lora
uv run jeff-train \
--initial-checkpoint checkpoints/jeff-0.8b \
--lora-rank 16 --lr 2e-4 --readout-lr 5e-6 \
--train data/mytask/train.jsonl --development data/mytask/dev.jsonl --temperature data/mytask/calibration.jsonl \
--run mytask-r16 --output checkpoints/mytaskThese are the settings the official adapters start from. An adapter uses its base checkpoint's prompt layout unless
you set --prompt-layout; the official v1.2 adapters keep the base's layout. Each official adapter's full command is
in its examples/<name>/ folder in the Jeff repository. A larger --lora-rank is the first thing to try if an
adapter falls short.
3. Measure it, three ways
Report the same test three ways, as every JeffHub page does. jeff-evaluate and jeff-latency take an adapter
folder as the checkpoint.
uv run jeff-evaluate --data data/mytask/test.jsonl --local --checkpoint checkpoints/mytask/final --output runs/eval/mytask.json
uv run jeff-latency --checkpoint checkpoints/mytask/finalWrite down the checkpoint each number came from. Say plainly what the adapter is not good at. Then submit it.
Next: Adapter kit
