Serving several adapters
One Jeff server, one base model, any number of adapters, chosen per request by name.
A Jeff server loads the base model once. With JEFF_ADAPTERS set to a folder, every subfolder of it is one adapter,
and the subfolder's name is the model name clients use.
JEFF_CHECKPOINT=checkpoints/jeff-0.8b JEFF_ADAPTERS=adapters/ PORT=8765 uv run --no-default-groups jeff-serveadapters/
legal-clauses/ → model="legal-clauses"
support-intents/ → model="support-intents"
spam/ → model="spam"Names
Adapter names use lower-case letters, digits, dots and dashes. The names jeff, jeff-latest and jeff-<base>
always mean the plain base model, so no adapter may use them. GET /v1/models and GET /health list every adapter
the server holds, with the number of options each takes.
Adding, replacing and removing adapters
Change the folder, then ask the server to catch up. No restart is needed.
curl -X POST localhost:8765/v1/adapters/reload
# {"added": ["spam"], "removed": [], "reloaded": ["legal-clauses"], "adapters": ["legal-clauses", "spam", "support-intents"]}New folders are loaded, removed ones dropped and changed ones reloaded. An adapter made for a different base model
is refused with a 422 that says so: each adapter records a hash of the exact base weights it was trained on. Without
JEFF_ADAPTERS the endpoint answers 409.
What an adapter folder holds
| File | What it is |
|---|---|
adapter_config.json, adapter_model.safetensors |
The LoRA weights, in the standard PEFT format |
readout.safetensors |
The adapter's own copy of the answer layer |
decision_config.json |
Its temperature, prompt layout, option limit, and the base checkpoint's hash |
Memory and speed
Each adapter adds a few tens of megabytes of GPU memory, and one request to an adapter takes somewhat longer than the base model alone. Switching between adapters from one request to the next costs nothing measurable. The measured numbers (median and 95th-percentile time per decision, GPU memory with one and with several adapters) are on the Results page, with the hardware they were measured on.
Both PyTorch (NVIDIA GPUs and CPU) and MLX (Apple silicon) serve adapters.
