Skip to content
JeffHub

Serving several adapters

One Jeff server, one base model, any number of adapters, chosen per request by name.

A Jeff server loads the base model once. With JEFF_ADAPTERS set to a folder, every subfolder of it is one adapter, and the subfolder's name is the model name clients use.

JEFF_CHECKPOINT=checkpoints/jeff-0.8b JEFF_ADAPTERS=adapters/ PORT=8765 uv run --no-default-groups jeff-serve
adapters/
  legal-clauses/     →  model="legal-clauses"
  support-intents/   →  model="support-intents"
  spam/              →  model="spam"

Names

Adapter names use lower-case letters, digits, dots and dashes. The names jeff, jeff-latest and jeff-<base> always mean the plain base model, so no adapter may use them. GET /v1/models and GET /health list every adapter the server holds, with the number of options each takes.

Adding, replacing and removing adapters

Change the folder, then ask the server to catch up. No restart is needed.

curl -X POST localhost:8765/v1/adapters/reload
# {"added": ["spam"], "removed": [], "reloaded": ["legal-clauses"], "adapters": ["legal-clauses", "spam", "support-intents"]}

New folders are loaded, removed ones dropped and changed ones reloaded. An adapter made for a different base model is refused with a 422 that says so: each adapter records a hash of the exact base weights it was trained on. Without JEFF_ADAPTERS the endpoint answers 409.

What an adapter folder holds

File What it is
adapter_config.json, adapter_model.safetensors The LoRA weights, in the standard PEFT format
readout.safetensors The adapter's own copy of the answer layer
decision_config.json Its temperature, prompt layout, option limit, and the base checkpoint's hash

Memory and speed

Each adapter adds a few tens of megabytes of GPU memory, and one request to an adapter takes somewhat longer than the base model alone. Switching between adapters from one request to the next costs nothing measurable. The measured numbers (median and 95th-percentile time per decision, GPU memory with one and with several adapters) are on the Results page, with the hardware they were measured on.

Both PyTorch (NVIDIA GPUs and CPU) and MLX (Apple silicon) serve adapters.

Next: Preparing requests in advance