Skip to content
JeffHub
RoutingOfficialReproducedVersion 0.1.0

toolsAgent tool choice

Picks the next tool for an AI agent to call, or says to answer directly or to ask the user for missing information.

Data: User messages generated with Qwen3.8-Max (hosted, Alibaba Cloud DashScope), used to finish in time; agents and tool lists by Qwen3.8-Flash-Next (local); checked by Qwen3.8-Flash and DeepSeek-V4-Flash (hosted).

Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap

Use it when

  • Your agent has a list of tools and must decide, at the start of each turn, which one to call next.
  • You want a quick first decision with a probability, so a low-confidence turn can be passed to a larger model.
  • Your tool list is your own; the adapter reads each tool's signature and description (2 to 150 tools in training).

Not a good fit when

  • You need the tool's arguments filled in. The adapter picks one tool; it does not write the call.
  • The task needs several tools in a row. The adapter chooses only the next step; ask again after each call.
  • Your tool list is so long that the request would pass 8,192 tokens, the longest row in training.

Request format

The state is an object with these fields, in this order. Only user_message changes from request to request, so it comes last and the rest can be prepared in advance.

State fieldChanges per requestWhat goes in it
agentNoOne sentence on what the agent is for.
conversationNoThe earlier turns, each with a role and text. May be empty; at most about 6 turns in training.
user_messageYesThe user's latest message.
QuestionTypeWhat it decides
toolChoiceWhich tool the agent should call next, or whether to answer directly or ask the user for missing information first.

Options: answer_directly and ask_user first, word for word, then your tools as t1, t2, … with each tool's signature and a one-line description.

  • Keep answer_directly and ask_user as the first two options, with exactly the wording in the example.
  • List your tools after them in a fixed order so the unchanging part of the request can be prepared in advance.
  • Choose answer_directly also when the user asks for something no tool can do; the agent then explains that it cannot.
  • Use the instructions below word for word; the adapter was trained mostly on them.

General rules for every request are in the request format guide.

Example

The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).

from jeff import Client
from jeff.client import choice_question

jeff = Client("http://localhost:8765", model="tools")

state = {
    "agent": "An assistant that manages the calendar and email of a small design studio.",
    "conversation": [
        {"role": "user", "text": "What's on my calendar tomorrow?"},
        {
            "role": "assistant",
            "text": "Tomorrow you have a client call with Harbour Books at 10:00 and a team review at 15:00.",
        },
    ],
    "user_message": "Please move the team review to Friday at the same time.",
}

answers = jeff.ask(state, {
    "tool": choice_question(
        {
            "answer_directly": "No tool is needed: answer the user directly from the conversation",
            "ask_user": "A tool is needed but required information is missing: ask the user first",
            "t1": "list_events(date): list the calendar events on a given day",
            "t2": "create_event(title, start, end, attendees): add a new event to the calendar",
            "t3": "move_event(event_id, new_start, new_end): move an existing event to another date or time",
            "t4": "send_email(to, subject, body): send an email now",
            "t5": "draft_email(to, subject, body): save an email as a draft without sending it",
        },
        "Which tool should the agent call next to handle the user's latest message? Use the conversation for context. If no tool is needed, choose answer directly. If a tool is needed but information it requires is missing, choose ask the user.",
    ),
})
print("tool", answers.choice("tool").key)

Response

{
  "model": "tools",
  "answers": {
    "tool": {
      "type": "choice",
      "probabilities": {
        "answer_directly": 0.04872739409107487,
        "ask_user": 0.9460192974217675,
        "t1": 0.0018965681163852716,
        "t2": 0.0003707438379073421,
        "t3": 0.0024392489346947615,
        "t4": 0.0003419324152816466,
        "t5": 0.00020481518288857154
      },
      "choice": "ask_user",
      "confidence": 0.9370225136587286
    }
  },
  "usage": {
    "input_tokens": 358,
    "output_tokens": 0,
    "orders": 1
  }
}

Results

On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters

  • tools5,157 test rows
    Qwen3.5-0.8B untrained
    18.0% · 0.064
    Jeff v1.2 0.8B alone
    57.8% · 0.091
    97.9% · 0.004

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

How sure is it, and is it right?

Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.

When this adapter says it is about 99.4% sure, it is right about 99.4% of the time (4,889 test rows).

Jeff v1.2 0.8B aloneJeff v1.2 0.8B + adapterperfectly calibrated
0%0%25%25%50%50%75%75%100%100%Stated confidenceRight answersJeff v1.2 0.8B alone: 17 rows stated 6% on average and were right 11.8% of the timeJeff v1.2 0.8B alone: 224 rows stated 11% on average and were right 17.9% of the timeJeff v1.2 0.8B alone: 381 rows stated 17% on average and were right 28.1% of the timeJeff v1.2 0.8B alone: 455 rows stated 23% on average and were right 35.8% of the timeJeff v1.2 0.8B alone: 481 rows stated 30% on average and were right 40.7% of the timeJeff v1.2 0.8B alone: 419 rows stated 37% on average and were right 52.3% of the timeJeff v1.2 0.8B alone: 389 rows stated 43% on average and were right 58.1% of the timeJeff v1.2 0.8B alone: 380 rows stated 50% on average and were right 57.9% of the timeJeff v1.2 0.8B alone: 365 rows stated 57% on average and were right 61.4% of the timeJeff v1.2 0.8B alone: 322 rows stated 63% on average and were right 71.4% of the timeJeff v1.2 0.8B alone: 342 rows stated 70% on average and were right 72.8% of the timeJeff v1.2 0.8B alone: 339 rows stated 77% on average and were right 75.8% of the timeJeff v1.2 0.8B alone: 346 rows stated 83% on average and were right 80.3% of the timeJeff v1.2 0.8B alone: 375 rows stated 90% on average and were right 80.3% of the timeJeff v1.2 0.8B alone: 322 rows stated 96% on average and were right 82.9% of the timeJeff v1.2 0.8B + adapter: 5 rows stated 37% on average and were right 60.0% of the timeJeff v1.2 0.8B + adapter: 7 rows stated 43% on average and were right 14.3% of the timeJeff v1.2 0.8B + adapter: 16 rows stated 50% on average and were right 31.3% of the timeJeff v1.2 0.8B + adapter: 16 rows stated 58% on average and were right 62.5% of the timeJeff v1.2 0.8B + adapter: 29 rows stated 62% on average and were right 62.1% of the timeJeff v1.2 0.8B + adapter: 29 rows stated 70% on average and were right 72.4% of the timeJeff v1.2 0.8B + adapter: 29 rows stated 77% on average and were right 69.0% of the timeJeff v1.2 0.8B + adapter: 41 rows stated 84% on average and were right 85.4% of the timeJeff v1.2 0.8B + adapter: 96 rows stated 90% on average and were right 83.3% of the timeJeff v1.2 0.8B + adapter: 4889 rows stated 99% on average and were right 99.4% of the time

Point at a dot for its numbers. Bigger dots hold more rows.

Calibration scores
ModelCalibration error (ECE)Brier scoreLog loss
Qwen3.5-0.8B untrained0.0640.9293.197
Jeff v1.2 0.8B alone0.0910.6021.657
Jeff v1.2 0.8B + adapter0.0040.0320.073

All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.

Where it gets things wrong

  • a listed option read as another listed option: 51 rows
  • ask_user read as a listed option: 21 rows
  • a listed option read as ask_user: 12 rows
  • answer_directly read as a listed option: 7 rows
  • answer_directly read as ask_user: 7 rows

The five commonest mistakes with the adapter, out of 5,157 test rows.

Accuracy per right answer
Right answerTest rowsAccuracy with the adapter
a listed option3,61198.1%
answer_directly1,03198.6%
ask_user51595.5%

“A listed option” pools the rows whose right answer is one of the options listed in that request (keys such as o3 or t1, whose meaning changes from row to row); “another listed option” is a different one of them.

Against Qwen3.8-27B

More accurate, 36× faster than Qwen3.8-27B alone.

Gain over Qwen3.8-27B alone: +7.7 points [+4.7, +11.0] (95% interval).

On 300 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 90.3% of the time at 11.27 s per query on average; with Jeff and this adapter answering first, it was right 98.0% at 309 ms.

This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.

Time per query and prompt length for both routes
RouteMeanMedian95th percentileTypical prompt
Qwen3.8-27B alone11.27 s9.61 s26.67 s1,296 tokens
Jeff + adapter, its own answer309 ms269 ms742 ms1,263 tokens

Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.

This task's threshold 0.00: accuracy 98.0%, 36.5× faster, 0.0% sent on to Qwen3.8-27B.

Accuracy98.0%
90.0%95.0%100.0%0.000.250.500.751.00
Speed-up36.5×
0.0×20.0×40.0×0.000.250.500.751.00
Sent on to Qwen3.8-27B0.0%
0.0%50.0%100.0%0.000.250.500.751.00

Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 300 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.

All tasks, and how this was measured

How it was trained

Training rows
43,123
Steps
674
Training time
131 min
Size as saved
41.5 MB

One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-tools-20260930-1532.

Source: jeff-finetunes/adapters/BASELINE.md

Not yet measured: BFCL v3, multiple and irrelevance; When2Call, test.

Data card

Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean

How the test set was held out
Requests to the 10% of the roughly 450 generated agents that were never trained on.
Training data
Training data not published.
Which models made the data, counted on the 39,203 training rows
What it didModelWhere it ranTraining rows
Wrote the user message teacher.messageQwen3.8-Maxhosted (Alibaba Cloud DashScope)39,203
Wrote the tool list teacher.toolsQwen3.8-Flash-Nextlocal (own hardware)39,203
Wrote the agent teacher.agentQwen3.8-Flash-Nextlocal (own hardware)39,203
Reworded the instructions teacher.instructionsQwen3.8-Flash-Nextlocal (own hardware)10,554
Judged the label teacher.judgeQwen3.8-Flashhosted (Alibaba Cloud DashScope)39,203
Checked the judged label judge.checker_modelQwen3.8-Flashhosted (Alibaba Cloud DashScope)39,203
Gave a second opinion on the label second_opinion.modelDeepSeek-V4-Flashhosted (DeepSeek)5,340
Judged the label teacher.judgeDeepSeek-V4-Flashhosted (DeepSeek)5,338
Judged the label (strict check) teacher.judge_strictDeepSeek-V4-Flashhosted (DeepSeek)3,920
Judged the label (strict check) teacher.judge_strictQwen3.8-Flashhosted (Alibaba Cloud DashScope)3,920
Judged the label teacher.judgeQwen3.8-Flash-Nextlocal (own hardware)1,555
Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total.
The terms of the hosted model providers are being checked for training and publication use.

Data and licence

The adapter is released under Apache-2.0. It was trained on:

  • Generated agents, tool sets and user messagesLicence: Released with the adapter under Apache-2.0 · Generated by Qwen3.8-Max (hosted) and Qwen3.8-Flash-Next (local); see the data card

    About 450 agents with their tool lists (including deliberate near-duplicate tools) and 39,203 training rows, written by language models and checked by code and a second pass (which models, and for how many rows, is in the data card).

Changelog

  1. 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.

Comments

Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.