toolsAgent tool choice
Picks the next tool for an AI agent to call, or says to answer directly or to ask the user for missing information.
Data: User messages generated with Qwen3.8-Max (hosted, Alibaba Cloud DashScope), used to finish in time; agents and tool lists by Qwen3.8-Flash-Next (local); checked by Qwen3.8-Flash and DeepSeek-V4-Flash (hosted).
Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap
Use it when
- Your agent has a list of tools and must decide, at the start of each turn, which one to call next.
- You want a quick first decision with a probability, so a low-confidence turn can be passed to a larger model.
- Your tool list is your own; the adapter reads each tool's signature and description (2 to 150 tools in training).
Not a good fit when
- You need the tool's arguments filled in. The adapter picks one tool; it does not write the call.
- The task needs several tools in a row. The adapter chooses only the next step; ask again after each call.
- Your tool list is so long that the request would pass 8,192 tokens, the longest row in training.
Request format
The state is an object with these fields, in this order. Only user_message changes from request to request, so it comes last and the rest can be prepared in advance.
| State field | Changes per request | What goes in it |
|---|---|---|
agent | No | One sentence on what the agent is for. |
conversation | No | The earlier turns, each with a role and text. May be empty; at most about 6 turns in training. |
user_message | Yes | The user's latest message. |
| Question | Type | What it decides |
|---|---|---|
tool | Choice | Which tool the agent should call next, or whether to answer directly or ask the user for missing information first. Options: |
- Keep
answer_directlyandask_useras the first two options, with exactly the wording in the example. - List your tools after them in a fixed order so the unchanging part of the request can be prepared in advance.
- Choose
answer_directlyalso when the user asks for something no tool can do; the agent then explains that it cannot. - Use the instructions below word for word; the adapter was trained mostly on them.
General rules for every request are in the request format guide.
Example
The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).
from jeff import Client
from jeff.client import choice_question
jeff = Client("http://localhost:8765", model="tools")
state = {
"agent": "An assistant that manages the calendar and email of a small design studio.",
"conversation": [
{"role": "user", "text": "What's on my calendar tomorrow?"},
{
"role": "assistant",
"text": "Tomorrow you have a client call with Harbour Books at 10:00 and a team review at 15:00.",
},
],
"user_message": "Please move the team review to Friday at the same time.",
}
answers = jeff.ask(state, {
"tool": choice_question(
{
"answer_directly": "No tool is needed: answer the user directly from the conversation",
"ask_user": "A tool is needed but required information is missing: ask the user first",
"t1": "list_events(date): list the calendar events on a given day",
"t2": "create_event(title, start, end, attendees): add a new event to the calendar",
"t3": "move_event(event_id, new_start, new_end): move an existing event to another date or time",
"t4": "send_email(to, subject, body): send an email now",
"t5": "draft_email(to, subject, body): save an email as a draft without sending it",
},
"Which tool should the agent call next to handle the user's latest message? Use the conversation for context. If no tool is needed, choose answer directly. If a tool is needed but information it requires is missing, choose ask the user.",
),
})
print("tool", answers.choice("tool").key)import { Client, choiceQuestion } from '@jeff/client';
const jeff = new Client({ url: 'http://localhost:8765', model: 'tools' });
const state = {
agent: 'An assistant that manages the calendar and email of a small design studio.',
conversation: [
{ role: 'user', text: 'What\'s on my calendar tomorrow?' },
{
role: 'assistant',
text: 'Tomorrow you have a client call with Harbour Books at 10:00 and a team review at 15:00.',
},
],
user_message: 'Please move the team review to Friday at the same time.',
};
const answers = await jeff.ask(state, {
tool: choiceQuestion(
{
answer_directly: 'No tool is needed: answer the user directly from the conversation',
ask_user: 'A tool is needed but required information is missing: ask the user first',
t1: 'list_events(date): list the calendar events on a given day',
t2: 'create_event(title, start, end, attendees): add a new event to the calendar',
t3: 'move_event(event_id, new_start, new_end): move an existing event to another date or time',
t4: 'send_email(to, subject, body): send an email now',
t5: 'draft_email(to, subject, body): save an email as a draft without sending it',
},
'Which tool should the agent call next to handle the user\'s latest message? Use the conversation for context. If no tool is needed, choose answer directly. If a tool is needed but information it requires is missing, choose ask the user.',
),
});
console.log('tool', answers.tool.key);curl -s http://localhost:8765/v1/systemone \
-H 'content-type: application/json' \
-d '{
"model": "tools",
"state": {
"agent": "An assistant that manages the calendar and email of a small design studio.",
"conversation": [
{
"role": "user",
"text": "What'\''s on my calendar tomorrow?"
},
{
"role": "assistant",
"text": "Tomorrow you have a client call with Harbour Books at 10:00 and a team review at 15:00."
}
],
"user_message": "Please move the team review to Friday at the same time."
},
"questions": {
"tool": {
"type": "choice",
"instructions": "Which tool should the agent call next to handle the user'\''s latest message? Use the conversation for context. If no tool is needed, choose answer directly. If a tool is needed but information it requires is missing, choose ask the user.",
"criteria": {
"answer_directly": "No tool is needed: answer the user directly from the conversation",
"ask_user": "A tool is needed but required information is missing: ask the user first",
"t1": "list_events(date): list the calendar events on a given day",
"t2": "create_event(title, start, end, attendees): add a new event to the calendar",
"t3": "move_event(event_id, new_start, new_end): move an existing event to another date or time",
"t4": "send_email(to, subject, body): send an email now",
"t5": "draft_email(to, subject, body): save an email as a draft without sending it"
}
}
}
}'Response
{
"model": "tools",
"answers": {
"tool": {
"type": "choice",
"probabilities": {
"answer_directly": 0.04872739409107487,
"ask_user": 0.9460192974217675,
"t1": 0.0018965681163852716,
"t2": 0.0003707438379073421,
"t3": 0.0024392489346947615,
"t4": 0.0003419324152816466,
"t5": 0.00020481518288857154
},
"choice": "ask_user",
"confidence": 0.9370225136587286
}
},
"usage": {
"input_tokens": 358,
"output_tokens": 0,
"orders": 1
}
}Results
On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters
| Test set | Test rows | Qwen3.5-0.8B untrained | Jeff v1.2 0.8B alone | Jeff v1.2 0.8B + adapter |
|---|---|---|---|---|
tools | 5,157 | 18.0% · 0.064 | 57.8% · 0.091 | 97.9% · 0.004 |
tools5,157 test rows- Qwen3.5-0.8B untrained
- 18.0% · 0.064
- Jeff v1.2 0.8B alone
- 57.8% · 0.091
- Jeff v1.2 0.8B + adapter
- 97.9% · 0.004
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
How sure is it, and is it right?
Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.
When this adapter says it is about 99.4% sure, it is right about 99.4% of the time (4,889 test rows).
Point at a dot for its numbers. Bigger dots hold more rows.
| Model | Calibration error (ECE) | Brier score | Log loss |
|---|---|---|---|
| Qwen3.5-0.8B untrained | 0.064 | 0.929 | 3.197 |
| Jeff v1.2 0.8B alone | 0.091 | 0.602 | 1.657 |
| Jeff v1.2 0.8B + adapter | 0.004 | 0.032 | 0.073 |
All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.
Where it gets things wrong
- a listed option read as another listed option: 51 rows
ask_userread as a listed option: 21 rows- a listed option read as
ask_user: 12 rows answer_directlyread as a listed option: 7 rowsanswer_directlyread asask_user: 7 rows
The five commonest mistakes with the adapter, out of 5,157 test rows.
| Right answer | Test rows | Accuracy with the adapter |
|---|---|---|
| a listed option | 3,611 | 98.1% |
answer_directly | 1,031 | 98.6% |
ask_user | 515 | 95.5% |
“A listed option” pools the rows whose right answer is one of the options listed in that request (keys such as o3 or t1, whose meaning changes from row to row); “another listed option” is a different one of them.
Against Qwen3.8-27B
More accurate, 36× faster than Qwen3.8-27B alone.
Gain over Qwen3.8-27B alone: +7.7 points [+4.7, +11.0] (95% interval).
On 300 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 90.3% of the time at 11.27 s per query on average; with Jeff and this adapter answering first, it was right 98.0% at 309 ms.
This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.
| Route | Mean | Median | 95th percentile | Typical prompt |
|---|---|---|---|---|
| Qwen3.8-27B alone | 11.27 s | 9.61 s | 26.67 s | 1,296 tokens |
| Jeff + adapter, its own answer | 309 ms | 269 ms | 742 ms | 1,263 tokens |
Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.
This task's threshold 0.00: accuracy 98.0%, 36.5× faster, 0.0% sent on to Qwen3.8-27B.
Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 300 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.
How it was trained
- Training rows
- 43,123
- Steps
- 674
- Training time
- 131 min
- Size as saved
- 41.5 MB
One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-tools-20260930-1532.
Source: jeff-finetunes/adapters/BASELINE.md
Not yet measured: BFCL v3, multiple and irrelevance; When2Call, test.
Data card
Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean
- Test set, so anyone can check the numbers
- Calibration rows, the rows its threshold is chosen on
- QA report, sanitised: the data-quality checks run before training
- How the test set was held out
- Requests to the 10% of the roughly 450 generated agents that were never trained on.
- Training data
- Training data not published.
- Which models made the data, counted on the 39,203 training rows
- Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total.
What it did Model Where it ran Training rows Wrote the user message teacher.messageQwen3.8-Max hosted (Alibaba Cloud DashScope) 39,203 Wrote the tool list teacher.toolsQwen3.8-Flash-Next local (own hardware) 39,203 Wrote the agent teacher.agentQwen3.8-Flash-Next local (own hardware) 39,203 Reworded the instructions teacher.instructionsQwen3.8-Flash-Next local (own hardware) 10,554 Judged the label teacher.judgeQwen3.8-Flash hosted (Alibaba Cloud DashScope) 39,203 Checked the judged label judge.checker_modelQwen3.8-Flash hosted (Alibaba Cloud DashScope) 39,203 Gave a second opinion on the label second_opinion.modelDeepSeek-V4-Flash hosted (DeepSeek) 5,340 Judged the label teacher.judgeDeepSeek-V4-Flash hosted (DeepSeek) 5,338 Judged the label (strict check) teacher.judge_strictDeepSeek-V4-Flash hosted (DeepSeek) 3,920 Judged the label (strict check) teacher.judge_strictQwen3.8-Flash hosted (Alibaba Cloud DashScope) 3,920 Judged the label teacher.judgeQwen3.8-Flash-Next local (own hardware) 1,555 - The terms of the hosted model providers are being checked for training and publication use.
Data and licence
The adapter is released under Apache-2.0. It was trained on:
- Generated agents, tool sets and user messagesLicence: Released with the adapter under Apache-2.0 · Generated by Qwen3.8-Max (hosted) and Qwen3.8-Flash-Next (local); see the data card
About 450 agents with their tool lists (including deliberate near-duplicate tools) and 39,203 training rows, written by language models and checked by code and a second pass (which models, and for how many rows, is in the data card).
Changelog
- 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.
Comments
Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.
