support-intentsCustomer request intents
Names what a customer or assistant user is asking for, from a fixed list of requests such as cancel_order or track_refund.
Data: Public data: Bitext, HWU64 and SNIPS
Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap
Use it when
- You run an online shop chat or a voice assistant and want each message matched to one request from a fixed list.
- Your requests are close to the ones in training (27 online-shop requests, 64 home-assistant requests, 7 voice-assistant requests), each described in one line.
- Messages are short and in English.
Not a good fit when
- Your list of requests is very different from the training lists. Try it, but measure first; the triage adapter reads your own team descriptions.
- A message holds several requests and you need all of them. The question picks one.
- You need the reply written. Jeff chooses between options; it does not generate text.
- Your messages are long emails or tickets. Training messages are short single requests.
Request format
The state is an object with these fields, in this order. Only message changes from request to request, so it comes last and the rest can be prepared in advance.
| State field | Changes per request | What goes in it |
|---|---|---|
service | No | One short phrase about the service the message is sent to. Training used three, one per data set, for example "Customer support chat of an online shop". |
message | Yes | The customer's or user's message as received. |
| Question | Type | What it decides |
|---|---|---|
intent | Choice | What the customer wants, as the request that best matches their message. Options: In training, each message's options were all the requests of its own data set: 27 for the online shop (Bitext), 64 for the home assistant (HWU64) and 7 for the voice assistant (SNIPS). Keys are snake_case names such as track_order, each with a one-line description. The full lists are BITEXT, HWU64 and SNIPS in descriptions.py in the source. |
- Use the option keys and descriptions from descriptions.py where they fit your service; the adapter was trained on them.
- Use the instructions below word for word; the adapter was trained mostly on them.
General rules for every request are in the request format guide.
Example
The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).
from jeff import Client
from jeff.client import choice_question
jeff = Client("http://localhost:8765", model="support-intents")
state = {
"service": "Customer support chat of an online shop",
"message": "hi, I sent back the jacket two weeks ago and still haven't seen the money. where is it?",
}
answers = jeff.ask(state, {
"intent": choice_question(
{
"track_refund": "Check the status of a refund they are expecting.",
"get_refund": "Get their money back for a purchase.",
"check_refund_policy": "Learn the refund policy and whether they qualify for a refund.",
"track_order": "Find out where their order is or its current status.",
"cancel_order": "Cancel an order they placed.",
"change_order": "Change an existing order (for example add, remove or swap items).",
"payment_issue": "Report or solve a problem with a payment.",
"complaint": "Make a complaint about the product, service or company.",
"contact_human_agent": "Talk to a human agent instead of an automated assistant.",
"delivery_period": "Find out when an order will arrive or how long delivery takes.",
},
"What does the customer want? Choose the request that best matches what the customer is asking for in their message.",
),
})
print("intent", answers.choice("intent").key)import { Client, choiceQuestion } from '@jeff/client';
const jeff = new Client({ url: 'http://localhost:8765', model: 'support-intents' });
const state = {
service: 'Customer support chat of an online shop',
message: 'hi, I sent back the jacket two weeks ago and still haven\'t seen the money. where is it?',
};
const answers = await jeff.ask(state, {
intent: choiceQuestion(
{
track_refund: 'Check the status of a refund they are expecting.',
get_refund: 'Get their money back for a purchase.',
check_refund_policy: 'Learn the refund policy and whether they qualify for a refund.',
track_order: 'Find out where their order is or its current status.',
cancel_order: 'Cancel an order they placed.',
change_order: 'Change an existing order (for example add, remove or swap items).',
payment_issue: 'Report or solve a problem with a payment.',
complaint: 'Make a complaint about the product, service or company.',
contact_human_agent: 'Talk to a human agent instead of an automated assistant.',
delivery_period: 'Find out when an order will arrive or how long delivery takes.',
},
'What does the customer want? Choose the request that best matches what the customer is asking for in their message.',
),
});
console.log('intent', answers.intent.key);curl -s http://localhost:8765/v1/systemone \
-H 'content-type: application/json' \
-d '{
"model": "support-intents",
"state": {
"service": "Customer support chat of an online shop",
"message": "hi, I sent back the jacket two weeks ago and still haven'\''t seen the money. where is it?"
},
"questions": {
"intent": {
"type": "choice",
"instructions": "What does the customer want? Choose the request that best matches what the customer is asking for in their message.",
"criteria": {
"track_refund": "Check the status of a refund they are expecting.",
"get_refund": "Get their money back for a purchase.",
"check_refund_policy": "Learn the refund policy and whether they qualify for a refund.",
"track_order": "Find out where their order is or its current status.",
"cancel_order": "Cancel an order they placed.",
"change_order": "Change an existing order (for example add, remove or swap items).",
"payment_issue": "Report or solve a problem with a payment.",
"complaint": "Make a complaint about the product, service or company.",
"contact_human_agent": "Talk to a human agent instead of an automated assistant.",
"delivery_period": "Find out when an order will arrive or how long delivery takes."
}
}
}
}'Response
{
"model": "support-intents",
"answers": {
"intent": {
"type": "choice",
"probabilities": {
"track_refund": 0.9028009763827044,
"get_refund": 0.013336566264330635,
"check_refund_policy": 0.00008649192014740069,
"track_order": 0.08268026999819446,
"cancel_order": 0.00022875498667774524,
"change_order": 0.00014981366970950356,
"payment_issue": 0.00027732248392003673,
"complaint": 0.00010125868711429669,
"contact_human_agent": 0.00022202538072712725,
"delivery_period": 0.0001165202264744504
},
"choice": "track_refund",
"confidence": 0.8920010848696716
}
},
"usage": {
"input_tokens": 297,
"output_tokens": 0,
"orders": 1
}
}Results
On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters
| Test set | Test rows | Qwen3.5-0.8B untrained | Jeff v1.2 0.8B alone | Jeff v1.2 0.8B + adapter |
|---|---|---|---|---|
support-intents | 5,577 | 33.9% · 0.162 | 85.1% · 0.080 | 96.8% · 0.006 |
support-intents5,577 test rows- Qwen3.5-0.8B untrained
- 33.9% · 0.162
- Jeff v1.2 0.8B alone
- 85.1% · 0.080
- Jeff v1.2 0.8B + adapter
- 96.8% · 0.006
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
How sure is it, and is it right?
Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.
When this adapter says it is about 99.4% sure, it is right about 99.4% of the time (4,985 test rows).
Point at a dot for its numbers. Bigger dots hold more rows.
| Model | Calibration error (ECE) | Brier score | Log loss |
|---|---|---|---|
| Qwen3.5-0.8B untrained | 0.162 | 0.864 | 3.097 |
| Jeff v1.2 0.8B alone | 0.080 | 0.229 | 0.538 |
| Jeff v1.2 0.8B + adapter | 0.006 | 0.052 | 0.124 |
All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.
Where it gets things wrong
general_quirkyread asnews_query: 7 rowsqa_factoidread asgeneral_quirky: 5 rowsgeneral_quirkyread asqa_factoid: 5 rowscalendar_queryread asgeneral_quirky: 5 rowssearch_screening_eventread assearch_creative_work: 5 rows
The five commonest mistakes with the adapter, out of 5,577 test rows.
| Right answer | Test rows | Accuracy with the adapter |
|---|---|---|
play_music | 220 | 96.4% |
calendar_set | 150 | 95.3% |
payment_issue | 130 | 100.0% |
check_refund_policy | 124 | 100.0% |
general_quirky | 115 | 70.4% |
complaint | 106 | 100.0% |
general_negate | 104 | 99.0% |
check_payment_methods | 104 | 100.0% |
delivery_period | 103 | 100.0% |
registration_problems | 101 | 100.0% |
Show all 97 answers
| Right answer | Test rows | Accuracy with the adapter |
|---|---|---|
play_music | 220 | 96.4% |
calendar_set | 150 | 95.3% |
payment_issue | 130 | 100.0% |
check_refund_policy | 124 | 100.0% |
general_quirky | 115 | 70.4% |
complaint | 106 | 100.0% |
general_negate | 104 | 99.0% |
check_payment_methods | 104 | 100.0% |
delivery_period | 103 | 100.0% |
registration_problems | 101 | 100.0% |
search_creative_work | 100 | 100.0% |
add_to_playlist | 100 | 100.0% |
book_restaurant | 100 | 100.0% |
qa_factoid | 99 | 85.9% |
get_weather | 99 | 100.0% |
search_screening_event | 98 | 94.9% |
rate_book | 98 | 99.0% |
change_shipping_address | 97 | 100.0% |
newsletter_subscription | 96 | 100.0% |
review | 95 | 100.0% |
get_invoice | 94 | 100.0% |
contact_human_agent | 93 | 100.0% |
check_cancellation_fee | 90 | 100.0% |
delete_account | 89 | 100.0% |
check_invoice | 88 | 100.0% |
recover_password | 88 | 100.0% |
weather_query | 88 | 100.0% |
contact_customer_service | 85 | 100.0% |
set_up_shipping_address | 84 | 100.0% |
calendar_query | 84 | 85.7% |
get_refund | 84 | 100.0% |
place_order | 82 | 100.0% |
create_account | 82 | 98.8% |
switch_account | 81 | 100.0% |
edit_account | 78 | 100.0% |
general_praise | 76 | 100.0% |
change_order | 75 | 98.7% |
email_sendemail | 75 | 93.3% |
news_query | 71 | 95.8% |
email_query | 69 | 98.6% |
general_affirm | 69 | 100.0% |
track_order | 64 | 100.0% |
general_explain | 64 | 100.0% |
general_repeat | 63 | 100.0% |
datetime_query | 63 | 95.2% |
delivery_options | 62 | 100.0% |
track_refund | 60 | 100.0% |
general_dontcare | 58 | 100.0% |
calendar_remove | 56 | 98.2% |
social_post | 53 | 94.3% |
general_confirm | 53 | 100.0% |
play_radio | 49 | 85.7% |
cooking_recipe | 47 | 95.7% |
qa_definition | 47 | 93.6% |
qa_currency | 44 | 100.0% |
play_podcasts | 41 | 95.1% |
general_commandstop | 38 | 100.0% |
lists_remove | 35 | 91.4% |
transport_query | 35 | 88.6% |
recommendation_events | 31 | 83.9% |
recommendation_locations | 30 | 86.7% |
lists_query | 30 | 93.3% |
qa_stock | 30 | 96.7% |
alarm_set | 30 | 96.7% |
music_query | 30 | 80.0% |
lists_createoradd | 27 | 92.6% |
iot_hue_lightoff | 26 | 100.0% |
cancel_order | 26 | 100.0% |
transport_ticket | 25 | 96.0% |
play_audiobook | 24 | 87.5% |
iot_cleaning | 23 | 100.0% |
email_querycontact | 23 | 87.0% |
play_game | 22 | 95.5% |
transport_taxi | 21 | 95.2% |
alarm_query | 21 | 90.5% |
takeaway_order | 20 | 70.0% |
iot_coffee | 20 | 100.0% |
music_likeness | 19 | 94.7% |
iot_hue_lightchange | 18 | 88.9% |
iot_hue_lightup | 17 | 94.1% |
takeaway_query | 16 | 93.8% |
audio_volume_mute | 15 | 100.0% |
social_query | 14 | 92.9% |
qa_maths | 13 | 84.6% |
alarm_remove | 13 | 100.0% |
general_joke | 12 | 100.0% |
email_addcontact | 12 | 91.7% |
iot_hue_lightdim | 11 | 100.0% |
transport_traffic | 11 | 90.9% |
audio_volume_down | 9 | 66.7% |
recommendation_movies | 9 | 66.7% |
iot_wemo_off | 8 | 75.0% |
iot_hue_lighton | 7 | 85.7% |
iot_wemo_on | 6 | 100.0% |
music_settings | 5 | 60.0% |
datetime_convert | 4 | 100.0% |
audio_volume_up | 3 | 100.0% |
By source
Look at hwu64 unseen: requests the base model never saw in its own training. That row is the honest test of how well the adapter generalises.
| Source | Test rows | Qwen3.5-0.8B untrained | Jeff v1.2 0.8B alone | Jeff v1.2 0.8B + adapter |
|---|---|---|---|---|
bitext | 2,361 | 43.6% · 0.260 | 84.0% · 0.082 | 99.9% · 0.001 |
hwu64 seen by the base | 897 | 19.5% · 0.084 | 88.5% · 0.069 | 92.6% · 0.017 |
hwu64 unseen | 1,624 | 17.5% · 0.082 | 83.5% · 0.103 | 93.6% · 0.013 |
snips | 695 | 57.6% · 0.222 | 87.9% · 0.042 | 98.8% · 0.007 |
bitext2,361 test rows- Qwen3.5-0.8B untrained
- 43.6% · 0.260
- Jeff v1.2 0.8B alone
- 84.0% · 0.082
- Jeff v1.2 0.8B + adapter
- 99.9% · 0.001
hwu64 seen by the base897 test rows- Qwen3.5-0.8B untrained
- 19.5% · 0.084
- Jeff v1.2 0.8B alone
- 88.5% · 0.069
- Jeff v1.2 0.8B + adapter
- 92.6% · 0.017
hwu64 unseen1,624 test rows- Qwen3.5-0.8B untrained
- 17.5% · 0.082
- Jeff v1.2 0.8B alone
- 83.5% · 0.103
- Jeff v1.2 0.8B + adapter
- 93.6% · 0.013
snips695 test rows- Qwen3.5-0.8B untrained
- 57.6% · 0.222
- Jeff v1.2 0.8B alone
- 87.9% · 0.042
- Jeff v1.2 0.8B + adapter
- 98.8% · 0.007
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
Against Qwen3.8-27B
More accurate, 55× faster than Qwen3.8-27B alone.
Gain over Qwen3.8-27B alone: +9.3 points [+5.3, +13.3] (95% interval).
On 300 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 86.0% of the time at 6.44 s per query on average; with Jeff and this adapter answering first, it was right 95.3% at 118 ms.
This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.
| Route | Mean | Median | 95th percentile | Typical prompt |
|---|---|---|---|---|
| Qwen3.8-27B alone | 6.44 s | 5.17 s | 9.20 s | 593 tokens |
| Jeff + adapter, its own answer | 118 ms | 91 ms | 168 ms | 560 tokens |
Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.
This task's threshold 0.00: accuracy 95.3%, 54.6× faster, 0.0% sent on to Qwen3.8-27B.
Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 300 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.
How it was trained
- Training rows
- 60,334
- Steps
- 943
- Training time
- 87 min
- Size as saved
- 41.5 MB
One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-support-intents-20260930-1209.
Source: jeff-finetunes/adapters/BASELINE.md
Data card
Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean
- Test set, so anyone can check the numbers
- Calibration rows, the rows its threshold is chosen on
- QA report, sanitised: the data-quality checks run before training
- How the test set was held out
- 10% of the Bitext and HWU64 messages, held out by a stable hash of the text, plus the official SNIPS June 2017 held-out files; never trained on.
- Training data
- Built from public data sets, listed under Data and licence.
- The source data sets are public (listed under Data and licence). A script to rebuild our rows from them will follow.
Data and licence
The adapter is released under Apache-2.0. It was trained on:
- Bitext customer support training data set (27 intents)Licence: CDLA-Sharing-1.0
Only the customer message and intent columns are used; the assistant response column is never read. Under CDLA, models trained on the data are Results, not Data.
- HWU64 (Liu et al. 2019, NLU-Evaluation-Data)Licence: CC-BY-4.0
Home-assistant requests. Four very small intents outside the usual 64 are left out.
- SNIPS custom intent engines benchmark, June 2017 (Coucke et al. 2018)Licence: CC0-1.0
Seven voice-assistant intents.
Changelog
- 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.
Comments
Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.
