Skip to content
JeffHub
SupportOfficialReproducedVersion 0.1.0

support-intentsCustomer request intents

Names what a customer or assistant user is asking for, from a fixed list of requests such as cancel_order or track_refund.

Data: Public data: Bitext, HWU64 and SNIPS

Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap

Use it when

  • You run an online shop chat or a voice assistant and want each message matched to one request from a fixed list.
  • Your requests are close to the ones in training (27 online-shop requests, 64 home-assistant requests, 7 voice-assistant requests), each described in one line.
  • Messages are short and in English.

Not a good fit when

  • Your list of requests is very different from the training lists. Try it, but measure first; the triage adapter reads your own team descriptions.
  • A message holds several requests and you need all of them. The question picks one.
  • You need the reply written. Jeff chooses between options; it does not generate text.
  • Your messages are long emails or tickets. Training messages are short single requests.

Request format

The state is an object with these fields, in this order. Only message changes from request to request, so it comes last and the rest can be prepared in advance.

State fieldChanges per requestWhat goes in it
serviceNoOne short phrase about the service the message is sent to. Training used three, one per data set, for example "Customer support chat of an online shop".
messageYesThe customer's or user's message as received.
QuestionTypeWhat it decides
intentChoiceWhat the customer wants, as the request that best matches their message.

Options: In training, each message's options were all the requests of its own data set: 27 for the online shop (Bitext), 64 for the home assistant (HWU64) and 7 for the voice assistant (SNIPS). Keys are snake_case names such as track_order, each with a one-line description. The full lists are BITEXT, HWU64 and SNIPS in descriptions.py in the source.

  • Use the option keys and descriptions from descriptions.py where they fit your service; the adapter was trained on them.
  • Use the instructions below word for word; the adapter was trained mostly on them.

General rules for every request are in the request format guide.

Example

The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).

from jeff import Client
from jeff.client import choice_question

jeff = Client("http://localhost:8765", model="support-intents")

state = {
    "service": "Customer support chat of an online shop",
    "message": "hi, I sent back the jacket two weeks ago and still haven't seen the money. where is it?",
}

answers = jeff.ask(state, {
    "intent": choice_question(
        {
            "track_refund": "Check the status of a refund they are expecting.",
            "get_refund": "Get their money back for a purchase.",
            "check_refund_policy": "Learn the refund policy and whether they qualify for a refund.",
            "track_order": "Find out where their order is or its current status.",
            "cancel_order": "Cancel an order they placed.",
            "change_order": "Change an existing order (for example add, remove or swap items).",
            "payment_issue": "Report or solve a problem with a payment.",
            "complaint": "Make a complaint about the product, service or company.",
            "contact_human_agent": "Talk to a human agent instead of an automated assistant.",
            "delivery_period": "Find out when an order will arrive or how long delivery takes.",
        },
        "What does the customer want? Choose the request that best matches what the customer is asking for in their message.",
    ),
})
print("intent", answers.choice("intent").key)

Response

{
  "model": "support-intents",
  "answers": {
    "intent": {
      "type": "choice",
      "probabilities": {
        "track_refund": 0.9028009763827044,
        "get_refund": 0.013336566264330635,
        "check_refund_policy": 0.00008649192014740069,
        "track_order": 0.08268026999819446,
        "cancel_order": 0.00022875498667774524,
        "change_order": 0.00014981366970950356,
        "payment_issue": 0.00027732248392003673,
        "complaint": 0.00010125868711429669,
        "contact_human_agent": 0.00022202538072712725,
        "delivery_period": 0.0001165202264744504
      },
      "choice": "track_refund",
      "confidence": 0.8920010848696716
    }
  },
  "usage": {
    "input_tokens": 297,
    "output_tokens": 0,
    "orders": 1
  }
}

Results

On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters

  • support-intents5,577 test rows
    Qwen3.5-0.8B untrained
    33.9% · 0.162
    Jeff v1.2 0.8B alone
    85.1% · 0.080
    96.8% · 0.006

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

How sure is it, and is it right?

Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.

When this adapter says it is about 99.4% sure, it is right about 99.4% of the time (4,985 test rows).

Jeff v1.2 0.8B aloneJeff v1.2 0.8B + adapterperfectly calibrated
0%0%25%25%50%50%75%75%100%100%Stated confidenceRight answersJeff v1.2 0.8B alone: 3 rows stated 6% on average and were right 0.0% of the timeJeff v1.2 0.8B alone: 25 rows stated 11% on average and were right 32.0% of the timeJeff v1.2 0.8B alone: 50 rows stated 17% on average and were right 30.0% of the timeJeff v1.2 0.8B alone: 91 rows stated 24% on average and were right 38.5% of the timeJeff v1.2 0.8B alone: 174 rows stated 30% on average and were right 51.1% of the timeJeff v1.2 0.8B alone: 202 rows stated 37% on average and were right 47.5% of the timeJeff v1.2 0.8B alone: 253 rows stated 43% on average and were right 65.2% of the timeJeff v1.2 0.8B alone: 288 rows stated 50% on average and were right 66.3% of the timeJeff v1.2 0.8B alone: 290 rows stated 57% on average and were right 67.6% of the timeJeff v1.2 0.8B alone: 276 rows stated 64% on average and were right 80.1% of the timeJeff v1.2 0.8B alone: 307 rows stated 70% on average and were right 78.5% of the timeJeff v1.2 0.8B alone: 329 rows stated 77% on average and were right 86.0% of the timeJeff v1.2 0.8B alone: 489 rows stated 83% on average and were right 92.2% of the timeJeff v1.2 0.8B alone: 814 rows stated 90% on average and were right 96.4% of the timeJeff v1.2 0.8B alone: 1986 rows stated 97% on average and were right 99.1% of the timeJeff v1.2 0.8B + adapter: 1 rows stated 19% on average and were right 100.0% of the timeJeff v1.2 0.8B + adapter: 2 rows stated 23% on average and were right 50.0% of the timeJeff v1.2 0.8B + adapter: 11 rows stated 30% on average and were right 27.3% of the timeJeff v1.2 0.8B + adapter: 13 rows stated 37% on average and were right 61.5% of the timeJeff v1.2 0.8B + adapter: 31 rows stated 44% on average and were right 51.6% of the timeJeff v1.2 0.8B + adapter: 52 rows stated 50% on average and were right 51.9% of the timeJeff v1.2 0.8B + adapter: 42 rows stated 56% on average and were right 66.7% of the timeJeff v1.2 0.8B + adapter: 57 rows stated 63% on average and were right 66.7% of the timeJeff v1.2 0.8B + adapter: 50 rows stated 70% on average and were right 74.0% of the timeJeff v1.2 0.8B + adapter: 68 rows stated 77% on average and were right 79.4% of the timeJeff v1.2 0.8B + adapter: 97 rows stated 83% on average and were right 86.6% of the timeJeff v1.2 0.8B + adapter: 168 rows stated 91% on average and were right 85.1% of the timeJeff v1.2 0.8B + adapter: 4985 rows stated 99% on average and were right 99.4% of the time

Point at a dot for its numbers. Bigger dots hold more rows.

Calibration scores
ModelCalibration error (ECE)Brier scoreLog loss
Qwen3.5-0.8B untrained0.1620.8643.097
Jeff v1.2 0.8B alone0.0800.2290.538
Jeff v1.2 0.8B + adapter0.0060.0520.124

All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.

Where it gets things wrong

  • general_quirky read as news_query: 7 rows
  • qa_factoid read as general_quirky: 5 rows
  • general_quirky read as qa_factoid: 5 rows
  • calendar_query read as general_quirky: 5 rows
  • search_screening_event read as search_creative_work: 5 rows

The five commonest mistakes with the adapter, out of 5,577 test rows.

Accuracy per right answer
Right answerTest rowsAccuracy with the adapter
play_music22096.4%
calendar_set15095.3%
payment_issue130100.0%
check_refund_policy124100.0%
general_quirky11570.4%
complaint106100.0%
general_negate10499.0%
check_payment_methods104100.0%
delivery_period103100.0%
registration_problems101100.0%
Show all 97 answers
Accuracy per right answer, all answers
Right answerTest rowsAccuracy with the adapter
play_music22096.4%
calendar_set15095.3%
payment_issue130100.0%
check_refund_policy124100.0%
general_quirky11570.4%
complaint106100.0%
general_negate10499.0%
check_payment_methods104100.0%
delivery_period103100.0%
registration_problems101100.0%
search_creative_work100100.0%
add_to_playlist100100.0%
book_restaurant100100.0%
qa_factoid9985.9%
get_weather99100.0%
search_screening_event9894.9%
rate_book9899.0%
change_shipping_address97100.0%
newsletter_subscription96100.0%
review95100.0%
get_invoice94100.0%
contact_human_agent93100.0%
check_cancellation_fee90100.0%
delete_account89100.0%
check_invoice88100.0%
recover_password88100.0%
weather_query88100.0%
contact_customer_service85100.0%
set_up_shipping_address84100.0%
calendar_query8485.7%
get_refund84100.0%
place_order82100.0%
create_account8298.8%
switch_account81100.0%
edit_account78100.0%
general_praise76100.0%
change_order7598.7%
email_sendemail7593.3%
news_query7195.8%
email_query6998.6%
general_affirm69100.0%
track_order64100.0%
general_explain64100.0%
general_repeat63100.0%
datetime_query6395.2%
delivery_options62100.0%
track_refund60100.0%
general_dontcare58100.0%
calendar_remove5698.2%
social_post5394.3%
general_confirm53100.0%
play_radio4985.7%
cooking_recipe4795.7%
qa_definition4793.6%
qa_currency44100.0%
play_podcasts4195.1%
general_commandstop38100.0%
lists_remove3591.4%
transport_query3588.6%
recommendation_events3183.9%
recommendation_locations3086.7%
lists_query3093.3%
qa_stock3096.7%
alarm_set3096.7%
music_query3080.0%
lists_createoradd2792.6%
iot_hue_lightoff26100.0%
cancel_order26100.0%
transport_ticket2596.0%
play_audiobook2487.5%
iot_cleaning23100.0%
email_querycontact2387.0%
play_game2295.5%
transport_taxi2195.2%
alarm_query2190.5%
takeaway_order2070.0%
iot_coffee20100.0%
music_likeness1994.7%
iot_hue_lightchange1888.9%
iot_hue_lightup1794.1%
takeaway_query1693.8%
audio_volume_mute15100.0%
social_query1492.9%
qa_maths1384.6%
alarm_remove13100.0%
general_joke12100.0%
email_addcontact1291.7%
iot_hue_lightdim11100.0%
transport_traffic1190.9%
audio_volume_down966.7%
recommendation_movies966.7%
iot_wemo_off875.0%
iot_hue_lighton785.7%
iot_wemo_on6100.0%
music_settings560.0%
datetime_convert4100.0%
audio_volume_up3100.0%

By source

Look at hwu64 unseen: requests the base model never saw in its own training. That row is the honest test of how well the adapter generalises.

  • bitext2,361 test rows
    Qwen3.5-0.8B untrained
    43.6% · 0.260
    Jeff v1.2 0.8B alone
    84.0% · 0.082
    99.9% · 0.001
  • hwu64 seen by the base897 test rows
    Qwen3.5-0.8B untrained
    19.5% · 0.084
    Jeff v1.2 0.8B alone
    88.5% · 0.069
    92.6% · 0.017
  • hwu64 unseen1,624 test rows
    Qwen3.5-0.8B untrained
    17.5% · 0.082
    Jeff v1.2 0.8B alone
    83.5% · 0.103
    93.6% · 0.013
  • snips695 test rows
    Qwen3.5-0.8B untrained
    57.6% · 0.222
    Jeff v1.2 0.8B alone
    87.9% · 0.042
    98.8% · 0.007

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

Against Qwen3.8-27B

More accurate, 55× faster than Qwen3.8-27B alone.

Gain over Qwen3.8-27B alone: +9.3 points [+5.3, +13.3] (95% interval).

On 300 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 86.0% of the time at 6.44 s per query on average; with Jeff and this adapter answering first, it was right 95.3% at 118 ms.

This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.

Time per query and prompt length for both routes
RouteMeanMedian95th percentileTypical prompt
Qwen3.8-27B alone6.44 s5.17 s9.20 s593 tokens
Jeff + adapter, its own answer118 ms91 ms168 ms560 tokens

Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.

This task's threshold 0.00: accuracy 95.3%, 54.6× faster, 0.0% sent on to Qwen3.8-27B.

Accuracy95.3%
85.0%92.5%100.0%0.000.250.500.751.00
Speed-up54.6×
0.0×30.0×60.0×0.000.250.500.751.00
Sent on to Qwen3.8-27B0.0%
0.0%50.0%100.0%0.000.250.500.751.00

Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 300 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.

All tasks, and how this was measured

How it was trained

Training rows
60,334
Steps
943
Training time
87 min
Size as saved
41.5 MB

One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-support-intents-20260930-1209.

Source: jeff-finetunes/adapters/BASELINE.md

Data card

Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean

How the test set was held out
10% of the Bitext and HWU64 messages, held out by a stable hash of the text, plus the official SNIPS June 2017 held-out files; never trained on.
Training data
Built from public data sets, listed under Data and licence.
The source data sets are public (listed under Data and licence). A script to rebuild our rows from them will follow.

Data and licence

The adapter is released under Apache-2.0. It was trained on:

Changelog

  1. 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.

Comments

Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.