triageSupport ticket triage
Routes a support ticket or email to a team and scores its urgency, sentiment and need for a person.
Data: Generated mainly with Qwen3.8-Max (hosted, Alibaba Cloud DashScope), used to finish in time; checked by Qwen3.8-Flash (hosted) and Qwen3.8-Flash-Next (local); urgency and sentiment also rated by DeepSeek-V4-Flash (hosted).
Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap
Use it when
- You receive tickets, emails or chat messages and need a team, an urgency and a "does a person need to see this?" flag for each.
- Your teams are described in a sentence each; the adapter reads the descriptions, so your team list can be your own (up to 40 teams in training).
- You want all four answers from one request.
Not a good fit when
- You need the reply written for you. Jeff chooses between options; it does not generate text.
- The decision depends on facts outside the message, such as order history or account status. Put them in the state, or decide in code.
- Messages mostly arrive in languages other than English. Training had some European-language messages, but most are English.
Request format
The state is an object with these fields, in this order. Only message changes from request to request, so it comes last and the rest can be prepared in advance.
| State field | Changes per request | What goes in it |
|---|---|---|
company | No | One sentence about the organisation the message is sent to. |
channel | No | Where the message came from, for example email, web form, chat, app review, social media reply or phone transcript. |
message | Yes | The ticket or email as received, including subject line, quoted thread and signature. |
| Question | Type | What it decides |
|---|---|---|
route | Choice | Which team should handle the message. If it covers several issues, the team for the most important one. Options: |
urgency | Score | How urgently the message needs a response. Options: Five levels, lowest first, from No action needed to Handle immediately. |
sentiment | Score | How the customer feels. Options: Five levels, from Very negative, angry or distressed to Very positive. |
needs_human | Yes or no | Whether a person must act on it now rather than an automatic reply (escalating complaints, legal or safety issues, requests an automatic system cannot resolve). |
- Ask all four questions in one request; they are answered together.
- Keep
otheras the first route option, word for word, and list the teams after it in a fixed order so the unchanging part of the request can be prepared in advance. - Use the instructions below word for word; the adapter was trained mostly on them.
General rules for every request are in the request format guide.
Example
The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).
from jeff import Client
from jeff.client import choice_question, score_question, yes_no_question
jeff = Client("http://localhost:8765", model="triage")
state = {
"company": "A mid-sized online furniture retailer shipping across Europe.",
"channel": "email",
"message": "Subject: Crushed wardrobe\n\nHi, the wardrobe I ordered (order 55120) arrived today and the box was crushed. Two doors are split. I want a replacement or my money back. This is the second time.\n\nAnna",
}
answers = jeff.ask(state, {
"route": choice_question(
{
"other": "Not for any of these teams",
"k1": "Deliveries: late, lost or damaged deliveries, and delivery bookings",
"k2": "Refunds and payments: refunds, charges and invoices",
"k3": "Assembly service: booking and complaints about assembly",
"k4": "Account access: logins, passwords and account details",
},
"Which team should handle this message? If it covers several issues, choose the team for the most important one.",
),
"urgency": score_question(
[
"No action needed",
"Can wait a few days",
"Handle within a day",
"Handle within hours",
"Handle immediately",
],
"How urgently does this need a response?",
),
"sentiment": score_question(
[
"Very negative, angry or distressed",
"Negative",
"Neutral",
"Positive",
"Very positive",
],
"How does the customer feel?",
),
"needs_human": yes_no_question("Does this need a person to act on it now, rather than an automatic reply? Answer yes for complaints that could escalate, legal or safety issues, or requests an automatic system cannot resolve."),
})
print("route", answers.choice("route").key)
print("urgency", answers.score("urgency").level)
print("sentiment", answers.score("sentiment").level)
print("needs_human", answers.yes_no("needs_human"))import { Client, choiceQuestion, scoreQuestion, yesNoQuestion } from '@jeff/client';
const jeff = new Client({ url: 'http://localhost:8765', model: 'triage' });
const state = {
company: 'A mid-sized online furniture retailer shipping across Europe.',
channel: 'email',
message: 'Subject: Crushed wardrobe\n\nHi, the wardrobe I ordered (order 55120) arrived today and the box was crushed. Two doors are split. I want a replacement or my money back. This is the second time.\n\nAnna',
};
const answers = await jeff.ask(state, {
route: choiceQuestion(
{
other: 'Not for any of these teams',
k1: 'Deliveries: late, lost or damaged deliveries, and delivery bookings',
k2: 'Refunds and payments: refunds, charges and invoices',
k3: 'Assembly service: booking and complaints about assembly',
k4: 'Account access: logins, passwords and account details',
},
'Which team should handle this message? If it covers several issues, choose the team for the most important one.',
),
urgency: scoreQuestion(
[
'No action needed',
'Can wait a few days',
'Handle within a day',
'Handle within hours',
'Handle immediately',
],
'How urgently does this need a response?',
),
sentiment: scoreQuestion(
[
'Very negative, angry or distressed',
'Negative',
'Neutral',
'Positive',
'Very positive',
],
'How does the customer feel?',
),
needs_human: yesNoQuestion('Does this need a person to act on it now, rather than an automatic reply? Answer yes for complaints that could escalate, legal or safety issues, or requests an automatic system cannot resolve.'),
});
console.log('route', answers.route.key);
console.log('urgency', answers.urgency.level);
console.log('sentiment', answers.sentiment.level);
console.log('needs_human', answers.needs_human);curl -s http://localhost:8765/v1/systemone \
-H 'content-type: application/json' \
-d '{
"model": "triage",
"state": {
"company": "A mid-sized online furniture retailer shipping across Europe.",
"channel": "email",
"message": "Subject: Crushed wardrobe\n\nHi, the wardrobe I ordered (order 55120) arrived today and the box was crushed. Two doors are split. I want a replacement or my money back. This is the second time.\n\nAnna"
},
"questions": {
"route": {
"type": "choice",
"instructions": "Which team should handle this message? If it covers several issues, choose the team for the most important one.",
"criteria": {
"other": "Not for any of these teams",
"k1": "Deliveries: late, lost or damaged deliveries, and delivery bookings",
"k2": "Refunds and payments: refunds, charges and invoices",
"k3": "Assembly service: booking and complaints about assembly",
"k4": "Account access: logins, passwords and account details"
}
},
"urgency": {
"type": "score",
"instructions": "How urgently does this need a response?",
"criteria": [
"No action needed",
"Can wait a few days",
"Handle within a day",
"Handle within hours",
"Handle immediately"
]
},
"sentiment": {
"type": "score",
"instructions": "How does the customer feel?",
"criteria": [
"Very negative, angry or distressed",
"Negative",
"Neutral",
"Positive",
"Very positive"
]
},
"needs_human": {
"type": "noul",
"instructions": "Does this need a person to act on it now, rather than an automatic reply? Answer yes for complaints that could escalate, legal or safety issues, or requests an automatic system cannot resolve."
}
}
}'Response
{
"model": "triage",
"answers": {
"route": {
"type": "choice",
"probabilities": {
"other": 0.0000033690992533612695,
"k1": 0.993102311783428,
"k2": 0.006892040008140672,
"k3": 0.0000019342346550391597,
"k4": 3.448745229582921e-7
},
"choice": "k1",
"confidence": 0.9913778897292849
},
"urgency": {
"type": "score",
"probabilities": {
"0": 0.000005131385466649408,
"1": 0.012269967207748434,
"2": 0.8941875079820598,
"3": 0.09306573101596023,
"4": 0.00047166240876484817
},
"legend": {
"0": "No action needed",
"1": "Can wait a few days",
"2": "Handle within a day",
"3": "Handle within hours",
"4": "Handle immediately"
},
"score": 2.081728825854808,
"confidence": 0.9114255951565237
},
"sentiment": {
"type": "score",
"probabilities": {
"0": 0.3118397015877296,
"1": 0.6808165893481221,
"2": 0.007336929713986554,
"3": 0.0000065898633241445205,
"4": 1.8948683766195162e-7
},
"legend": {
"0": "Very negative, angry or distressed",
"1": "Negative",
"2": "Neutral",
"3": "Positive",
"4": "Very positive"
},
"score": 0.6955109763134183,
"confidence": 0.7340080170926021
},
"needs_human": {
"type": "noul",
"noul": 0.999893985270675
}
},
"usage": {
"input_tokens": 800,
"output_tokens": 0,
"orders": 1
}
}Results
On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters
| Test set | Test rows | Qwen3.5-0.8B untrained | Jeff v1.2 0.8B alone | Jeff v1.2 0.8B + adapter |
|---|---|---|---|---|
triage | 7,256 | 44.4% · 0.107 | 67.1% · 0.026 | 91.8% · 0.009 |
triage7,256 test rows- Qwen3.5-0.8B untrained
- 44.4% · 0.107
- Jeff v1.2 0.8B alone
- 67.1% · 0.026
- Jeff v1.2 0.8B + adapter
- 91.8% · 0.009
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
How sure is it, and is it right?
Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.
When this adapter says it is about 91% sure, it is right about 91% of the time (611 test rows).
Point at a dot for its numbers. Bigger dots hold more rows.
| Model | Calibration error (ECE) | Brier score | Log loss |
|---|---|---|---|
| Qwen3.5-0.8B untrained | 0.107 | 0.672 | 1.378 |
| Jeff v1.2 0.8B alone | 0.026 | 0.440 | 0.848 |
| Jeff v1.2 0.8B + adapter | 0.009 | 0.119 | 0.212 |
All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.
Where it gets things wrong
4read as3: 90 rows2read as1: 80 rows3read as4: 74 rows3read as2: 73 rows1read as0: 57 rows
The five commonest mistakes with the adapter, out of 7,256 test rows.
| Right answer | Test rows | Accuracy with the adapter |
|---|---|---|
| a listed option | 1,705 | 96.9% |
| yes | 1,044 | 98.9% |
4 | 1,005 | 90.7% |
0 | 993 | 96.5% |
| no | 770 | 98.1% |
2 | 621 | 80.0% |
3 | 603 | 75.5% |
1 | 406 | 72.4% |
other | 109 | 94.5% |
“A listed option” pools the rows whose right answer is one of the options listed in that request (keys such as o3 or t1, whose meaning changes from row to row); “another listed option” is a different one of them.
Against Qwen3.8-27B
More accurate, 59× faster than Qwen3.8-27B alone.
Gain over Qwen3.8-27B alone: +9.7 points [+4.7, +14.7] (95% interval).
On 300 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 81.3% of the time at 3.60 s per query on average; with Jeff and this adapter answering first, it was right 91.0% at 61 ms.
This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.
| Route | Mean | Median | 95th percentile | Typical prompt |
|---|---|---|---|---|
| Qwen3.8-27B alone | 3.60 s | 3.03 s | 8.38 s | 352 tokens |
| Jeff + adapter, its own answer | 61 ms | 49 ms | 132 ms | 319 tokens |
Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.
This task's threshold 0.00: accuracy 91.0%, 59.1× faster, 0.0% sent on to Qwen3.8-27B.
Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 300 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.
How it was trained
- Training rows
- 73,185
- Steps
- 1,144
- Training time
- 116 min
- Size as saved
- 41.5 MB
One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-triage-20260930-1803.
Source: jeff-finetunes/adapters/BASELINE.md
Not yet measured: typed-decisions, customer_service; Banking77, test split.
Data card
Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean
- Test set, so anyone can check the numbers
- Calibration rows, the rows its threshold is chosen on
- QA report, sanitised: the data-quality checks run before training
- How the test set was held out
- Messages to the 10% of the roughly 300 generated organisations that were never trained on, scored on all four questions (route, urgency, sentiment and whether a person is needed).
- Training data
- Training data not published.
- Which models made the data, counted on the 66,532 training rows
- Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. Urgency and sentiment targets are the mean of the generation plan and the two model ratings.
What it did Model Where it ran Training rows Wrote the message teacher.writerQwen3.8-Max hosted (Alibaba Cloud DashScope) 65,368 Wrote the message teacher.writerQwen3.8-Flash-Next local (own hardware) 1,164 Edited the message to remove giveaway cues decorrelated_byQwen3.8-Flash hosted (Alibaba Cloud DashScope) 34,128 Rated urgency and sentiment score_ratersQwen3.8-Max hosted (Alibaba Cloud DashScope) 66,532 Rated urgency and sentiment score_ratersDeepSeek-V4-Flash hosted (DeepSeek) 66,532 Checked the labels teacher.checkerQwen3.8-Flash hosted (Alibaba Cloud DashScope) 34,348 Checked the labels teacher.checkerQwen3.8-Flash-Next local (own hardware) 32,184 - The terms of the hosted model providers are being checked for training and publication use.
Data and licence
The adapter is released under Apache-2.0. It was trained on:
- Generated organisations and messagesLicence: Released with the adapter under Apache-2.0 · Generated by Qwen3.8-Max (hosted), with some rows by Qwen3.8-Flash-Next (local); see the data card
About 300 organisations and 40,000 messages written by language models and checked by a second blind pass (which models, and for how many rows, is in the data card). Urgency and sentiment targets are the mean of three ratings.
Changelog
- 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.
Comments
Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.
