Skip to content
JeffHub
SupportOfficialReproducedVersion 0.1.0

triageSupport ticket triage

Routes a support ticket or email to a team and scores its urgency, sentiment and need for a person.

Data: Generated mainly with Qwen3.8-Max (hosted, Alibaba Cloud DashScope), used to finish in time; checked by Qwen3.8-Flash (hosted) and Qwen3.8-Flash-Next (local); urgency and sentiment also rated by DeepSeek-V4-Flash (hosted).

Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap

Use it when

  • You receive tickets, emails or chat messages and need a team, an urgency and a "does a person need to see this?" flag for each.
  • Your teams are described in a sentence each; the adapter reads the descriptions, so your team list can be your own (up to 40 teams in training).
  • You want all four answers from one request.

Not a good fit when

  • You need the reply written for you. Jeff chooses between options; it does not generate text.
  • The decision depends on facts outside the message, such as order history or account status. Put them in the state, or decide in code.
  • Messages mostly arrive in languages other than English. Training had some European-language messages, but most are English.

Request format

The state is an object with these fields, in this order. Only message changes from request to request, so it comes last and the rest can be prepared in advance.

State fieldChanges per requestWhat goes in it
companyNoOne sentence about the organisation the message is sent to.
channelNoWhere the message came from, for example email, web form, chat, app review, social media reply or phone transcript.
messageYesThe ticket or email as received, including subject line, quoted thread and signature.
QuestionTypeWhat it decides
routeChoiceWhich team should handle the message. If it covers several issues, the team for the most important one.

Options: other first (Not for any of these teams), then your teams as k1, k2, … with one line each on what the team handles.

urgencyScoreHow urgently the message needs a response.

Options: Five levels, lowest first, from No action needed to Handle immediately.

sentimentScoreHow the customer feels.

Options: Five levels, from Very negative, angry or distressed to Very positive.

needs_humanYes or noWhether a person must act on it now rather than an automatic reply (escalating complaints, legal or safety issues, requests an automatic system cannot resolve).
  • Ask all four questions in one request; they are answered together.
  • Keep other as the first route option, word for word, and list the teams after it in a fixed order so the unchanging part of the request can be prepared in advance.
  • Use the instructions below word for word; the adapter was trained mostly on them.

General rules for every request are in the request format guide.

Example

The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).

from jeff import Client
from jeff.client import choice_question, score_question, yes_no_question

jeff = Client("http://localhost:8765", model="triage")

state = {
    "company": "A mid-sized online furniture retailer shipping across Europe.",
    "channel": "email",
    "message": "Subject: Crushed wardrobe\n\nHi, the wardrobe I ordered (order 55120) arrived today and the box was crushed. Two doors are split. I want a replacement or my money back. This is the second time.\n\nAnna",
}

answers = jeff.ask(state, {
    "route": choice_question(
        {
            "other": "Not for any of these teams",
            "k1": "Deliveries: late, lost or damaged deliveries, and delivery bookings",
            "k2": "Refunds and payments: refunds, charges and invoices",
            "k3": "Assembly service: booking and complaints about assembly",
            "k4": "Account access: logins, passwords and account details",
        },
        "Which team should handle this message? If it covers several issues, choose the team for the most important one.",
    ),
    "urgency": score_question(
        [
            "No action needed",
            "Can wait a few days",
            "Handle within a day",
            "Handle within hours",
            "Handle immediately",
        ],
        "How urgently does this need a response?",
    ),
    "sentiment": score_question(
        [
            "Very negative, angry or distressed",
            "Negative",
            "Neutral",
            "Positive",
            "Very positive",
        ],
        "How does the customer feel?",
    ),
    "needs_human": yes_no_question("Does this need a person to act on it now, rather than an automatic reply? Answer yes for complaints that could escalate, legal or safety issues, or requests an automatic system cannot resolve."),
})
print("route", answers.choice("route").key)
print("urgency", answers.score("urgency").level)
print("sentiment", answers.score("sentiment").level)
print("needs_human", answers.yes_no("needs_human"))

Response

{
  "model": "triage",
  "answers": {
    "route": {
      "type": "choice",
      "probabilities": {
        "other": 0.0000033690992533612695,
        "k1": 0.993102311783428,
        "k2": 0.006892040008140672,
        "k3": 0.0000019342346550391597,
        "k4": 3.448745229582921e-7
      },
      "choice": "k1",
      "confidence": 0.9913778897292849
    },
    "urgency": {
      "type": "score",
      "probabilities": {
        "0": 0.000005131385466649408,
        "1": 0.012269967207748434,
        "2": 0.8941875079820598,
        "3": 0.09306573101596023,
        "4": 0.00047166240876484817
      },
      "legend": {
        "0": "No action needed",
        "1": "Can wait a few days",
        "2": "Handle within a day",
        "3": "Handle within hours",
        "4": "Handle immediately"
      },
      "score": 2.081728825854808,
      "confidence": 0.9114255951565237
    },
    "sentiment": {
      "type": "score",
      "probabilities": {
        "0": 0.3118397015877296,
        "1": 0.6808165893481221,
        "2": 0.007336929713986554,
        "3": 0.0000065898633241445205,
        "4": 1.8948683766195162e-7
      },
      "legend": {
        "0": "Very negative, angry or distressed",
        "1": "Negative",
        "2": "Neutral",
        "3": "Positive",
        "4": "Very positive"
      },
      "score": 0.6955109763134183,
      "confidence": 0.7340080170926021
    },
    "needs_human": {
      "type": "noul",
      "noul": 0.999893985270675
    }
  },
  "usage": {
    "input_tokens": 800,
    "output_tokens": 0,
    "orders": 1
  }
}

Results

On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters

  • triage7,256 test rows
    Qwen3.5-0.8B untrained
    44.4% · 0.107
    Jeff v1.2 0.8B alone
    67.1% · 0.026
    91.8% · 0.009

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

How sure is it, and is it right?

Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.

When this adapter says it is about 91% sure, it is right about 91% of the time (611 test rows).

Jeff v1.2 0.8B aloneJeff v1.2 0.8B + adapterperfectly calibrated
0%0%25%25%50%50%75%75%100%100%Stated confidenceRight answersJeff v1.2 0.8B alone: 31 rows stated 18% on average and were right 22.6% of the timeJeff v1.2 0.8B alone: 106 rows stated 24% on average and were right 33.0% of the timeJeff v1.2 0.8B alone: 334 rows stated 30% on average and were right 24.0% of the timeJeff v1.2 0.8B alone: 488 rows stated 37% on average and were right 36.9% of the timeJeff v1.2 0.8B alone: 602 rows stated 43% on average and were right 43.4% of the timeJeff v1.2 0.8B alone: 749 rows stated 50% on average and were right 51.7% of the timeJeff v1.2 0.8B alone: 828 rows stated 57% on average and were right 63.4% of the timeJeff v1.2 0.8B alone: 677 rows stated 63% on average and were right 66.9% of the timeJeff v1.2 0.8B alone: 585 rows stated 70% on average and were right 70.8% of the timeJeff v1.2 0.8B alone: 650 rows stated 77% on average and were right 77.8% of the timeJeff v1.2 0.8B alone: 696 rows stated 84% on average and were right 83.3% of the timeJeff v1.2 0.8B alone: 891 rows stated 90% on average and were right 93.5% of the timeJeff v1.2 0.8B alone: 619 rows stated 96% on average and were right 98.7% of the timeJeff v1.2 0.8B + adapter: 1 rows stated 31% on average and were right 0.0% of the timeJeff v1.2 0.8B + adapter: 2 rows stated 37% on average and were right 50.0% of the timeJeff v1.2 0.8B + adapter: 8 rows stated 44% on average and were right 50.0% of the timeJeff v1.2 0.8B + adapter: 122 rows stated 51% on average and were right 37.7% of the timeJeff v1.2 0.8B + adapter: 294 rows stated 57% on average and were right 58.5% of the timeJeff v1.2 0.8B + adapter: 213 rows stated 63% on average and were right 65.7% of the timeJeff v1.2 0.8B + adapter: 290 rows stated 70% on average and were right 71.7% of the timeJeff v1.2 0.8B + adapter: 318 rows stated 77% on average and were right 77.4% of the timeJeff v1.2 0.8B + adapter: 435 rows stated 84% on average and were right 86.7% of the timeJeff v1.2 0.8B + adapter: 611 rows stated 91% on average and were right 90.8% of the timeJeff v1.2 0.8B + adapter: 4962 rows stated 99% on average and were right 98.9% of the time

Point at a dot for its numbers. Bigger dots hold more rows.

Calibration scores
ModelCalibration error (ECE)Brier scoreLog loss
Qwen3.5-0.8B untrained0.1070.6721.378
Jeff v1.2 0.8B alone0.0260.4400.848
Jeff v1.2 0.8B + adapter0.0090.1190.212

All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.

Where it gets things wrong

  • 4 read as 3: 90 rows
  • 2 read as 1: 80 rows
  • 3 read as 4: 74 rows
  • 3 read as 2: 73 rows
  • 1 read as 0: 57 rows

The five commonest mistakes with the adapter, out of 7,256 test rows.

Accuracy per right answer
Right answerTest rowsAccuracy with the adapter
a listed option1,70596.9%
yes1,04498.9%
41,00590.7%
099396.5%
no77098.1%
262180.0%
360375.5%
140672.4%
other10994.5%

“A listed option” pools the rows whose right answer is one of the options listed in that request (keys such as o3 or t1, whose meaning changes from row to row); “another listed option” is a different one of them.

Against Qwen3.8-27B

More accurate, 59× faster than Qwen3.8-27B alone.

Gain over Qwen3.8-27B alone: +9.7 points [+4.7, +14.7] (95% interval).

On 300 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 81.3% of the time at 3.60 s per query on average; with Jeff and this adapter answering first, it was right 91.0% at 61 ms.

This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.

Time per query and prompt length for both routes
RouteMeanMedian95th percentileTypical prompt
Qwen3.8-27B alone3.60 s3.03 s8.38 s352 tokens
Jeff + adapter, its own answer61 ms49 ms132 ms319 tokens

Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.

This task's threshold 0.00: accuracy 91.0%, 59.1× faster, 0.0% sent on to Qwen3.8-27B.

Accuracy91.0%
80.0%90.0%100.0%0.000.250.500.751.00
Speed-up59.1×
0.0×30.0×60.0×0.000.250.500.751.00
Sent on to Qwen3.8-27B0.0%
0.0%50.0%100.0%0.000.250.500.751.00

Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 300 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.

All tasks, and how this was measured

How it was trained

Training rows
73,185
Steps
1,144
Training time
116 min
Size as saved
41.5 MB

One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-triage-20260930-1803.

Source: jeff-finetunes/adapters/BASELINE.md

Not yet measured: typed-decisions, customer_service; Banking77, test split.

Data card

Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean

How the test set was held out
Messages to the 10% of the roughly 300 generated organisations that were never trained on, scored on all four questions (route, urgency, sentiment and whether a person is needed).
Training data
Training data not published.
Which models made the data, counted on the 66,532 training rows
What it didModelWhere it ranTraining rows
Wrote the message teacher.writerQwen3.8-Maxhosted (Alibaba Cloud DashScope)65,368
Wrote the message teacher.writerQwen3.8-Flash-Nextlocal (own hardware)1,164
Edited the message to remove giveaway cues decorrelated_byQwen3.8-Flashhosted (Alibaba Cloud DashScope)34,128
Rated urgency and sentiment score_ratersQwen3.8-Maxhosted (Alibaba Cloud DashScope)66,532
Rated urgency and sentiment score_ratersDeepSeek-V4-Flashhosted (DeepSeek)66,532
Checked the labels teacher.checkerQwen3.8-Flashhosted (Alibaba Cloud DashScope)34,348
Checked the labels teacher.checkerQwen3.8-Flash-Nextlocal (own hardware)32,184
Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. Urgency and sentiment targets are the mean of the generation plan and the two model ratings.
The terms of the hosted model providers are being checked for training and publication use.

Data and licence

The adapter is released under Apache-2.0. It was trained on:

  • Generated organisations and messagesLicence: Released with the adapter under Apache-2.0 · Generated by Qwen3.8-Max (hosted), with some rows by Qwen3.8-Flash-Next (local); see the data card

    About 300 organisations and 40,000 messages written by language models and checked by a second blind pass (which models, and for how many rows, is in the data card). Urgency and sentiment targets are the mean of three ratings.

Changelog

  1. 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.

Comments

Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.