Skip to content
JeffHub
LanguageOfficialReproducedVersion 0.1.0

emotionEmotion in short comments

Picks the strongest of 27 emotions, or neutral, in a short comment or message.

Data: Public data: GoEmotions (Reddit comments)

Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap

Use it when

  • You want a finer reading of feeling than positive or negative, for example gratitude, confusion or disappointment.
  • Your texts are short, informal English comments or messages.
  • You are happy to read the result as a spread of probabilities. Many comments carry more than one emotion.

Not a good fit when

  • Your texts are long, formal or not in English. All training text is English Reddit comments.
  • You need a clinical or safety judgement, such as risk of self-harm. The adapter only names emotions.
  • You need every emotion in a text listed separately. The question asks for the one expressed most.

Request format

The state is an object with these fields, in this order. Only comment changes from request to request, so it comes last and the rest can be prepared in advance.

State fieldChanges per requestWhat goes in it
sourceNoOne short phrase on where the text comes from. In training this was always "Reddit comment ([NAME] and [RELIGION] replace removed names)".
commentYesThe comment or message to read.
QuestionTypeWhat it decides
emotionChoiceWhich emotion the comment expresses most strongly, or neutral if it expresses no particular emotion.

Options: 28 options: 27 emotions and neutral, keyed by their GoEmotions names (admiration, amusement, anger and so on), each with a one-line description. The full list is GOEMOTIONS in descriptions.py in the source.

  • Use the option keys and descriptions from descriptions.py; the adapter was trained on them.
  • Comments with several gold emotions were trained with the probability spread evenly over them, so a split answer is expected, not a fault.
  • Use the instructions below word for word; the adapter was trained mostly on them.

General rules for every request are in the request format guide.

Example

The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).

from jeff import Client
from jeff.client import choice_question

jeff = Client("http://localhost:8765", model="emotion")

state = {
    "source": "Reddit comment ([NAME] and [RELIGION] replace removed names)",
    "comment": "Thanks so much for the tip, I had no idea the library lent out tools. Saved me a fortune.",
}

answers = jeff.ask(state, {
    "emotion": choice_question(
        {
            "gratitude": "Gratitude: feeling thankful or appreciative.",
            "joy": "Joy: a feeling of pleasure and happiness.",
            "surprise": "Surprise: being astonished or startled by something unexpected.",
            "realization": "Realization: becoming aware of something.",
            "admiration": "Admiration: finding something impressive or worthy of respect.",
            "relief": "Relief: reassurance and relaxation after anxiety or distress ends.",
            "annoyance": "Annoyance: mild anger or irritation.",
            "confusion": "Confusion: lack of understanding or uncertainty.",
            "sadness": "Sadness: feeling sorrow or unhappiness.",
            "neutral": "Neutral: no particular emotion is expressed.",
        },
        "Which emotion does the comment express most? Choose the emotion that is strongest in the comment, or neutral if it expresses no particular emotion.",
    ),
})
print("emotion", answers.choice("emotion").key)

Response

{
  "model": "emotion",
  "answers": {
    "emotion": {
      "type": "choice",
      "probabilities": {
        "gratitude": 0.9377514192297772,
        "joy": 0.0048510179566744254,
        "surprise": 0.0013096331149660347,
        "realization": 0.00906897121104442,
        "admiration": 0.005445419730876873,
        "relief": 0.00815921427750748,
        "annoyance": 0.0002489432112212598,
        "confusion": 0.03209981177197163,
        "sadness": 0.00015036246233088356,
        "neutral": 0.0009152070336297481
      },
      "choice": "gratitude",
      "confidence": 0.9308349102553081
    }
  },
  "usage": {
    "input_tokens": 285,
    "output_tokens": 0,
    "orders": 1
  }
}

Results

On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters

  • emotion5,408 test rows
    Qwen3.5-0.8B untrained
    12.5% · 0.044
    Jeff v1.2 0.8B alone
    32.2% · 0.038
    60.6% · 0.020

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

How sure is it, and is it right?

Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.

When this adapter says it is about 50% sure, it is right about 52% of the time (676 test rows).

Jeff v1.2 0.8B aloneJeff v1.2 0.8B + adapterperfectly calibrated
0%0%25%25%50%50%75%75%100%100%Stated confidenceRight answersJeff v1.2 0.8B alone: 86 rows stated 12% on average and were right 12.8% of the timeJeff v1.2 0.8B alone: 738 rows stated 17% on average and were right 20.6% of the timeJeff v1.2 0.8B alone: 1169 rows stated 23% on average and were right 26.2% of the timeJeff v1.2 0.8B alone: 1107 rows stated 30% on average and were right 27.6% of the timeJeff v1.2 0.8B alone: 822 rows stated 36% on average and were right 32.5% of the timeJeff v1.2 0.8B alone: 572 rows stated 43% on average and were right 38.3% of the timeJeff v1.2 0.8B alone: 341 rows stated 50% on average and were right 44.0% of the timeJeff v1.2 0.8B alone: 225 rows stated 57% on average and were right 50.7% of the timeJeff v1.2 0.8B alone: 138 rows stated 63% on average and were right 56.5% of the timeJeff v1.2 0.8B alone: 103 rows stated 70% on average and were right 59.2% of the timeJeff v1.2 0.8B alone: 58 rows stated 76% on average and were right 65.5% of the timeJeff v1.2 0.8B alone: 34 rows stated 83% on average and were right 85.3% of the timeJeff v1.2 0.8B alone: 13 rows stated 88% on average and were right 84.6% of the timeJeff v1.2 0.8B alone: 2 rows stated 96% on average and were right 100.0% of the timeJeff v1.2 0.8B + adapter: 7 rows stated 19% on average and were right 28.6% of the timeJeff v1.2 0.8B + adapter: 142 rows stated 24% on average and were right 25.4% of the timeJeff v1.2 0.8B + adapter: 347 rows stated 30% on average and were right 30.0% of the timeJeff v1.2 0.8B + adapter: 539 rows stated 37% on average and were right 41.6% of the timeJeff v1.2 0.8B + adapter: 654 rows stated 43% on average and were right 46.2% of the timeJeff v1.2 0.8B + adapter: 676 rows stated 50% on average and were right 52.4% of the timeJeff v1.2 0.8B + adapter: 578 rows stated 56% on average and were right 56.4% of the timeJeff v1.2 0.8B + adapter: 522 rows stated 63% on average and were right 64.9% of the timeJeff v1.2 0.8B + adapter: 440 rows stated 70% on average and were right 72.5% of the timeJeff v1.2 0.8B + adapter: 462 rows stated 77% on average and were right 77.3% of the timeJeff v1.2 0.8B + adapter: 407 rows stated 83% on average and were right 80.1% of the timeJeff v1.2 0.8B + adapter: 360 rows stated 90% on average and were right 91.1% of the timeJeff v1.2 0.8B + adapter: 274 rows stated 96% on average and were right 95.3% of the time

Point at a dot for its numbers. Bigger dots hold more rows.

Calibration scores
ModelCalibration error (ECE)Brier scoreLog loss
Qwen3.5-0.8B untrained0.0440.9533.232
Jeff v1.2 0.8B alone0.0380.8172.241
Jeff v1.2 0.8B + adapter0.0200.5391.236

All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.

Where it gets things wrong

  • approval read as neutral: 127 rows
  • annoyance read as neutral: 122 rows
  • disapproval read as neutral: 105 rows
  • neutral read as curiosity: 80 rows
  • curiosity read as neutral: 63 rows

The five commonest mistakes with the adapter, out of 5,408 test rows.

Accuracy per right answer
Right answerTest rowsAccuracy with the adapter
neutral1,68278.5%
admiration42273.5%
gratitude29490.1%
approval28126.7%
annoyance26119.5%
disapproval23233.2%
amusement22784.6%
curiosity22554.7%
love19882.8%
anger16648.2%
Show all 28 answers
Accuracy per right answer, all answers
Right answerTest rowsAccuracy with the adapter
neutral1,68278.5%
admiration42273.5%
gratitude29490.1%
approval28126.7%
annoyance26119.5%
disapproval23233.2%
amusement22784.6%
curiosity22554.7%
love19882.8%
anger16648.2%
optimism14653.4%
sadness12953.5%
confusion12834.4%
joy12857.0%
realization12417.7%
caring11028.2%
disappointment10924.8%
surprise10854.6%
disgust9645.8%
excitement7829.5%
fear7273.6%
desire6944.9%
remorse5172.5%
embarrassment3046.7%
nervousness1752.9%
pride1315.4%
relief825.0%
grief450.0%

Against Qwen3.8-27B

More accurate, 42× faster than Qwen3.8-27B alone.

Gain over Qwen3.8-27B alone: +25.0 points [+20.0, +30.0] (95% interval).

On 500 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 35.6% of the time at 4.79 s per query on average; with Jeff and this adapter answering first, it was right 60.6% at 113 ms.

This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.

Time per query and prompt length for both routes
RouteMeanMedian95th percentileTypical prompt
Qwen3.8-27B alone4.79 s4.83 s5.23 s617 tokens
Jeff + adapter, its own answer113 ms113 ms117 ms584 tokens

Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.

This task's threshold 0.00: accuracy 60.6%, 42.5× faster, 0.0% sent on to Qwen3.8-27B.

Accuracy60.6%
35.0%67.5%100.0%0.000.250.500.751.00
Speed-up42.5×
0.0×25.0×50.0×0.000.250.500.751.00
Sent on to Qwen3.8-27B0.0%
0.0%50.0%100.0%0.000.250.500.751.00

Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 500 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.

All tasks, and how this was measured

How it was trained

Training rows
50,674
Steps
792
Training time
81 min
Size as saved
41.5 MB

One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-emotion-20260930-0623.

Source: jeff-finetunes/adapters/BASELINE.md

Data card

Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean

How the test set was held out
The official GoEmotions test split (simplified configuration); never trained on.
Training data
Built from public data sets, listed under Data and licence.
The source data sets are public (listed under Data and licence). A script to rebuild our rows from them will follow.

Data and licence

The adapter is released under Apache-2.0. It was trained on:

Changelog

  1. 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.

Comments

Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.