emotionEmotion in short comments
Picks the strongest of 27 emotions, or neutral, in a short comment or message.
Data: Public data: GoEmotions (Reddit comments)
Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap
Use it when
- You want a finer reading of feeling than positive or negative, for example gratitude, confusion or disappointment.
- Your texts are short, informal English comments or messages.
- You are happy to read the result as a spread of probabilities. Many comments carry more than one emotion.
Not a good fit when
- Your texts are long, formal or not in English. All training text is English Reddit comments.
- You need a clinical or safety judgement, such as risk of self-harm. The adapter only names emotions.
- You need every emotion in a text listed separately. The question asks for the one expressed most.
Request format
The state is an object with these fields, in this order. Only comment changes from request to request, so it comes last and the rest can be prepared in advance.
| State field | Changes per request | What goes in it |
|---|---|---|
source | No | One short phrase on where the text comes from. In training this was always "Reddit comment ([NAME] and [RELIGION] replace removed names)". |
comment | Yes | The comment or message to read. |
| Question | Type | What it decides |
|---|---|---|
emotion | Choice | Which emotion the comment expresses most strongly, or neutral if it expresses no particular emotion. Options: 28 options: 27 emotions and neutral, keyed by their GoEmotions names (admiration, amusement, anger and so on), each with a one-line description. The full list is GOEMOTIONS in descriptions.py in the source. |
- Use the option keys and descriptions from descriptions.py; the adapter was trained on them.
- Comments with several gold emotions were trained with the probability spread evenly over them, so a split answer is expected, not a fault.
- Use the instructions below word for word; the adapter was trained mostly on them.
General rules for every request are in the request format guide.
Example
The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).
from jeff import Client
from jeff.client import choice_question
jeff = Client("http://localhost:8765", model="emotion")
state = {
"source": "Reddit comment ([NAME] and [RELIGION] replace removed names)",
"comment": "Thanks so much for the tip, I had no idea the library lent out tools. Saved me a fortune.",
}
answers = jeff.ask(state, {
"emotion": choice_question(
{
"gratitude": "Gratitude: feeling thankful or appreciative.",
"joy": "Joy: a feeling of pleasure and happiness.",
"surprise": "Surprise: being astonished or startled by something unexpected.",
"realization": "Realization: becoming aware of something.",
"admiration": "Admiration: finding something impressive or worthy of respect.",
"relief": "Relief: reassurance and relaxation after anxiety or distress ends.",
"annoyance": "Annoyance: mild anger or irritation.",
"confusion": "Confusion: lack of understanding or uncertainty.",
"sadness": "Sadness: feeling sorrow or unhappiness.",
"neutral": "Neutral: no particular emotion is expressed.",
},
"Which emotion does the comment express most? Choose the emotion that is strongest in the comment, or neutral if it expresses no particular emotion.",
),
})
print("emotion", answers.choice("emotion").key)import { Client, choiceQuestion } from '@jeff/client';
const jeff = new Client({ url: 'http://localhost:8765', model: 'emotion' });
const state = {
source: 'Reddit comment ([NAME] and [RELIGION] replace removed names)',
comment: 'Thanks so much for the tip, I had no idea the library lent out tools. Saved me a fortune.',
};
const answers = await jeff.ask(state, {
emotion: choiceQuestion(
{
gratitude: 'Gratitude: feeling thankful or appreciative.',
joy: 'Joy: a feeling of pleasure and happiness.',
surprise: 'Surprise: being astonished or startled by something unexpected.',
realization: 'Realization: becoming aware of something.',
admiration: 'Admiration: finding something impressive or worthy of respect.',
relief: 'Relief: reassurance and relaxation after anxiety or distress ends.',
annoyance: 'Annoyance: mild anger or irritation.',
confusion: 'Confusion: lack of understanding or uncertainty.',
sadness: 'Sadness: feeling sorrow or unhappiness.',
neutral: 'Neutral: no particular emotion is expressed.',
},
'Which emotion does the comment express most? Choose the emotion that is strongest in the comment, or neutral if it expresses no particular emotion.',
),
});
console.log('emotion', answers.emotion.key);curl -s http://localhost:8765/v1/systemone \
-H 'content-type: application/json' \
-d '{
"model": "emotion",
"state": {
"source": "Reddit comment ([NAME] and [RELIGION] replace removed names)",
"comment": "Thanks so much for the tip, I had no idea the library lent out tools. Saved me a fortune."
},
"questions": {
"emotion": {
"type": "choice",
"instructions": "Which emotion does the comment express most? Choose the emotion that is strongest in the comment, or neutral if it expresses no particular emotion.",
"criteria": {
"gratitude": "Gratitude: feeling thankful or appreciative.",
"joy": "Joy: a feeling of pleasure and happiness.",
"surprise": "Surprise: being astonished or startled by something unexpected.",
"realization": "Realization: becoming aware of something.",
"admiration": "Admiration: finding something impressive or worthy of respect.",
"relief": "Relief: reassurance and relaxation after anxiety or distress ends.",
"annoyance": "Annoyance: mild anger or irritation.",
"confusion": "Confusion: lack of understanding or uncertainty.",
"sadness": "Sadness: feeling sorrow or unhappiness.",
"neutral": "Neutral: no particular emotion is expressed."
}
}
}
}'Response
{
"model": "emotion",
"answers": {
"emotion": {
"type": "choice",
"probabilities": {
"gratitude": 0.9377514192297772,
"joy": 0.0048510179566744254,
"surprise": 0.0013096331149660347,
"realization": 0.00906897121104442,
"admiration": 0.005445419730876873,
"relief": 0.00815921427750748,
"annoyance": 0.0002489432112212598,
"confusion": 0.03209981177197163,
"sadness": 0.00015036246233088356,
"neutral": 0.0009152070336297481
},
"choice": "gratitude",
"confidence": 0.9308349102553081
}
},
"usage": {
"input_tokens": 285,
"output_tokens": 0,
"orders": 1
}
}Results
On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters
| Test set | Test rows | Qwen3.5-0.8B untrained | Jeff v1.2 0.8B alone | Jeff v1.2 0.8B + adapter |
|---|---|---|---|---|
emotion | 5,408 | 12.5% · 0.044 | 32.2% · 0.038 | 60.6% · 0.020 |
emotion5,408 test rows- Qwen3.5-0.8B untrained
- 12.5% · 0.044
- Jeff v1.2 0.8B alone
- 32.2% · 0.038
- Jeff v1.2 0.8B + adapter
- 60.6% · 0.020
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
How sure is it, and is it right?
Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.
When this adapter says it is about 50% sure, it is right about 52% of the time (676 test rows).
Point at a dot for its numbers. Bigger dots hold more rows.
| Model | Calibration error (ECE) | Brier score | Log loss |
|---|---|---|---|
| Qwen3.5-0.8B untrained | 0.044 | 0.953 | 3.232 |
| Jeff v1.2 0.8B alone | 0.038 | 0.817 | 2.241 |
| Jeff v1.2 0.8B + adapter | 0.020 | 0.539 | 1.236 |
All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.
Where it gets things wrong
approvalread asneutral: 127 rowsannoyanceread asneutral: 122 rowsdisapprovalread asneutral: 105 rowsneutralread ascuriosity: 80 rowscuriosityread asneutral: 63 rows
The five commonest mistakes with the adapter, out of 5,408 test rows.
| Right answer | Test rows | Accuracy with the adapter |
|---|---|---|
neutral | 1,682 | 78.5% |
admiration | 422 | 73.5% |
gratitude | 294 | 90.1% |
approval | 281 | 26.7% |
annoyance | 261 | 19.5% |
disapproval | 232 | 33.2% |
amusement | 227 | 84.6% |
curiosity | 225 | 54.7% |
love | 198 | 82.8% |
anger | 166 | 48.2% |
Show all 28 answers
| Right answer | Test rows | Accuracy with the adapter |
|---|---|---|
neutral | 1,682 | 78.5% |
admiration | 422 | 73.5% |
gratitude | 294 | 90.1% |
approval | 281 | 26.7% |
annoyance | 261 | 19.5% |
disapproval | 232 | 33.2% |
amusement | 227 | 84.6% |
curiosity | 225 | 54.7% |
love | 198 | 82.8% |
anger | 166 | 48.2% |
optimism | 146 | 53.4% |
sadness | 129 | 53.5% |
confusion | 128 | 34.4% |
joy | 128 | 57.0% |
realization | 124 | 17.7% |
caring | 110 | 28.2% |
disappointment | 109 | 24.8% |
surprise | 108 | 54.6% |
disgust | 96 | 45.8% |
excitement | 78 | 29.5% |
fear | 72 | 73.6% |
desire | 69 | 44.9% |
remorse | 51 | 72.5% |
embarrassment | 30 | 46.7% |
nervousness | 17 | 52.9% |
pride | 13 | 15.4% |
relief | 8 | 25.0% |
grief | 4 | 50.0% |
Against Qwen3.8-27B
More accurate, 42× faster than Qwen3.8-27B alone.
Gain over Qwen3.8-27B alone: +25.0 points [+20.0, +30.0] (95% interval).
On 500 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 35.6% of the time at 4.79 s per query on average; with Jeff and this adapter answering first, it was right 60.6% at 113 ms.
This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.
| Route | Mean | Median | 95th percentile | Typical prompt |
|---|---|---|---|---|
| Qwen3.8-27B alone | 4.79 s | 4.83 s | 5.23 s | 617 tokens |
| Jeff + adapter, its own answer | 113 ms | 113 ms | 117 ms | 584 tokens |
Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.
This task's threshold 0.00: accuracy 60.6%, 42.5× faster, 0.0% sent on to Qwen3.8-27B.
Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 500 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.
How it was trained
- Training rows
- 50,674
- Steps
- 792
- Training time
- 81 min
- Size as saved
- 41.5 MB
One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-emotion-20260930-0623.
Source: jeff-finetunes/adapters/BASELINE.md
Data card
Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean
- Test set, so anyone can check the numbers
- Calibration rows, the rows its threshold is chosen on
- QA report, sanitised: the data-quality checks run before training
- How the test set was held out
- The official GoEmotions test split (simplified configuration); never trained on.
- Training data
- Built from public data sets, listed under Data and licence.
- The source data sets are public (listed under Data and licence). A script to rebuild our rows from them will follow.
Data and licence
The adapter is released under Apache-2.0. It was trained on:
- GoEmotions, simplified configurationLicence: Apache-2.0
English Reddit comments with names removed. Option descriptions were written by hand from the label names.
Changelog
- 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.
Comments
Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.
