QA report: emotion
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.
QA: emotion
Checked 2026-09-30 11:12 by adapters/qa/qa.py (READY file ready-emotion.json, public v1 (READY 2026-09-29T23:56)).
Verdict: PASS WITH NOTES (for v2)
Findings
- No shortcut. A surface-only model sits at the majority baseline (31.1% accuracy, 5% balanced against 3.6% chance). The strong phrases ("sorry" for remorse, "thanks" for gratitude, "lol" for amusement) are the meaning of the labels. The GoEmotions placeholders "[NAME]" and "[RELIGION]" appear in 13.9% of rows, spread evenly across classes (9–17%).
- Label noise is built into the source. GoEmotions is multi-label, and v1 keeps one label per comment. Near-identical comments carry different labels: "I'm sorry for your loss" is labelled caring, grief, remorse and amusement. The BoW model's balanced accuracy is only 32%, which fits a noisy task.
v2: turn the raters' labels into soft targets (the counts are in
source.gold_labelsandupstream_ids), and/or merge rare, confusable classes. - The class balance is heavily skewed: neutral is 31% of test. Duplicates are negligible (0.1%), and held-out near duplicates are 0.1–0.2%.
Automatic flags (for the reviewer to judge; not all are problems)
- formatting 'ends with ?' differs by class: curiosity 51%, confusion 32%, neutral 5%, surprise 5% …
- formatting 'ends with .' differs by class: disapproval 62%, realization 58%, approval 57%, disappointment 56% …
- formatting 'ends with !' differs by class: excitement 29%, gratitude 23%, joy 15%, admiration 15% …
- formatting 'no end punctuation' differs by class: amusement 54%, grief 43%, sadness 41%, love 40% …
- 157 strong phrase flags (see list): review whether they are meaning or leakage
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 46,067 | 46067 | train.jsonl |
| dev | 1,000 | 1000 | dev.jsonl |
| calibration | 1,000 | 1000 | calibration.jsonl |
| test | 5,408 | 5408 | test.jsonl |
- Train sha256:
30867b255a4e4ca1eb9ec1f3116245aebacd6cf702a0a9adc0e71c87cd1c4bcd(READY file gives no checksum) - Main text field (the text the phrase and length checks use):
state.comment. - Label classes: admiration, amusement, anger, annoyance, approval, caring, confusion, curiosity, desire, disappointment, disapproval, disgust, embarrassment, excitement, fear, gratitude, grief, joy, love, nervousness, neutral, optimism, pride, realization, relief, remorse, sadness, surprise (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): emotion.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| admiration | 7.8% | 7.4% | 7.7% | 7.8% | 3,590 |
| amusement | 4.6% | 5.7% | 5.4% | 4.2% | 2,102 |
| anger | 2.9% | 3.0% | 2.9% | 3.1% | 1,357 |
| annoyance | 4.5% | 3.8% | 3.7% | 4.8% | 2,082 |
| approval | 5.6% | 5.7% | 7.6% | 5.2% | 2,566 |
| caring | 2.0% | 2.5% | 2.0% | 2.0% | 910 |
| confusion | 2.6% | 2.5% | 1.7% | 2.4% | 1,180 |
| curiosity | 4.1% | 2.8% | 5.4% | 4.2% | 1,872 |
| desire | 1.2% | 0.5% | 1.7% | 1.3% | 542 |
| disappointment | 2.3% | 2.3% | 1.8% | 2.0% | 1,061 |
| disapproval | 3.9% | 4.6% | 5.3% | 4.3% | 1,818 |
| disgust | 1.5% | 1.3% | 1.2% | 1.8% | 675 |
| embarrassment | 0.6% | 0.6% | 0.4% | 0.6% | 260 |
| excitement | 1.5% | 0.9% | 1.6% | 1.4% | 693 |
| fear | 1.2% | 1.0% | 1.1% | 1.3% | 570 |
| gratitude | 4.9% | 6.0% | 5.3% | 5.4% | 2,279 |
| grief | 0.1% | 0.2% | 0.0% | 0.1% | 67 |
| joy | 2.6% | 2.2% | 2.7% | 2.4% | 1,188 |
| love | 3.9% | 3.5% | 3.4% | 3.7% | 1,808 |
| nervousness | 0.3% | 0.2% | 0.4% | 0.3% | 130 |
| neutral | 31.3% | 33.0% | 28.1% | 31.1% | 14,417 |
| optimism | 2.8% | 2.8% | 3.5% | 2.7% | 1,293 |
| pride | 0.2% | 0.4% | 0.3% | 0.2% | 81 |
| realization | 1.9% | 2.8% | 2.0% | 2.3% | 865 |
| relief | 0.3% | 0.0% | 0.3% | 0.1% | 126 |
| remorse | 1.0% | 0.7% | 0.8% | 0.9% | 474 |
| sadness | 2.5% | 1.3% | 2.0% | 2.4% | 1,129 |
| surprise | 2.0% | 2.3% | 1.7% | 2.0% | 932 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| emotion | 100.0% | 100.0% | 100.0% | 100.0% | 46,067 |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Options per choice row
| split | min | median | p99 | max |
|---|---|---|---|---|
| train | 28 | 28 | 28 | 28 |
| dev | 28 | 28 | 28 | 28 |
| calibration | 28 | 28 | 28 | 28 |
| test | 28 | 28 | 28 | 28 |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | estimate: characters / 3 (upper bound for English) | 669 | 701 | 870 | 0 |
| dev | estimate: characters / 3 (upper bound for English) | 669 | 705 | 710 | 0 |
| calibration | estimate: characters / 3 (upper bound for English) | 669 | 701 | 708 | 0 |
| test | estimate: characters / 3 (upper bound for English) | 668 | 700 | 713 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| comment, source | 46,067 (100.0%) |
Instructions (train)
- Canonical (the most common text) 70.1%, reworded 26.9% (20 distinct rewordings), none 3.0%. Target about 70 / 27 / 3.
- Canonical text: "Which emotion does the comment express most? Choose the emotion that is strongest in the comment, or neutral if it expresses no particular emotion."
sourceinstruction tag: canonical 70.1%, none 3.0%, variant-18 1.4%, variant-14 1.4%, variant-12 1.4%, variant-19 1.4%
| class | canonical | none |
|---|---|---|
| admiration | 70.1% | 2.9% |
| amusement | 69.0% | 2.8% |
| anger | 68.3% | 3.8% |
| annoyance | 69.5% | 3.1% |
| approval | 70.6% | 3.5% |
| caring | 72.0% | 1.9% |
| confusion | 71.9% | 2.6% |
| curiosity | 68.9% | 2.6% |
| desire | 71.0% | 2.6% |
| disappointment | 69.5% | 3.6% |
| disapproval | 71.6% | 2.6% |
| disgust | 72.1% | 3.0% |
| embarrassment | 71.9% | 2.7% |
| excitement | 69.0% | 4.2% |
| fear | 69.5% | 3.9% |
| gratitude | 69.1% | 2.6% |
| grief | 68.7% | 4.5% |
| joy | 70.2% | 2.9% |
| love | 70.2% | 3.5% |
| nervousness | 66.2% | 4.6% |
| neutral | 70.0% | 3.0% |
| optimism | 71.0% | 3.4% |
| pride | 72.8% | 3.7% |
| realization | 70.1% | 3.8% |
| relief | 67.5% | 4.8% |
| remorse | 66.9% | 3.6% |
| sadness | 72.7% | 2.5% |
| surprise | 72.2% | 2.7% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 46,067 train rows; models are scored on the full test file (5,408 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | admiration | 3590 | 20 | 58 | 117 | 63 |
| train | amusement | 2102 | 22 | 63 | 113 | 65 |
| train | anger | 1357 | 18 | 58 | 121 | 64 |
| train | annoyance | 2082 | 28 | 73 | 126 | 75 |
| train | approval | 2566 | 26 | 71 | 121 | 72 |
| train | caring | 910 | 25 | 72 | 122 | 73 |
| train | confusion | 1180 | 31 | 76 | 121 | 75 |
| train | curiosity | 1872 | 27 | 66 | 120 | 70 |
| train | desire | 542 | 31 | 69 | 118 | 70 |
| train | disappointment | 1061 | 33 | 74 | 124 | 76 |
| train | disapproval | 1818 | 30 | 73 | 124 | 75 |
| train | disgust | 675 | 25 | 67 | 118 | 69 |
| train | embarrassment | 260 | 33 | 74 | 122 | 76 |
| train | excitement | 693 | 20 | 55 | 111 | 60 |
| train | fear | 570 | 24 | 69 | 121 | 72 |
| train | gratitude | 2279 | 22 | 59 | 114 | 64 |
| train | grief | 67 | 25 | 57 | 114 | 64 |
| train | joy | 1188 | 23 | 64 | 116 | 67 |
| train | love | 1808 | 20 | 60 | 115 | 64 |
| train | nervousness | 130 | 25 | 77 | 124 | 75 |
| train | neutral | 14417 | 20 | 64 | 121 | 67 |
| train | optimism | 1293 | 36 | 80 | 124 | 80 |
| train | pride | 81 | 26 | 54 | 111 | 61 |
| train | realization | 865 | 36 | 77 | 126 | 79 |
| train | relief | 126 | 27 | 64 | 109 | 66 |
| train | remorse | 474 | 29 | 69 | 122 | 72 |
| train | sadness | 1129 | 22 | 69 | 119 | 69 |
| train | surprise | 932 | 25 | 65 | 118 | 69 |
| test | admiration | 422 | 19 | 60 | 110 | 62 |
| test | amusement | 227 | 24 | 61 | 111 | 64 |
| test | anger | 166 | 18 | 63 | 124 | 67 |
| test | annoyance | 261 | 26 | 73 | 126 | 75 |
| test | approval | 281 | 25 | 70 | 123 | 72 |
| test | caring | 110 | 22 | 68 | 125 | 71 |
| test | confusion | 128 | 32 | 76 | 118 | 75 |
| test | curiosity | 225 | 22 | 60 | 112 | 63 |
| test | desire | 69 | 38 | 78 | 114 | 73 |
| test | disappointment | 109 | 32 | 71 | 127 | 76 |
| test | disapproval | 232 | 28 | 72 | 122 | 74 |
| test | disgust | 96 | 26 | 67 | 123 | 72 |
| test | embarrassment | 30 | 31 | 88 | 128 | 79 |
| test | excitement | 78 | 20 | 50 | 117 | 61 |
| test | fear | 72 | 25 | 63 | 118 | 68 |
| test | gratitude | 294 | 19 | 58 | 112 | 61 |
| test | grief | 4 | 36 | 75 | 82 | 64 |
| test | joy | 128 | 23 | 63 | 115 | 67 |
| test | love | 198 | 20 | 60 | 112 | 62 |
| test | nervousness | 17 | 35 | 71 | 126 | 76 |
| test | neutral | 1682 | 19 | 64 | 119 | 66 |
| test | optimism | 146 | 31 | 75 | 121 | 74 |
| test | pride | 13 | 35 | 41 | 93 | 55 |
| test | realization | 124 | 36 | 83 | 121 | 79 |
| test | relief | 8 | 32 | 59 | 92 | 61 |
| test | remorse | 51 | 27 | 56 | 104 | 61 |
| test | sadness | 129 | 18 | 66 | 124 | 68 |
| test | surprise | 108 | 22 | 61 | 113 | 64 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| emotion | 46067 | 66 | 69 | 61 | 28 |
Correct option: longest / shortest / position / key
For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):
| split | rows | correct is longest | correct is shortest | chance (1/listed) | mean relative position (0 first, 1 last; 0.5 expected) | position fifths |
|---|
Correct key and position by option count
| split | options | rows | mean options | top correct keys | most common position (0-based) |
|---|---|---|---|---|---|
| train | 11-30 | 46067 | 28.0 | neutral 31.3%, admiration 7.8%, approval 5.6%, gratitude 4.9%, amusement 4.6% | 2 (4.2%) |
| test | 11-30 | 5408 | 28.0 | neutral 31.1%, admiration 7.8%, gratitude 5.4%, approval 5.2%, annoyance 4.8% | 1 (4.5%) |
Option count by label class (train)
| class | rows | min | median | mean | max |
|---|---|---|---|---|---|
| admiration | 3590 | 28 | 28 | 28.0 | 28 |
| amusement | 2102 | 28 | 28 | 28.0 | 28 |
| anger | 1357 | 28 | 28 | 28.0 | 28 |
| annoyance | 2082 | 28 | 28 | 28.0 | 28 |
| approval | 2566 | 28 | 28 | 28.0 | 28 |
| caring | 910 | 28 | 28 | 28.0 | 28 |
| confusion | 1180 | 28 | 28 | 28.0 | 28 |
| curiosity | 1872 | 28 | 28 | 28.0 | 28 |
| desire | 542 | 28 | 28 | 28.0 | 28 |
| disappointment | 1061 | 28 | 28 | 28.0 | 28 |
| disapproval | 1818 | 28 | 28 | 28.0 | 28 |
| disgust | 675 | 28 | 28 | 28.0 | 28 |
| embarrassment | 260 | 28 | 28 | 28.0 | 28 |
| excitement | 693 | 28 | 28 | 28.0 | 28 |
| fear | 570 | 28 | 28 | 28.0 | 28 |
| gratitude | 2279 | 28 | 28 | 28.0 | 28 |
| grief | 67 | 28 | 28 | 28.0 | 28 |
| joy | 1188 | 28 | 28 | 28.0 | 28 |
| love | 1808 | 28 | 28 | 28.0 | 28 |
| nervousness | 130 | 28 | 28 | 28.0 | 28 |
| neutral | 14417 | 28 | 28 | 28.0 | 28 |
| optimism | 1293 | 28 | 28 | 28.0 | 28 |
| pride | 81 | 28 | 28 | 28.0 | 28 |
| realization | 865 | 28 | 28 | 28.0 | 28 |
| relief | 126 | 28 | 28 | 28.0 | 28 |
| remorse | 474 | 28 | 28 | 28.0 | 28 |
| sadness | 1129 | 28 | 28 | 28.0 | 28 |
| surprise | 932 | 28 | 28 | 28.0 | 28 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 31.3%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| target_kind | 2 | 31.3% | hard: neutral 35.6%, admiration 7.4%; soft: neutral 9.9%, admiration 9.6% |
| instructions | 22 | 31.3% | canonical: neutral 31.2%, admiration 7.8%; none: neutral 30.9%, admiration 7.4%; variant-18: neutral 31.7%, admiration 7.5%; variant-14: neutral 31.6%, admiration 9.5%; variant-12: neutral 32.5%, admiration 7.8%; variant-19: neutral 29.4%, admiration 8.3% |
| split | 2 | 31.3% | train: neutral 31.3%, admiration 7.8%; validation: neutral 31.6%, admiration 7.5% |
| copies | 9 | 31.3% | 1: neutral 31.3%, admiration 7.8%; 2: neutral 28.8%, gratitude 15.8%; 3: neutral 27.5%, gratitude 15.0%; 4: gratitude 27.3%, neutral 27.3%; 5: neutral 44.4%, surprise 11.1%; 6: neutral 50.0%, joy 25.0% |
Formatting by label class (main text, share of rows)
| feature | lowest classes | highest classes | |
|---|---|---|---|
| ends with ? | pride 0%, relief 1%, joy 1% | curiosity 51%, confusion 32%, neutral 5% | gap |
| ends with . | curiosity 25%, amusement 34%, excitement 35% | disapproval 62%, realization 58%, approval 57% | gap |
| ends with ! | grief 1%, remorse 3%, confusion 3% | excitement 29%, gratitude 23%, joy 15% | gap |
| no end punctuation | curiosity 18%, confusion 25%, caring 27% | amusement 54%, grief 43%, sadness 41% | gap |
| starts lowercase | relief 4%, desire 4%, realization 4% | grief 9%, amusement 9%, nervousness 8% | |
| all lowercase | pride 1%, relief 2%, desire 2% | embarrassment 7%, nervousness 6%, grief 6% | |
| has a digit | anger 4%, caring 5%, gratitude 5% | disappointment 12%, surprise 12%, realization 12% | |
| has newline | admiration 0%, amusement 0%, anger 0% | surprise 0%, sadness 0%, remorse 0% | |
| has quotes | relief 1%, grief 1%, optimism 2% | annoyance 6%, amusement 5%, embarrassment 5% | |
| has markup (HTML/markdown) | caring 0%, desire 0%, disgust 0% | anger 1%, relief 1%, joy 1% | |
| has URL | admiration 0%, amusement 0%, anger 0% | surprise 0%, sadness 0%, remorse 0% | |
| non-ASCII | neutral 12%, anger 12%, desire 13% | relief 21%, disapproval 21%, remorse 20% | |
| non-Latin script | admiration 0%, anger 0%, annoyance 0% | embarrassment 0%, disapproval 0%, surprise 0% | |
| emoji | pride 0%, relief 0%, anger 1% | grief 4%, love 4%, sadness 4% | |
| ALL-CAPS word (4+) | embarrassment 8%, nervousness 9%, gratitude 10% | love 23%, anger 23%, neutral 21% | |
| contains ' - ' or — | grief 0%, pride 0%, sadness 0% | nervousness 2%, gratitude 1%, embarrassment 1% |
Same, by row kind
| feature | emotion |
|---|---|
| ends with ? | 6% |
| ends with . | 48% |
| ends with ! | 9% |
| no end punctuation | 34% |
| starts lowercase | 6% |
| all lowercase | 4% |
| has a digit | 9% |
| has newline | 0% |
| has quotes | 4% |
| has markup (HTML/markdown) | 0% |
| has URL | 0% |
| non-ASCII | 15% |
| non-Latin script | 0% |
| emoji | 2% |
| ALL-CAPS word (4+) | 17% |
| contains ' - ' or — | 1% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, admiration: great z=42 14.1% vs 0.8%; good z=32 14.1% vs 2.9%; awesome z=29 6.7% vs 0.2%; amazing z=28 6.0% vs 0.3%; nice z=24 5.4% vs 0.6%; beautiful z=22 4.1% vs 0.1%; best z=20 5.3% vs 1.0%; pretty z=20 5.4% vs 1.1%; cute z=18 2.6% vs 0.2%; appreciate z=16 2.2% vs 0.2%; looks z=14 3.7% vs 1.0%; cool z=13 2.6% vs 0.6%; fantastic z=12 1.2% vs 0.1%; job z=12 2.0% vs 0.4%; wow z=11 2.9% vs 1.0%; is z=11 23.0% vs 17.3%; wonderful z=11 1.0% vs 0.1%; incredible z=10 0.8% vs 0.0%; excellent z=10 0.8% vs 0.0%; interesting z=10 1.8% vs 0.5%
Words, amusement: lol z=64 42.6% vs 0.9%; haha z=33 11.0% vs 0.3%; funny z=30 9.1% vs 0.3%; lmao z=23 5.6% vs 0.3%; fun z=22 6.2% vs 0.6%; hilarious z=18 3.2% vs 0.1%; joke z=17 3.6% vs 0.3%; laugh z=17 2.9% vs 0.2%; hahaha z=16 2.8% vs 0.1%; laughed z=14 2.1% vs 0.0%; laughing z=12 1.6% vs 0.1%; ha z=8 1.0% vs 0.1%; loud z=7 0.8% vs 0.1%; kidding z=7 0.6% vs 0.0%; jokes z=7 0.7% vs 0.1%; joking z=7 0.5% vs 0.0%; hahahaha z=7 0.5% vs 0.0%; funnier z=7 0.5% vs 0.0%; entertaining z=6 0.5% vs 0.0%; actually z=6 3.2% vs 1.5%
Words, anger: fuck z=36 15.0% vs 0.5%; hate z=27 10.2% vs 0.6%; fucking z=26 8.8% vs 0.5%; angry z=15 2.8% vs 0.1%; stupid z=14 3.8% vs 0.5%; dare z=14 2.2% vs 0.1%; shut z=12 2.1% vs 0.1%; hell z=12 3.2% vs 0.5%; shit z=9 3.2% vs 0.7%; asshole z=9 1.0% vs 0.0%; fucked z=9 1.1% vs 0.1%; wtf z=9 1.1% vs 0.1%; bitch z=9 1.1% vs 0.1%; idiot z=9 1.4% vs 0.2%; bastard z=9 0.9% vs 0.0%; bullshit z=8 1.0% vs 0.1%; stop z=8 2.7% vs 0.7%; kill z=8 1.5% vs 0.3%; suck z=7 1.0% vs 0.1%; off z=7 3.2% vs 1.3%
Words, annoyance: stupid z=16 3.8% vs 0.4%; annoying z=13 1.9% vs 0.0%; fucking z=13 3.7% vs 0.6%; shit z=12 3.6% vs 0.7%; damn z=11 3.7% vs 0.9%; dumb z=10 1.7% vs 0.2%; idiot z=10 1.4% vs 0.1%; fuck z=9 3.1% vs 0.8%; sucks z=9 1.7% vs 0.3%; annoyed z=8 0.8% vs 0.0%; idiots z=8 0.9% vs 0.1%; weird z=8 2.2% vs 0.5%; stop z=8 2.5% vs 0.7%; pissed z=8 0.9% vs 0.1%; hate z=7 2.6% vs 0.8%; hell z=7 1.8% vs 0.5%; ridiculous z=7 0.9% vs 0.1%; asshole z=7 0.6% vs 0.0%; fool z=6 0.5% vs 0.0%; frustrating z=6 0.5% vs 0.0%
Words, approval: agree z=26 6.4% vs 0.2%; yes z=16 5.1% vs 0.9%; yeah z=13 5.7% vs 1.7%; right z=12 5.9% vs 1.9%; agreed z=12 1.4% vs 0.1%; true z=11 2.7% vs 0.6%; correct z=9 1.2% vs 0.1%; ok z=9 2.4% vs 0.6%; sure z=9 3.8% vs 1.4%; exactly z=8 1.8% vs 0.4%; yep z=7 0.9% vs 0.2%; definitely z=7 1.8% vs 0.6%; fine z=7 1.1% vs 0.3%; fair z=6 0.9% vs 0.2%; but z=6 12.0% vs 8.2%; free z=5 1.1% vs 0.4%; prefer z=5 0.5% vs 0.1%; it's z=5 6.5% vs 4.2%; with z=5 9.9% vs 6.9%; absolutely z=5 1.1% vs 0.4%
Words, caring: worry z=17 4.8% vs 0.2%; you z=16 47.3% vs 19.1%; yourself z=16 6.2% vs 0.5%; your z=14 19.1% vs 5.5%; stay z=14 4.5% vs 0.3%; safe z=13 3.2% vs 0.2%; luck z=13 4.8% vs 0.5%; help z=13 5.8% vs 0.8%; careful z=11 2.1% vs 0.0%; concerned z=11 1.8% vs 0.0%; care z=10 3.7% vs 0.5%; bless z=10 1.8% vs 0.1%; better z=9 5.9% vs 1.5%; get z=9 11.0% vs 4.0%; take z=9 4.8% vs 1.1%; need z=8 5.2% vs 1.3%; keep z=8 4.0% vs 0.9%; strong z=8 1.9% vs 0.2%; please z=7 3.2% vs 0.7%; praying z=7 0.9% vs 0.0%
Words, confusion: sure z=19 10.2% vs 1.3%; confused z=19 5.9% vs 0.0%; why z=17 11.5% vs 2.0%; or z=15 12.5% vs 3.0%; what z=14 17.6% vs 5.6%; understand z=14 4.6% vs 0.5%; don't z=12 10.9% vs 3.1%; know z=12 10.6% vs 3.0%; idea z=10 3.5% vs 0.5%; idk z=10 2.1% vs 0.2%; not z=10 17.7% vs 7.9%; how z=8 9.4% vs 3.7%; confusing z=8 0.9% vs 0.0%; maybe z=8 3.9% vs 1.0%; confusion z=7 0.8% vs 0.0%; don z=7 4.4% vs 1.4%; doubt z=7 1.2% vs 0.1%; clue z=6 0.7% vs 0.0%; do z=6 8.6% vs 4.2%; i z=6 45.8% vs 31.7%
Words, curiosity: what z=25 20.9% vs 5.2%; curious z=21 6.5% vs 0.1%; why z=17 9.0% vs 2.0%; how z=17 12.1% vs 3.4%; did z=16 7.8% vs 1.8%; do z=14 11.4% vs 4.0%; you z=13 33.0% vs 19.1%; does z=12 4.3% vs 1.0%; what's z=11 2.0% vs 0.2%; where z=11 4.0% vs 1.0%; are z=10 13.2% vs 6.8%; curiosity z=9 1.0% vs 0.0%; interesting z=8 2.2% vs 0.5%; anyone z=7 2.4% vs 0.7%; wonder z=7 1.4% vs 0.3%; explain z=7 0.9% vs 0.1%; wondering z=6 0.8% vs 0.1%; question z=6 1.2% vs 0.3%; any z=6 3.2% vs 1.4%; or z=6 5.7% vs 3.1%
Words, desire: wish z=44 39.7% vs 0.5%; want z=17 14.8% vs 1.7%; i z=15 74.5% vs 31.5%; could z=14 11.6% vs 1.5%; wanted z=12 4.8% vs 0.3%; need z=9 7.2% vs 1.3%; wanna z=7 2.2% vs 0.2%; i'd z=6 3.7% vs 0.7%; would z=6 10.3% vs 3.8%; hope z=6 5.9% vs 1.8%; more z=6 8.1% vs 3.0%; dream z=5 1.3% vs 0.1%; hadn't z=5 0.7% vs 0.0%; pray z=5 0.9% vs 0.1%; had z=5 5.9% vs 2.2%; to z=5 37.8% vs 24.7%; desire z=5 0.6% vs 0.0%; please z=4 2.6% vs 0.7%; expecting z=4 0.9% vs 0.1%; blunder z=4 0.4% vs 0.0%
Words, disappointment: disappointed z=15 3.3% vs 0.1%; bad z=14 9.0% vs 1.6%; disappointing z=9 2.4% vs 0.0%; upset z=9 1.6% vs 0.1%; depressing z=9 1.1% vs 0.0%; lost z=9 2.5% vs 0.3%; miss z=9 2.3% vs 0.3%; unfortunately z=8 1.8% vs 0.2%; disappointment z=8 1.2% vs 0.0%; missed z=8 1.7% vs 0.2%; tried z=7 1.8% vs 0.3%; hurt z=7 1.7% vs 0.2%; game z=7 4.2% vs 1.3%; upsetting z=6 0.7% vs 0.0%; lose z=6 1.5% vs 0.3%; poor z=6 1.9% vs 0.4%; disaster z=6 0.6% vs 0.0%; sucks z=5 1.4% vs 0.3%; worse z=5 1.6% vs 0.4%; boring z=5 0.8% vs 0.1%
Words, disapproval: not z=22 24.8% vs 7.5%; no z=18 14.4% vs 3.9%; don't z=16 11.1% vs 2.9%; t z=15 12.9% vs 4.0%; don z=13 5.9% vs 1.3%; disagree z=12 1.6% vs 0.1%; nope z=10 1.3% vs 0.1%; doesn't z=9 3.3% vs 0.9%; wrong z=9 3.1% vs 0.8%; can't z=9 3.9% vs 1.2%; nah z=8 1.5% vs 0.2%; isn't z=8 2.6% vs 0.7%; unpopular z=8 0.7% vs 0.0%; think z=8 7.0% vs 3.1%; dont z=8 2.1% vs 0.5%; doesn z=7 1.5% vs 0.3%; bad z=6 4.0% vs 1.7%; refuse z=6 0.5% vs 0.0%; illegal z=6 0.7% vs 0.1%; opinion z=6 1.1% vs 0.2%
Words, disgust: awful z=23 9.6% vs 0.2%; disgusting z=22 13.5% vs 0.0%; worst z=22 9.2% vs 0.3%; weird z=17 7.6% vs 0.5%; worse z=15 5.8% vs 0.3%; nasty z=11 2.2% vs 0.0%; ugly z=10 2.5% vs 0.1%; gross z=10 2.1% vs 0.1%; creepy z=10 2.2% vs 0.1%; terrible z=8 2.7% vs 0.4%; ugh z=7 1.6% vs 0.1%; is z=7 28.4% vs 17.6%; horrible z=7 1.9% vs 0.3%; hideous z=6 0.7% vs 0.0%; dirty z=6 0.9% vs 0.1%; disgusted z=6 1.2% vs 0.0%; disgust z=6 0.6% vs 0.0%; ever z=6 3.4% vs 1.0%; filthy z=6 0.6% vs 0.0%; bad z=5 4.7% vs 1.7%
Words, embarrassment: shame z=18 11.5% vs 0.2%; awkward z=18 10.4% vs 0.1%; embarrassing z=16 11.5% vs 0.0%; ashamed z=12 5.0% vs 0.0%; embarrassed z=12 4.6% vs 0.0%; embarrassment z=10 5.0% vs 0.0%; oops z=8 1.9% vs 0.0%; weird z=7 5.4% vs 0.6%; uncomfortable z=7 2.3% vs 0.1%; cringy z=7 1.5% vs 0.0%; shamed z=5 0.8% vs 0.0%; temporarily z=5 0.8% vs 0.0%; millionaires z=5 0.8% vs 0.0%; forgot z=5 2.3% vs 0.3%; translate z=4 0.8% vs 0.0%; shaming z=4 0.8% vs 0.0%; cringey z=4 0.8% vs 0.0%; feel z=4 5.8% vs 1.6%; socially z=4 0.8% vs 0.0%; abused z=4 0.8% vs 0.0%
Words, excitement: excited z=25 11.1% vs 0.1%; wait z=16 6.6% vs 0.6%; wow z=15 7.8% vs 1.0%; interesting z=14 5.8% vs 0.5%; happy z=13 7.4% vs 1.1%; cake z=11 2.6% vs 0.1%; birthday z=11 2.3% vs 0.1%; exciting z=10 1.6% vs 0.0%; yay z=9 1.4% vs 0.0%; new z=9 4.8% vs 1.0%; omg z=7 2.5% vs 0.4%; interested z=7 1.6% vs 0.2%; excitement z=7 0.9% vs 0.0%; day z=7 4.0% vs 1.2%; year z=6 3.8% vs 1.1%; crazy z=6 2.0% vs 0.4%; amazing z=6 2.7% vs 0.7%; stoked z=6 0.6% vs 0.0%; cheers z=6 1.3% vs 0.2%; cakeday z=6 0.7% vs 0.0%
Words, fear: afraid z=23 10.7% vs 0.1%; scared z=22 11.1% vs 0.1%; terrible z=21 10.0% vs 0.3%; horrible z=19 7.5% vs 0.2%; scary z=18 6.8% vs 0.0%; terrifying z=15 5.3% vs 0.0%; fear z=15 4.4% vs 0.1%; dangerous z=11 2.6% vs 0.1%; worried z=11 3.0% vs 0.1%; creepy z=10 2.8% vs 0.1%; cringe z=10 2.8% vs 0.1%; horrifying z=10 1.9% vs 0.0%; horror z=9 1.8% vs 0.0%; nightmare z=8 1.6% vs 0.1%; scares z=8 2.3% vs 0.0%; me z=8 15.6% vs 6.3%; horrific z=7 1.1% vs 0.0%; terrified z=6 1.6% vs 0.0%; horribly z=6 0.9% vs 0.0%; frightening z=6 0.7% vs 0.0%
Words, gratitude: thanks z=63 47.3% vs 0.5%; thank z=56 38.2% vs 0.4%; for z=37 39.7% vs 11.1%; you z=29 45.1% vs 18.3%; sharing z=17 2.9% vs 0.1%; advice z=16 2.5% vs 0.1%; appreciate z=15 2.7% vs 0.2%; i'll z=15 3.5% vs 0.5%; much z=13 6.6% vs 2.1%; welcome z=13 2.0% vs 0.2%; info z=12 1.5% vs 0.1%; congrats z=11 1.5% vs 0.2%; glad z=11 3.4% vs 1.0%; helpful z=10 1.1% vs 0.1%; very z=9 4.4% vs 1.7%; ll z=9 2.2% vs 0.5%; response z=9 1.2% vs 0.2%; luck z=9 2.0% vs 0.5%; tip z=9 0.8% vs 0.1%; reply z=8 1.1% vs 0.1%
Words, grief: died z=15 22.4% vs 0.1%; rip z=10 10.4% vs 0.1%; loss z=9 10.4% vs 0.2%; dead z=8 10.4% vs 0.2%; death z=7 9.0% vs 0.2%; heart z=6 7.5% vs 0.2%; deaths z=6 3.0% vs 0.0%; peace z=6 4.5% vs 0.1%; condolences z=6 3.0% vs 0.0%; sooner z=5 3.0% vs 0.0%; sorry z=5 14.9% vs 1.7%; disease z=5 3.0% vs 0.0%; passed z=4 3.0% vs 0.1%; lost z=4 6.0% vs 0.4%; lands z=4 1.5% vs 0.0%; depression z=4 3.0% vs 0.1%; 2013 z=4 1.5% vs 0.0%; grandfather z=4 1.5% vs 0.0%; windshield z=4 1.5% vs 0.0%; chills z=4 1.5% vs 0.0%
Words, joy: happy z=40 21.2% vs 0.7%; glad z=35 17.3% vs 0.6%; enjoy z=26 8.7% vs 0.2%; fun z=18 7.0% vs 0.7%; enjoyed z=14 2.7% vs 0.0%; cheers z=12 2.2% vs 0.1%; enjoying z=11 1.6% vs 0.0%; cake z=11 1.9% vs 0.1%; day z=10 4.9% vs 1.1%; joy z=10 1.3% vs 0.0%; i'm z=10 10.4% vs 4.0%; happiness z=8 1.0% vs 0.1%; cakeday z=7 0.8% vs 0.0%; so z=7 13.6% vs 7.6%; m z=7 5.4% vs 2.1%; smile z=7 1.1% vs 0.1%; happily z=7 0.6% vs 0.0%; made z=6 3.3% vs 1.1%; birthday z=6 0.9% vs 0.1%; gladly z=6 0.5% vs 0.0%
Words, love: love z=81 69.0% vs 1.6%; i z=28 66.3% vs 30.6%; loved z=22 5.2% vs 0.2%; favorite z=17 4.0% vs 0.4%; loves z=17 3.0% vs 0.1%; like z=12 14.5% vs 7.0%; loving z=12 1.4% vs 0.1%; my z=10 14.7% vs 7.8%; i'd z=9 2.7% vs 0.7%; name z=7 19.9% vs 14.0%; lovely z=7 0.8% vs 0.1%; favourite z=6 0.7% vs 0.1%; liked z=6 0.9% vs 0.2%; d z=6 2.0% vs 0.7%; this z=6 18.4% vs 13.9%; much z=6 4.2% vs 2.2%; song z=5 0.8% vs 0.2%; sweet z=5 1.0% vs 0.3%; it z=5 21.7% vs 17.2%; how z=5 6.0% vs 3.7%
Words, nervousness: worried z=15 13.8% vs 0.1%; nervous z=14 11.5% vs 0.0%; anxiety z=12 9.2% vs 0.1%; anxious z=11 6.9% vs 0.0%; worrying z=9 4.6% vs 0.0%; worry z=7 5.4% vs 0.2%; career z=6 3.1% vs 0.1%; stressed z=5 1.5% vs 0.0%; scary z=5 3.1% vs 0.1%; paranoid z=5 1.5% vs 0.0%; breathing z=5 1.5% vs 0.0%; nightmares z=5 1.5% vs 0.0%; about z=5 16.2% vs 4.3%; that'll z=4 1.5% vs 0.0%; panic z=4 1.5% vs 0.0%; me z=4 20.0% vs 6.4%; edgy z=4 1.5% vs 0.0%; pick z=4 3.1% vs 0.2%; i'm z=4 14.6% vs 4.2%; standing z=4 1.5% vs 0.0%
Words, neutral: they z=15 8.2% vs 5.0%; name z=15 17.3% vs 12.8%; he z=9 6.9% vs 5.0%; in z=8 14.1% vs 12.1%; on z=7 8.8% vs 7.4%; only z=7 2.8% vs 1.9%; by z=7 2.6% vs 1.8%; the z=7 33.8% vs 32.3%; she z=6 3.6% vs 2.7%; his z=6 3.3% vs 2.5%; or z=6 3.8% vs 3.0%; their z=6 2.5% vs 1.9%; left z=6 0.8% vs 0.4%; 1 z=5 0.9% vs 0.5%; r z=5 0.9% vs 0.5%; from z=5 3.9% vs 3.2%; who z=5 2.8% vs 2.2%; then z=5 2.4% vs 1.8%; go z=5 2.5% vs 1.9%; there's z=5 0.7% vs 0.4%
Words, optimism: hope z=51 36.0% vs 0.8%; luck z=21 7.3% vs 0.4%; hopefully z=21 6.8% vs 0.1%; hoping z=19 5.1% vs 0.1%; will z=13 10.2% vs 2.5%; good z=10 10.5% vs 3.6%; soon z=10 2.2% vs 0.2%; better z=9 5.6% vs 1.5%; wish z=7 3.4% vs 0.8%; can z=7 10.1% vs 4.2%; next z=7 2.7% vs 0.6%; i z=7 50.7% vs 31.5%; best z=7 4.3% vs 1.3%; win z=7 1.9% vs 0.3%; future z=7 1.5% vs 0.2%; goes z=7 1.6% vs 0.3%; probably z=6 3.3% vs 1.0%; get z=6 8.8% vs 4.0%; hopeful z=6 0.5% vs 0.0%; optimistic z=6 0.5% vs 0.0%
Words, pride: proud z=23 38.3% vs 0.1%; pride z=9 6.2% vs 0.0%; jersey z=5 3.7% vs 0.1%; winner z=5 2.5% vs 0.0%; myself z=5 6.2% vs 0.5%; am z=4 9.9% vs 1.3%; butterfly z=4 1.2% vs 0.0%; cubs z=4 1.2% vs 0.0%; european z=4 1.2% vs 0.0%; honored z=4 1.2% vs 0.0%; leaking z=4 1.2% vs 0.0%; integrity z=4 1.2% vs 0.0%; congrats z=4 3.7% vs 0.2%; corn z=4 1.2% vs 0.0%; playthrough z=4 1.2% vs 0.0%; accomplishment z=4 1.2% vs 0.0%; senator z=4 1.2% vs 0.0%; athletic z=4 1.2% vs 0.0%; 2006 z=4 1.2% vs 0.0%; glory z=4 1.2% vs 0.0%
Words, realization: realize z=18 5.8% vs 0.2%; realized z=16 4.2% vs 0.0%; thought z=10 6.1% vs 1.2%; forgot z=10 2.7% vs 0.2%; was z=9 19.3% vs 8.2%; reminds z=8 1.7% vs 0.1%; noticed z=8 1.8% vs 0.2%; until z=8 3.1% vs 0.6%; realizing z=7 0.9% vs 0.0%; realised z=7 0.9% vs 0.0%; figured z=7 1.0% vs 0.1%; ago z=7 2.7% vs 0.5%; realise z=7 0.8% vs 0.0%; didn't z=6 4.3% vs 1.2%; realization z=6 0.6% vs 0.0%; i z=6 48.7% vs 31.7%; aware z=5 0.9% vs 0.1%; wrong z=5 2.9% vs 0.8%; that z=5 28.6% vs 17.9%; ve z=5 2.7% vs 0.8%
Words, relief: glad z=14 22.2% vs 1.0%; relief z=9 4.8% vs 0.0%; god z=8 7.1% vs 0.3%; relieved z=8 3.2% vs 0.0%; least z=8 10.3% vs 0.8%; finally z=6 4.8% vs 0.2%; relax z=6 2.4% vs 0.0%; whew z=5 1.6% vs 0.0%; feel z=5 9.5% vs 1.6%; cool z=5 6.3% vs 0.8%; solved z=5 1.6% vs 0.0%; thank z=5 11.1% vs 2.3%; thankfully z=5 1.6% vs 0.0%; mac z=5 1.6% vs 0.0%; safe z=4 3.2% vs 0.2%; could've z=4 1.6% vs 0.0%; goodness z=4 1.6% vs 0.0%; m z=4 9.5% vs 2.2%; at z=4 15.1% vs 4.7%; oof z=4 1.6% vs 0.1%
Words, remorse: sorry z=59 75.9% vs 0.9%; regret z=16 5.5% vs 0.0%; apologies z=12 3.4% vs 0.0%; m z=11 11.4% vs 2.1%; i'm z=10 15.2% vs 4.1%; apologize z=9 1.9% vs 0.0%; guilty z=8 1.7% vs 0.0%; meant z=8 3.0% vs 0.2%; guilt z=7 1.3% vs 0.0%; i z=7 54.4% vs 31.8%; loss z=7 2.1% vs 0.2%; im z=6 3.4% vs 0.5%; am z=6 5.1% vs 1.3%; apology z=5 0.6% vs 0.0%; misunderstanding z=5 0.6% vs 0.0%; happened z=5 2.5% vs 0.5%; sincerely z=5 0.6% vs 0.0%; didn z=5 2.7% vs 0.6%; misread z=5 0.6% vs 0.0%; should've z=5 0.6% vs 0.0%
Words, sadness: sad z=35 17.3% vs 0.2%; sorry z=23 12.8% vs 1.4%; sadly z=18 5.0% vs 0.0%; miss z=14 3.5% vs 0.3%; hurts z=13 2.4% vs 0.1%; painful z=13 2.3% vs 0.0%; poor z=13 3.5% vs 0.3%; feel z=12 6.9% vs 1.5%; crying z=12 2.2% vs 0.1%; pain z=11 2.3% vs 0.2%; bad z=11 6.8% vs 1.6%; cry z=11 1.9% vs 0.1%; loss z=9 1.8% vs 0.1%; lonely z=8 1.0% vs 0.0%; my z=8 15.1% vs 7.9%; hurt z=8 1.7% vs 0.2%; i'm z=8 9.2% vs 4.1%; lost z=7 1.9% vs 0.3%; hard z=7 3.2% vs 0.8%; sick z=7 1.3% vs 0.2%
Words, surprise: wow z=31 17.0% vs 0.8%; surprised z=30 14.2% vs 0.1%; wonder z=22 7.6% vs 0.2%; omg z=18 6.1% vs 0.3%; wondering z=14 3.2% vs 0.1%; shocked z=14 3.2% vs 0.0%; oh z=14 9.4% vs 1.9%; surprise z=14 2.8% vs 0.1%; believe z=13 5.2% vs 0.7%; surprising z=10 1.4% vs 0.0%; unexpected z=8 1.1% vs 0.0%; wondered z=8 1.1% vs 0.0%; strange z=8 1.3% vs 0.1%; was z=8 16.1% vs 8.2%; amazed z=8 0.9% vs 0.0%; unbelievable z=8 0.9% vs 0.0%; god z=7 2.0% vs 0.3%; shocking z=7 0.8% vs 0.0%; how z=7 8.5% vs 3.7%; holy z=7 1.5% vs 0.2%
2–4-word phrases, admiration: a great z=22 3.9% vs 0.2%; the best z=19 3.8% vs 0.5%; a good z=15 2.8% vs 0.5%; i appreciate z=13 1.4% vs 0.1%; this is z=13 6.0% vs 2.5%; is amazing z=12 1.2% vs 0.0%; is awesome z=12 1.3% vs 0.0%; is great z=12 1.1% vs 0.1%; what a z=11 2.1% vs 0.5%; is the best z=11 0.9% vs 0.1%; good job z=10 0.9% vs 0.0%; such a z=10 1.6% vs 0.4%; so good z=10 0.8% vs 0.1%; a beautiful z=10 0.8% vs 0.0%; pretty good z=9 0.7% vs 0.0%; is a great z=9 0.7% vs 0.0%; so cute z=9 0.6% vs 0.0%; an amazing z=9 0.6% vs 0.0%; a pretty z=8 0.7% vs 0.1%; a nice z=8 0.8% vs 0.1%
2–4-word phrases, amusement: lol i z=16 2.7% vs 0.1%; haha i z=12 1.4% vs 0.0%; me laugh z=11 1.3% vs 0.0%; a joke z=10 1.3% vs 0.1%; made me laugh z=10 1.0% vs 0.0%; i laughed z=10 1.2% vs 0.0%; it lol z=9 0.9% vs 0.0%; fun to z=8 0.7% vs 0.0%; so hard z=8 1.0% vs 0.1%; is hilarious z=8 0.8% vs 0.0%; name lol z=8 0.7% vs 0.0%; lol you z=8 0.7% vs 0.0%; made me z=8 1.4% vs 0.3%; out loud z=8 0.6% vs 0.0%; it's funny z=8 0.6% vs 0.0%; funny how z=8 0.6% vs 0.0%; at this z=7 1.2% vs 0.2%; lol the z=7 0.6% vs 0.0%; lol this z=7 0.7% vs 0.0%; the joke z=7 0.6% vs 0.1%
2–4-word phrases, anger: i hate z=22 6.0% vs 0.3%; the fuck z=19 4.2% vs 0.1%; how dare z=13 2.0% vs 0.0%; what the z=12 2.6% vs 0.3%; what the fuck z=11 1.4% vs 0.0%; the hell z=11 1.5% vs 0.1%; dare you z=11 1.3% vs 0.0%; fuck you z=10 1.4% vs 0.0%; fuck is z=10 1.2% vs 0.0%; fuck off z=10 1.3% vs 0.0%; shut up z=10 1.1% vs 0.0%; how dare you z=10 1.3% vs 0.0%; the fuck is z=10 1.1% vs 0.0%; fuck the z=9 1.2% vs 0.0%; the fucking z=9 0.9% vs 0.0%; a fucking z=9 1.0% vs 0.1%; hate the z=8 0.9% vs 0.0%; you fucking z=8 0.7% vs 0.0%; hate that z=8 0.7% vs 0.0%; what the hell z=8 0.7% vs 0.0%
2–4-word phrases, annoyance: an idiot z=9 0.9% vs 0.0%; a fucking z=7 0.7% vs 0.1%; that sucks z=7 0.5% vs 0.0%; the hell z=6 0.8% vs 0.1%; an asshole z=6 0.4% vs 0.0%; a weird z=6 0.6% vs 0.1%; i hate z=6 1.3% vs 0.4%; can't even z=6 0.4% vs 0.0%; stupid to z=6 0.3% vs 0.0%; what the hell z=5 0.4% vs 0.0%; shut up z=5 0.4% vs 0.1%; is stupid z=5 0.3% vs 0.0%; fuck that z=5 0.3% vs 0.0%; a fuck z=5 0.3% vs 0.0%; waste of z=5 0.3% vs 0.0%; makes no z=5 0.3% vs 0.0%; as hell z=5 0.5% vs 0.1%; tired of z=5 0.3% vs 0.0%; a god z=5 0.3% vs 0.0%; a stupid z=5 0.3% vs 0.0%
2–4-word phrases, approval: i agree z=18 3.8% vs 0.1%; agree with z=14 1.8% vs 0.1%; i agree with z=10 1.0% vs 0.0%; yeah i z=8 1.5% vs 0.3%; agree with you z=8 0.6% vs 0.0%; you're right z=8 0.7% vs 0.1%; pretty sure z=7 0.7% vs 0.1%; you re right z=7 0.6% vs 0.0%; right i z=7 0.7% vs 0.1%; re right z=7 0.6% vs 0.0%; agree that z=7 0.5% vs 0.0%; with you z=7 1.1% vs 0.2%; i think z=6 3.3% vs 1.5%; yes i z=6 0.8% vs 0.2%; agree but z=6 0.4% vs 0.0%; for sure z=6 0.7% vs 0.1%; agree i z=6 0.5% vs 0.0%; yes that z=6 0.4% vs 0.0%; i believe z=6 0.6% vs 0.1%; i completely z=6 0.4% vs 0.0%
2–4-word phrases, caring: don't worry z=12 2.5% vs 0.0%; for you z=12 4.7% vs 0.6%; good luck z=11 3.6% vs 0.3%; you need z=11 2.9% vs 0.2%; take care z=10 1.6% vs 0.0%; be careful z=10 1.5% vs 0.0%; stay strong z=9 1.5% vs 0.0%; stay safe z=9 1.8% vs 0.0%; name bless z=8 1.2% vs 0.0%; care of z=8 1.3% vs 0.1%; keep your z=8 1.1% vs 0.0%; you have z=8 3.8% vs 0.8%; if you z=8 4.9% vs 1.2%; t worry z=8 1.0% vs 0.0%; don t worry z=8 1.0% vs 0.0%; don't be z=7 1.0% vs 0.0%; take care of z=7 0.9% vs 0.0%; you feel z=7 1.4% vs 0.1%; get some z=7 1.1% vs 0.1%; no worries z=7 0.9% vs 0.0%
2–4-word phrases, confusion: not sure z=23 7.8% vs 0.2%; i don't z=15 7.7% vs 1.0%; don't know z=14 4.1% vs 0.3%; i don't know z=13 3.1% vs 0.2%; no idea z=13 2.9% vs 0.1%; sure what z=12 2.3% vs 0.0%; i'm not sure z=11 1.9% vs 0.0%; don t know z=11 2.1% vs 0.1%; not sure what z=11 1.7% vs 0.0%; have no idea z=11 1.8% vs 0.1%; i don t know z=10 1.7% vs 0.1%; i have no idea z=10 1.6% vs 0.0%; i have no z=10 1.8% vs 0.1%; t know z=10 2.4% vs 0.2%; m not sure z=10 1.4% vs 0.0%; i m not sure z=10 1.4% vs 0.0%; sure if z=10 1.4% vs 0.0%; i don z=9 3.5% vs 0.5%; i don t z=9 3.5% vs 0.5%; know what z=9 2.6% vs 0.3%
2–4-word phrases, curiosity: do you z=22 6.5% vs 0.5%; are you z=18 4.9% vs 0.5%; did you z=16 3.2% vs 0.2%; is it z=14 2.7% vs 0.2%; is this z=13 2.7% vs 0.3%; can you z=11 1.7% vs 0.1%; what are z=11 1.4% vs 0.1%; is that z=10 2.8% vs 0.5%; do you have z=10 1.2% vs 0.1%; just curious z=10 1.4% vs 0.0%; have you z=10 1.3% vs 0.1%; how do z=9 1.1% vs 0.1%; you think z=9 1.8% vs 0.3%; how did z=9 1.0% vs 0.0%; what do z=9 1.1% vs 0.1%; why is z=9 1.1% vs 0.1%; what are you z=9 0.9% vs 0.0%; how is z=9 0.9% vs 0.0%; is there z=9 1.2% vs 0.1%; do you think z=9 1.0% vs 0.1%
2–4-word phrases, desire: i wish z=36 29.9% vs 0.2%; wish i z=24 13.1% vs 0.1%; i wish i z=22 10.9% vs 0.1%; wish i could z=17 6.6% vs 0.1%; i want z=17 8.5% vs 0.4%; i could z=16 7.7% vs 0.3%; i wish i could z=16 5.5% vs 0.1%; i want to z=12 4.2% vs 0.2%; want to z=10 6.6% vs 0.8%; wish we z=10 2.6% vs 0.0%; i just want z=9 2.2% vs 0.1%; wish i had z=9 2.0% vs 0.0%; i wanna z=9 2.2% vs 0.1%; wish the z=9 2.4% vs 0.0%; i need z=9 3.0% vs 0.2%; just want z=9 2.2% vs 0.1%; wish they z=8 1.7% vs 0.0%; wish it z=8 1.7% vs 0.0%; wish name z=8 1.3% vs 0.0%; i wish name z=8 1.3% vs 0.0%
2–4-word phrases, disappointment: my bad z=8 1.4% vs 0.1%; i miss z=8 1.5% vs 0.1%; i tried z=7 1.0% vs 0.1%; too bad z=7 1.1% vs 0.1%; disappointed in z=7 0.8% vs 0.0%; bad for z=6 1.0% vs 0.1%; missed it z=6 0.6% vs 0.0%; i missed z=6 0.8% vs 0.1%; so bad z=6 0.9% vs 0.1%; i tried to z=6 0.6% vs 0.0%; a disaster z=6 0.5% vs 0.0%; i lost z=5 0.6% vs 0.0%; oh no z=5 0.8% vs 0.1%; it didn z=5 0.5% vs 0.0%; it didn t z=5 0.5% vs 0.0%; tried to z=5 0.8% vs 0.1%; but no z=5 0.6% vs 0.0%; so disappointed z=5 0.4% vs 0.0%; filled with z=5 0.5% vs 0.0%; this game z=5 0.9% vs 0.2%
2–4-word phrases, disapproval: i don't z=14 5.6% vs 1.0%; don t z=13 5.9% vs 1.3%; don't think z=12 2.3% vs 0.2%; i don z=12 3.2% vs 0.5%; i don t z=12 3.2% vs 0.5%; that's not z=12 1.8% vs 0.1%; is not z=12 2.7% vs 0.4%; i don't think z=11 2.0% vs 0.2%; not a z=10 2.5% vs 0.4%; it's not z=10 2.3% vs 0.4%; i dont z=9 1.5% vs 0.2%; don t think z=8 1.1% vs 0.1%; no it z=8 0.8% vs 0.0%; t think z=8 1.2% vs 0.1%; don t want z=8 0.7% vs 0.0%; disagree with z=8 0.7% vs 0.0%; don't like z=8 0.9% vs 0.1%; i don t think z=7 0.9% vs 0.1%; not true z=7 0.6% vs 0.0%; no way z=7 0.9% vs 0.1%
2–4-word phrases, disgust: the worst z=18 6.2% vs 0.2%; is the worst z=10 1.6% vs 0.0%; is disgusting z=9 1.9% vs 0.0%; is awful z=8 1.2% vs 0.0%; worse than z=8 1.5% vs 0.1%; i've ever z=8 1.5% vs 0.1%; an awful z=7 1.0% vs 0.0%; even worse z=7 1.0% vs 0.0%; a weird z=7 1.2% vs 0.1%; i've ever seen z=7 1.0% vs 0.0%; weird to z=7 0.9% vs 0.0%; this is the worst z=7 0.7% vs 0.0%; it's weird z=6 0.7% vs 0.0%; ever seen z=6 1.2% vs 0.1%; i hate z=6 2.2% vs 0.4%; disgusting and z=6 0.7% vs 0.0%; that's awful z=6 0.6% vs 0.0%; be worse z=6 0.6% vs 0.0%; absolute worst z=6 0.6% vs 0.0%; is weird z=6 0.6% vs 0.0%
2–4-word phrases, embarrassment: a shame z=8 2.7% vs 0.1%; be ashamed z=7 1.5% vs 0.0%; my bad z=6 2.3% vs 0.1%; s a shame z=6 1.2% vs 0.0%; i wear z=6 1.2% vs 0.0%; ashamed of z=6 1.2% vs 0.0%; ashamed to z=6 1.2% vs 0.0%; of shame z=6 1.2% vs 0.0%; it s a shame z=6 1.2% vs 0.0%; an embarrassment z=5 1.5% vs 0.0%; my body z=5 1.2% vs 0.0%; sorry i z=5 2.3% vs 0.2%; second hand z=5 0.8% vs 0.0%; even know that z=5 0.8% vs 0.0%; it was weird z=5 0.8% vs 0.0%; body and z=5 0.8% vs 0.0%; was weird z=5 0.8% vs 0.0%; to much z=5 0.8% vs 0.0%; an easy z=5 0.8% vs 0.0%; be ashamed of z=5 0.8% vs 0.0%
2–4-word phrases, excitement: wait to z=14 3.2% vs 0.0%; can't wait z=13 3.0% vs 0.1%; t wait z=12 2.6% vs 0.0%; can t wait z=12 2.6% vs 0.0%; excited to z=12 2.7% vs 0.0%; wait for z=11 2.3% vs 0.1%; excited for z=11 3.3% vs 0.0%; happy new z=10 2.0% vs 0.1%; cake day z=10 2.0% vs 0.1%; so excited z=10 1.9% vs 0.0%; happy new year z=10 1.7% vs 0.1%; excited to see z=9 1.6% vs 0.0%; happy birthday z=9 1.4% vs 0.0%; happy cake day z=9 1.6% vs 0.1%; happy cake z=9 1.6% vs 0.1%; can t wait to z=9 1.4% vs 0.0%; t wait to z=9 1.4% vs 0.0%; wait to see z=9 1.3% vs 0.0%; i can't wait z=9 1.3% vs 0.0%; new year z=9 1.7% vs 0.1%
2–4-word phrases, fear: a terrible z=13 3.3% vs 0.1%; afraid of z=13 3.3% vs 0.0%; afraid to z=11 2.3% vs 0.0%; scared of z=10 2.5% vs 0.0%; i'm afraid z=10 2.6% vs 0.0%; a horrible z=9 1.9% vs 0.1%; scared to z=8 2.8% vs 0.0%; is terrible z=8 1.4% vs 0.0%; worried about z=8 1.6% vs 0.1%; fear of z=8 1.2% vs 0.0%; i'm scared z=7 1.8% vs 0.0%; terrible idea z=7 1.2% vs 0.0%; a terrible idea z=7 1.2% vs 0.0%; was afraid z=7 1.2% vs 0.0%; out of me z=7 0.9% vs 0.0%; be afraid z=6 0.9% vs 0.0%; afraid i z=6 1.1% vs 0.0%; is terrifying z=6 1.1% vs 0.0%; is scary z=6 0.9% vs 0.0%; more worried about z=6 0.7% vs 0.0%
2–4-word phrases, gratitude: thank you z=50 35.2% vs 0.3%; thanks for z=37 18.9% vs 0.2%; you for z=30 10.0% vs 0.2%; for the z=29 12.9% vs 1.4%; thank you for z=27 10.0% vs 0.1%; thanks for the z=24 8.9% vs 0.1%; you i z=18 3.6% vs 0.2%; for your z=17 3.7% vs 0.3%; thanks i z=17 3.1% vs 0.1%; you so z=16 2.9% vs 0.1%; thank you i z=16 3.5% vs 0.0%; for sharing z=15 2.7% vs 0.1%; you so much z=15 2.7% vs 0.0%; thank you so z=14 2.9% vs 0.0%; thank you so much z=14 2.7% vs 0.0%; so much z=13 3.7% vs 0.6%; thanks for sharing z=12 1.6% vs 0.0%; you for the z=12 2.2% vs 0.0%; you for your z=12 1.8% vs 0.0%; i will z=12 2.5% vs 0.3%
2–4-word phrases, grief: your loss z=8 7.5% vs 0.1%; sorry about your z=7 4.5% vs 0.0%; sorry for z=7 9.0% vs 0.2%; sorry about z=6 4.5% vs 0.0%; passed away z=6 3.0% vs 0.0%; my condolences z=6 3.0% vs 0.0%; your friend z=6 4.5% vs 0.1%; for your loss z=5 4.5% vs 0.1%; sorry for your loss z=5 4.5% vs 0.1%; he died z=5 3.0% vs 0.0%; very sorry z=5 3.0% vs 0.0%; my only z=5 3.0% vs 0.0%; sorry for your z=5 4.5% vs 0.1%; this may z=5 3.0% vs 0.0%; about your z=5 4.5% vs 0.1%; to death z=5 3.0% vs 0.0%; two years z=5 3.0% vs 0.0%; after he z=5 3.0% vs 0.0%; to die z=4 3.0% vs 0.1%; he could z=4 3.0% vs 0.1%
2–4-word phrases, joy: glad you z=17 4.0% vs 0.1%; so happy z=16 3.6% vs 0.1%; glad to z=15 3.1% vs 0.1%; happy to z=14 2.9% vs 0.1%; i'm glad z=14 3.1% vs 0.1%; happy for z=13 2.5% vs 0.1%; so glad z=13 2.3% vs 0.1%; glad i z=12 2.1% vs 0.1%; i enjoy z=11 1.9% vs 0.0%; cake day z=11 1.7% vs 0.1%; enjoy it z=11 1.6% vs 0.0%; happy cake day z=11 1.6% vs 0.0%; happy cake z=11 1.6% vs 0.0%; me happy z=10 1.5% vs 0.0%; glad to see z=10 1.5% vs 0.0%; i enjoyed z=10 1.7% vs 0.0%; be happy z=10 1.5% vs 0.1%; enjoy the z=10 1.3% vs 0.0%; happy for you z=10 1.3% vs 0.0%; m glad z=10 1.4% vs 0.1%
2–4-word phrases, love: i love z=54 36.9% vs 0.5%; love it z=26 8.0% vs 0.2%; i like z=25 8.2% vs 0.5%; love the z=23 6.2% vs 0.2%; love this z=22 6.1% vs 0.1%; love to z=21 5.1% vs 0.1%; i love it z=19 4.6% vs 0.1%; would love z=17 3.4% vs 0.1%; i love this z=17 4.1% vs 0.0%; i love the z=17 3.4% vs 0.1%; love that z=16 3.4% vs 0.0%; love you z=16 3.3% vs 0.0%; love how z=16 3.5% vs 0.0%; love name z=15 2.8% vs 0.1%; my favorite z=15 3.3% vs 0.2%; i love how z=14 2.9% vs 0.0%; would love to z=14 2.2% vs 0.1%; i'd love z=14 2.2% vs 0.0%; i loved z=14 2.2% vs 0.1%; i would love z=13 2.0% vs 0.0%
2–4-word phrases, nervousness: i'm worried z=7 3.1% vs 0.0%; worried that z=7 3.1% vs 0.0%; worried about z=7 3.8% vs 0.1%; worried about the z=6 2.3% vs 0.0%; i'm worried that z=6 2.3% vs 0.0%; worried for z=6 3.1% vs 0.0%; to pick z=6 2.3% vs 0.0%; very worried z=5 1.5% vs 0.0%; get older z=5 1.5% vs 0.0%; so don't z=5 1.5% vs 0.0%; be worried z=5 1.5% vs 0.0%; so worried z=5 1.5% vs 0.0%; i can feel z=5 1.5% vs 0.0%; more worried about z=5 1.5% vs 0.0%; i have this z=5 1.5% vs 0.0%; more worried z=5 1.5% vs 0.0%; me off z=5 1.5% vs 0.0%; worrying about z=5 1.5% vs 0.0%; can feel z=4 1.5% vs 0.0%; makes me z=4 4.6% vs 0.5%
2–4-word phrases, neutral: name and z=8 1.4% vs 0.8%; in the z=7 3.7% vs 3.0%; and name z=7 1.2% vs 0.7%; name is z=6 1.8% vs 1.3%; name name z=6 0.6% vs 0.2%; they are z=6 1.0% vs 0.6%; name and name z=6 0.8% vs 0.5%; if they z=6 0.6% vs 0.3%; with name z=6 0.6% vs 0.3%; because they z=6 0.4% vs 0.2%; they don't z=6 0.3% vs 0.1%; on the z=5 1.9% vs 1.5%; name was z=5 0.7% vs 0.4%; name got z=5 0.2% vs 0.0%; he was z=5 0.9% vs 0.6%; looks like z=5 0.8% vs 0.5%; they can z=5 0.3% vs 0.1%; because she z=5 0.2% vs 0.0%; does not z=5 0.2% vs 0.1%; on name z=5 0.2% vs 0.1%
2–4-word phrases, optimism: i hope z=36 18.7% vs 0.4%; hope you z=22 7.1% vs 0.2%; good luck z=18 5.5% vs 0.3%; i hope you z=18 4.7% vs 0.1%; really hope z=12 2.2% vs 0.0%; hope it z=12 2.1% vs 0.0%; hope that z=12 2.1% vs 0.0%; i really hope z=12 2.1% vs 0.0%; hope he z=12 2.0% vs 0.0%; hope this z=11 1.9% vs 0.0%; hope the z=11 1.8% vs 0.0%; hope they z=11 1.7% vs 0.0%; hope she z=9 1.3% vs 0.0%; hope we z=9 1.2% vs 0.0%; hope you're z=9 1.2% vs 0.0%; was hoping z=9 1.2% vs 0.0%; hope for z=9 1.2% vs 0.0%; hope your z=9 1.2% vs 0.0%; i hope it z=9 1.2% vs 0.0%; just hope z=9 1.2% vs 0.0%
2–4-word phrases, pride: proud of z=16 21.0% vs 0.1%; so proud z=13 13.6% vs 0.0%; so proud of z=10 8.6% vs 0.0%; proud of you z=10 8.6% vs 0.1%; be proud z=8 4.9% vs 0.0%; of you z=7 8.6% vs 0.2%; i'm so proud z=7 4.9% vs 0.0%; so proud of you z=7 3.7% vs 0.0%; should be proud z=6 3.7% vs 0.0%; i'm so proud of z=6 3.7% vs 0.0%; of you and z=6 2.5% vs 0.0%; m proud z=6 2.5% vs 0.0%; i m proud z=6 2.5% vs 0.0%; proud to z=6 2.5% vs 0.0%; i'm proud z=5 2.5% vs 0.0%; be proud of z=5 2.5% vs 0.0%; this community z=5 2.5% vs 0.0%; i am just z=5 2.5% vs 0.0%; been so z=5 2.5% vs 0.0%; am just z=5 2.5% vs 0.0%
2–4-word phrases, realization: i realized z=10 2.2% vs 0.0%; realize that z=9 1.6% vs 0.1%; i forgot z=9 1.6% vs 0.1%; i thought z=8 3.7% vs 0.6%; i didn't z=8 2.9% vs 0.4%; didn't realize z=8 1.2% vs 0.0%; reminds me z=8 1.6% vs 0.1%; you realize z=8 1.0% vs 0.0%; i didn't realize z=8 1.0% vs 0.0%; just realized z=8 1.0% vs 0.0%; reminds me of z=8 1.5% vs 0.1%; me of z=8 1.7% vs 0.1%; i was z=7 6.4% vs 1.8%; i thought it z=7 1.5% vs 0.1%; it took z=7 0.9% vs 0.0%; realized that z=6 0.8% vs 0.0%; realize it z=6 0.7% vs 0.0%; thought it z=6 1.6% vs 0.2%; to realize z=6 0.7% vs 0.0%; but then i z=6 0.7% vs 0.0%
2–4-word phrases, relief: thank god z=10 5.6% vs 0.0%; m glad z=9 5.6% vs 0.1%; i m glad z=9 5.6% vs 0.1%; glad i z=8 5.6% vs 0.1%; at least z=7 10.3% vs 0.7%; m glad i z=7 3.2% vs 0.0%; i m glad i z=7 3.2% vs 0.0%; have to worry z=6 2.4% vs 0.0%; have to worry about z=6 2.4% vs 0.0%; glad i'm not z=6 2.4% vs 0.0%; makes me feel z=6 3.2% vs 0.1%; glad i'm z=6 2.4% vs 0.0%; someone i z=6 2.4% vs 0.0%; this makes me feel z=6 2.4% vs 0.0%; thank name z=6 3.2% vs 0.1%; so glad z=6 4.0% vs 0.1%; better about z=6 2.4% vs 0.0%; the only one z=6 4.0% vs 0.1%; don't have to z=6 2.4% vs 0.0%; to worry about z=6 2.4% vs 0.0%
2–4-word phrases, remorse: sorry i z=20 10.1% vs 0.1%; so sorry z=20 10.1% vs 0.1%; i'm sorry z=19 9.3% vs 0.1%; sorry for z=19 8.6% vs 0.1%; i m sorry z=16 6.1% vs 0.1%; m sorry z=16 6.1% vs 0.1%; sorry but z=15 5.5% vs 0.1%; sorry you z=14 5.1% vs 0.1%; i m so sorry z=13 4.0% vs 0.0%; m so sorry z=13 4.0% vs 0.0%; sorry to z=12 3.8% vs 0.1%; m so z=12 4.2% vs 0.1%; i m so z=12 4.2% vs 0.1%; sorry that z=11 3.0% vs 0.1%; i'm so sorry z=10 2.7% vs 0.0%; sorry for the z=10 2.5% vs 0.0%; sorry for your z=10 2.5% vs 0.1%; i m z=10 11.4% vs 2.1%; i regret z=9 2.3% vs 0.0%; i meant z=9 2.5% vs 0.1%
2–4-word phrases, sadness: so sad z=12 2.1% vs 0.0%; i feel z=12 4.3% vs 0.6%; bad for z=12 2.0% vs 0.1%; i miss z=11 2.0% vs 0.1%; sorry for z=11 2.2% vs 0.2%; feel bad z=11 1.7% vs 0.1%; i'm sorry z=10 2.2% vs 0.2%; feel bad for z=10 1.5% vs 0.0%; sad that z=10 1.7% vs 0.0%; sorry to z=9 1.4% vs 0.1%; so sorry z=9 1.9% vs 0.2%; me sad z=9 1.3% vs 0.0%; i m sorry z=9 1.6% vs 0.1%; m sorry z=9 1.6% vs 0.1%; sorry for your z=9 1.2% vs 0.1%; sorry that z=9 1.2% vs 0.1%; i feel bad z=8 1.0% vs 0.0%; for your loss z=8 1.1% vs 0.0%; sorry for your loss z=8 1.1% vs 0.0%; your loss z=8 1.1% vs 0.1%
2–4-word phrases, surprise: i wonder z=18 4.7% vs 0.1%; wow i z=13 2.9% vs 0.1%; wonder if z=13 2.5% vs 0.0%; can't believe z=12 2.1% vs 0.1%; wonder how z=12 2.0% vs 0.0%; i wonder if z=11 1.8% vs 0.0%; oh my z=11 2.4% vs 0.2%; t believe z=11 1.7% vs 0.0%; was wondering z=10 1.7% vs 0.0%; be surprised z=10 1.8% vs 0.1%; can t believe z=9 1.4% vs 0.0%; wonder why z=9 1.3% vs 0.0%; wow you z=9 1.3% vs 0.0%; i wonder how z=9 1.2% vs 0.0%; surprised if z=9 1.2% vs 0.0%; a surprise z=9 1.2% vs 0.0%; i'm surprised z=8 2.5% vs 0.0%; i was wondering z=8 1.5% vs 0.0%; wonder what z=8 1.1% vs 0.0%; i can't believe z=8 1.1% vs 0.0%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- remorse:
sorry75.9% vs 0.9% - love:
love69.0% vs 1.6% - gratitude:
thanks47.3% vs 0.5% - amusement:
lol42.6% vs 0.9% - desire:
wish39.7% vs 0.5% - pride:
proud38.3% vs 0.1% - gratitude:
thank38.2% vs 0.4% - love:
i love36.9% vs 0.5% - optimism:
hope36.0% vs 0.8% - gratitude:
thank you35.2% vs 0.3% - desire:
i wish29.9% vs 0.2% - grief:
died22.4% vs 0.1% - relief:
glad22.2% vs 1.0% - joy:
happy21.2% vs 0.7% - pride:
proud of21.0% vs 0.1% - gratitude:
thanks for18.9% vs 0.2% - optimism:
i hope18.7% vs 0.4% - sadness:
sad17.3% vs 0.2% - joy:
glad17.3% vs 0.6% - surprise:
wow17.0% vs 0.8% - anger:
fuck15.0% vs 0.5% - grief:
sorry14.9% vs 1.7% - desire:
want14.8% vs 1.7% - surprise:
surprised14.2% vs 0.1% - admiration:
good14.1% vs 2.9% - admiration:
great14.1% vs 0.8% - nervousness:
worried13.8% vs 0.1% - pride:
so proud13.6% vs 0.0% - disgust:
disgusting13.5% vs 0.0% - desire:
wish i13.1% vs 0.1% - gratitude:
for the12.9% vs 1.4% - sadness:
sorry12.8% vs 1.4% - confusion:
or12.5% vs 3.0% - desire:
could11.6% vs 1.5% - embarrassment:
shame11.5% vs 0.2% - embarrassment:
embarrassing11.5% vs 0.0% - nervousness:
nervous11.5% vs 0.0% - confusion:
why11.5% vs 2.0% - remorse:
m11.4% vs 2.1% - remorse:
i m11.4% vs 2.1%
Shortcut models
Predicting the label class on test (5,408 rows). Chance 3.6%, majority class ('neutral') 31.1%; balanced chance 3.6%.
| model (logistic regression, trained on the train sample) | test accuracy | balanced accuracy (mean recall) |
|---|---|---|
| bag of words, whole state (words and word pairs) | 53.2% | 32.0% |
bag of words, main text only (comment) |
41.7% | 20.5% |
| surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) | 31.1% | 5.0% |
| surface features of the main text only | 31.0% | 4.8% |
Strongest single surface features (logistic regression on one feature, balanced accuracy on test):
| feature | accuracy | balanced accuracy |
|---|---|---|
| ends_? | 31.3% | 5.2% |
| count_? | 31.0% | 3.8% |
| count_! | 30.8% | 3.6% |
| chars(log) | 31.1% | 3.6% |
| words(log) | 31.1% | 3.6% |
| upper_ratio | 31.1% | 3.6% |
| digit_ratio | 31.1% | 3.6% |
| nonlatin | 31.1% | 3.6% |
| emoji | 31.1% | 3.6% |
| html_tag | 31.1% | 3.6% |
Other state fields alone (predicting the label class on test from one field, without the main text):
| field | treated as | accuracy | balanced accuracy |
|---|---|---|---|
| source | categorical, 1 values | 31.1% | 3.6% |
No-meaning option picker: a logistic ranker scores each option from its position, length, key type, fixed-option identity and shape (commas, brackets, capitals), never reading the state or the option's words, and picks the top option per row.
- Test accuracy 31.1% against uniform chance 3.6% (this includes the fixed options, whose share is a class prior).
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
- Train rows identical in the whole prompt (state, options, instructions): 0.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 0 (0.0%) | |
| calibration | 0 (0.0%) | |
| test | 0 (0.0%) |
Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)
- Train: 39 near-duplicate pairs; 73 rows (0.2%) sit in 35 clusters; largest cluster 4; excess rows (cluster size − 1) 38 (0.1%).
- Clusters with more than one label class: 15 (32 rows).
- ×4: "I'm sorry for your loss" → caring 1, grief 1, remorse 1, amusement 1
- ×2: "He's hot" → neutral 1, admiration 1
- ×2: "[NAME]. What a time to be alive." → neutral 1, approval 1
- ×2: "because it’s funny" → amusement 1, confusion 1
- ×2: "Oh that’s terrifying" → surprise 1, fear 1
- Held-out rows with a near duplicate in train: dev 2 (0.2%), calibration 0 (0.0%), test 6 (0.1%)
- train "Weird flex but okay" ~ test "Weird flex but ok" (J=0.87)
- train "What? That doesn’t answer my question." ~ test "That doesn't answer my question." (J=0.80)
- train "Love the username" ~ test "Love the username <3" (J=0.87)
- train "Thanks a bunch!" ~ test "Thanks a bunch <3" (J=0.83)
- train "Yes YES #YES" ~ test "* yes * yes * yes * yes" (J=1.00)
Largest train clusters:
- ×4: "I'm sorry for your loss"
- ×3: "HAHAHAHA!!!!!"
- ×2: "He's hot"
- ×2: "RemindMe! 7 days"
- ×2: "RemindMe! 3 Days"
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 240 |
| dev | 0 | 5 |
| calibration | 0 | 6 |
| test | 0 | 40 |
Very short train examples: "He's hot" (neutral); "by [NAME]" (neutral); "Hey now!" (approval); "$390 CAD." (neutral); "[NAME] 3!" (neutral); "Go you!" (neutral); "Shame." (neutral); "Wow. Yes" (approval); "funny as." (amusement); "Or both!" (neutral); "Lol, wut" (amusement); "My bad." (neutral)
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 6,405 (13.9%) | admiration 14.5%, amusement 13.0%, anger 13.9%, annoyance 12.3%, approval 11.4%, caring 9.2%, confusion 12.9%, curiosity 13.1%, desire 17.0%, disappointment 13.9%, disapproval 10.6%, disgust 11.4%, embarrassment 7.3%, excitement 13.0%, fear 13.5%, gratitude 7.9%, grief 14.9%, joy 9.3%, love 19.4%, nervousness 7.7%, neutral 17.0%, optimism 12.9%, pride 11.1%, realization 10.3%, relief 10.3%, remorse 9.7%, sadness 10.2%, surprise 15.9% |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 1 (0.0%) | annoyance 0.0% |
| chat preamble ('Here is/are...', 'Sure!') | 39 (0.1%) | admiration 0.1%, approval 0.4%, curiosity 0.1%, disappointment 0.1%, disapproval 0.1%, excitement 0.3%, gratitude 0.0%, joy 0.1%, neutral 0.1%, optimism 0.1%, sadness 0.2% |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 0 (0.0%) | |
| encoding garbage (mojibake/replacement char) | 0 (0.0%) | |
| HTML tag | 3 (0.0%) | approval 0.0%, neutral 0.0% |
| HTML entity | 0 (0.0%) | |
| base64-like run (40+ chars) | 11 (0.0%) | admiration 0.0%, amusement 0.0%, anger 0.1%, confusion 0.1%, curiosity 0.1%, neutral 0.0%, surprise 0.1% |
| URL | 0 (0.0%) |
placeholder [NAME]-style:
emotion:go_emotions:train:efeo8w6:emotion(admiration): Unreal. [NAME]- Expert level |emotion:go_emotions:train:eeezcc7:emotion(disapproval): [NAME] would've had to get lucky on that one. Can't give up that opportunity. |emotion:go_emotions:train:eeajm3e:emotion(approval): I bet he’s not going to visit their hall of fame in order to avoid meeting w/ [NAME]'As an AI' / refusal:
emotion:go_emotions:train:ef9ebro:emotion(annoyance): I cannot help but get a sense of schadenfreude when a seemingly shitty person ruins their own career right before my eyes. ☕️chat preamble ('Here is/are...', 'Sure!'):
emotion:go_emotions:train:eebqp3y:emotion(neutral): Sure, there are edge cases where this is less clear, but in the vast majority of conflict, this is true. |emotion:go_emotions:train:ee7fxah:emotion(neutral): Sure, its just that it was completely irrelevant to my reply |emotion:go_emotions:train:ed1m9qs:emotion(approval): Sure, they can help nurse [NAME] back to health.HTML tag:
emotion:go_emotions:train:ee5v5uy:emotion(neutral): <<creepy loui face.jpg>> |emotion:go_emotions:train:edfd4z0:emotion(approval): Yes, definitely classy. <rolls eyes> |emotion:go_emotions:train:ef8b1z0:emotion(neutral): If by "better", you mean "more fun", then I'd go with the second one. <smile>base64-like run (40+ chars):
emotion:go_emotions:train:ef9cbbq:emotion(surprise): Oh no, my secret identity as a neurosurgeon/entomologist/radiologist/anesthetist/dentist has been discovered! I knew I should have given le… |emotion:go_emotions:train:edrtk8u:emotion(amusement): I did anyway hahahahahahahahahahahahahahahahahahahahahahahahahahhahahahhahahaa fucking rekt XD |emotion:go_emotions:train:ef59tsi:emotion(neutral): Omg ur like soooooooooooooooooooooooooooooooooooooooooooooooooo quirky!Possibly cut off: 1 of 4 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: admiration 100.0%, neutral 0.0%
emotion:go_emotions:train:edy8hly:emotion: …765434567654323454323456543345678987654323456789876565656565656565656565656565454545654565454323456765432345678765456 IQ
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- none
6. Samples
20 random train rows per kind: emotion-samples.txt. Reading notes are in the findings above.
