Skip to content
JeffHub

QA report: emotion

The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.

The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.

QA: emotion

Checked 2026-09-30 11:12 by adapters/qa/qa.py (READY file ready-emotion.json, public v1 (READY 2026-09-29T23:56)).

Verdict: PASS WITH NOTES (for v2)

Findings

  1. No shortcut. A surface-only model sits at the majority baseline (31.1% accuracy, 5% balanced against 3.6% chance). The strong phrases ("sorry" for remorse, "thanks" for gratitude, "lol" for amusement) are the meaning of the labels. The GoEmotions placeholders "[NAME]" and "[RELIGION]" appear in 13.9% of rows, spread evenly across classes (9–17%).
  2. Label noise is built into the source. GoEmotions is multi-label, and v1 keeps one label per comment. Near-identical comments carry different labels: "I'm sorry for your loss" is labelled caring, grief, remorse and amusement. The BoW model's balanced accuracy is only 32%, which fits a noisy task. v2: turn the raters' labels into soft targets (the counts are in source.gold_labels and upstream_ids), and/or merge rare, confusable classes.
  3. The class balance is heavily skewed: neutral is 31% of test. Duplicates are negligible (0.1%), and held-out near duplicates are 0.1–0.2%.

Automatic flags (for the reviewer to judge; not all are problems)

  • formatting 'ends with ?' differs by class: curiosity 51%, confusion 32%, neutral 5%, surprise 5% …
  • formatting 'ends with .' differs by class: disapproval 62%, realization 58%, approval 57%, disappointment 56% …
  • formatting 'ends with !' differs by class: excitement 29%, gratitude 23%, joy 15%, admiration 15% …
  • formatting 'no end punctuation' differs by class: amusement 54%, grief 43%, sadness 41%, love 40% …
  • 157 strong phrase flags (see list): review whether they are meaning or leakage

Data checked

split rows families file
train 46,067 46067 train.jsonl
dev 1,000 1000 dev.jsonl
calibration 1,000 1000 calibration.jsonl
test 5,408 5408 test.jsonl
  • Train sha256: 30867b255a4e4ca1eb9ec1f3116245aebacd6cf702a0a9adc0e71c87cd1c4bcd (READY file gives no checksum)
  • Main text field (the text the phrase and length checks use): state.comment.
  • Label classes: admiration, amusement, anger, annoyance, approval, caring, confusion, curiosity, desire, disappointment, disapproval, disgust, embarrassment, excitement, fear, gratitude, grief, joy, love, nervousness, neutral, optimism, pride, realization, relief, remorse, sadness, surprise (<listed option> = one of the per-row listed options such as t3 or o12). Row kinds (source.kind): emotion.

3. Balance

Label class share per split

lclass train dev calibration test train rows
admiration 7.8% 7.4% 7.7% 7.8% 3,590
amusement 4.6% 5.7% 5.4% 4.2% 2,102
anger 2.9% 3.0% 2.9% 3.1% 1,357
annoyance 4.5% 3.8% 3.7% 4.8% 2,082
approval 5.6% 5.7% 7.6% 5.2% 2,566
caring 2.0% 2.5% 2.0% 2.0% 910
confusion 2.6% 2.5% 1.7% 2.4% 1,180
curiosity 4.1% 2.8% 5.4% 4.2% 1,872
desire 1.2% 0.5% 1.7% 1.3% 542
disappointment 2.3% 2.3% 1.8% 2.0% 1,061
disapproval 3.9% 4.6% 5.3% 4.3% 1,818
disgust 1.5% 1.3% 1.2% 1.8% 675
embarrassment 0.6% 0.6% 0.4% 0.6% 260
excitement 1.5% 0.9% 1.6% 1.4% 693
fear 1.2% 1.0% 1.1% 1.3% 570
gratitude 4.9% 6.0% 5.3% 5.4% 2,279
grief 0.1% 0.2% 0.0% 0.1% 67
joy 2.6% 2.2% 2.7% 2.4% 1,188
love 3.9% 3.5% 3.4% 3.7% 1,808
nervousness 0.3% 0.2% 0.4% 0.3% 130
neutral 31.3% 33.0% 28.1% 31.1% 14,417
optimism 2.8% 2.8% 3.5% 2.7% 1,293
pride 0.2% 0.4% 0.3% 0.2% 81
realization 1.9% 2.8% 2.0% 2.3% 865
relief 0.3% 0.0% 0.3% 0.1% 126
remorse 1.0% 0.7% 0.8% 0.9% 474
sadness 2.5% 1.3% 2.0% 2.4% 1,129
surprise 2.0% 2.3% 1.7% 2.0% 932

Row kind share per split

kind train dev calibration test train rows
emotion 100.0% 100.0% 100.0% 100.0% 46,067

4. Format

split row-level format problems
train none
dev none
calibration none
test none

Options per choice row

split min median p99 max
train 28 28 28 28
dev 28 28 28 28
calibration 28 28 28 28
test 28 28 28 28

Prompt length in tokens

split measure median p99 max > 8192
train estimate: characters / 3 (upper bound for English) 669 701 870 0
dev estimate: characters / 3 (upper bound for English) 669 705 710 0
calibration estimate: characters / 3 (upper bound for English) 669 701 708 0
test estimate: characters / 3 (upper bound for English) 668 700 713 0

State key sets (train)

keys rows
comment, source 46,067 (100.0%)

Instructions (train)

  • Canonical (the most common text) 70.1%, reworded 26.9% (20 distinct rewordings), none 3.0%. Target about 70 / 27 / 3.
  • Canonical text: "Which emotion does the comment express most? Choose the emotion that is strongest in the comment, or neutral if it expresses no particular emotion."
  • source instruction tag: canonical 70.1%, none 3.0%, variant-18 1.4%, variant-14 1.4%, variant-12 1.4%, variant-19 1.4%
class canonical none
admiration 70.1% 2.9%
amusement 69.0% 2.8%
anger 68.3% 3.8%
annoyance 69.5% 3.1%
approval 70.6% 3.5%
caring 72.0% 1.9%
confusion 71.9% 2.6%
curiosity 68.9% 2.6%
desire 71.0% 2.6%
disappointment 69.5% 3.6%
disapproval 71.6% 2.6%
disgust 72.1% 3.0%
embarrassment 71.9% 2.7%
excitement 69.0% 4.2%
fear 69.5% 3.9%
gratitude 69.1% 2.6%
grief 68.7% 4.5%
joy 70.2% 2.9%
love 70.2% 3.5%
nervousness 66.2% 4.6%
neutral 70.0% 3.0%
optimism 71.0% 3.4%
pride 72.8% 3.7%
realization 70.1% 3.8%
relief 67.5% 4.8%
remorse 66.9% 3.6%
sadness 72.7% 2.5%
surprise 72.2% 2.7%

1. Shortcuts

Phrase statistics and models use a label-stratified sample of 46,067 train rows; models are scored on the full test file (5,408 rows).

Text length by label class (main text, characters)

split class rows p10 median p90 mean
train admiration 3590 20 58 117 63
train amusement 2102 22 63 113 65
train anger 1357 18 58 121 64
train annoyance 2082 28 73 126 75
train approval 2566 26 71 121 72
train caring 910 25 72 122 73
train confusion 1180 31 76 121 75
train curiosity 1872 27 66 120 70
train desire 542 31 69 118 70
train disappointment 1061 33 74 124 76
train disapproval 1818 30 73 124 75
train disgust 675 25 67 118 69
train embarrassment 260 33 74 122 76
train excitement 693 20 55 111 60
train fear 570 24 69 121 72
train gratitude 2279 22 59 114 64
train grief 67 25 57 114 64
train joy 1188 23 64 116 67
train love 1808 20 60 115 64
train nervousness 130 25 77 124 75
train neutral 14417 20 64 121 67
train optimism 1293 36 80 124 80
train pride 81 26 54 111 61
train realization 865 36 77 126 79
train relief 126 27 64 109 66
train remorse 474 29 69 122 72
train sadness 1129 22 69 119 69
train surprise 932 25 65 118 69
test admiration 422 19 60 110 62
test amusement 227 24 61 111 64
test anger 166 18 63 124 67
test annoyance 261 26 73 126 75
test approval 281 25 70 123 72
test caring 110 22 68 125 71
test confusion 128 32 76 118 75
test curiosity 225 22 60 112 63
test desire 69 38 78 114 73
test disappointment 109 32 71 127 76
test disapproval 232 28 72 122 74
test disgust 96 26 67 123 72
test embarrassment 30 31 88 128 79
test excitement 78 20 50 117 61
test fear 72 25 63 118 68
test gratitude 294 19 58 112 61
test grief 4 36 75 82 64
test joy 128 23 63 115 67
test love 198 20 60 112 62
test nervousness 17 35 71 126 76
test neutral 1682 19 64 119 66
test optimism 146 31 75 121 74
test pride 13 35 41 93 55
test realization 124 36 83 121 79
test relief 8 32 59 92 61
test remorse 51 27 56 104 61
test sadness 129 18 66 124 68
test surprise 108 22 61 113 64

By row kind (train): main-text length, length of the rest of the state, options

kind rows median chars mean chars median other-state chars median options
emotion 46067 66 69 61 28

Correct option: longest / shortest / position / key

For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):

split rows correct is longest correct is shortest chance (1/listed) mean relative position (0 first, 1 last; 0.5 expected) position fifths

Correct key and position by option count

split options rows mean options top correct keys most common position (0-based)
train 11-30 46067 28.0 neutral 31.3%, admiration 7.8%, approval 5.6%, gratitude 4.9%, amusement 4.6% 2 (4.2%)
test 11-30 5408 28.0 neutral 31.1%, admiration 7.8%, gratitude 5.4%, approval 5.2%, annoyance 4.8% 1 (4.5%)

Option count by label class (train)

class rows min median mean max
admiration 3590 28 28 28.0 28
amusement 2102 28 28 28.0 28
anger 1357 28 28 28.0 28
annoyance 2082 28 28 28.0 28
approval 2566 28 28 28.0 28
caring 910 28 28 28.0 28
confusion 1180 28 28 28.0 28
curiosity 1872 28 28 28.0 28
desire 542 28 28 28.0 28
disappointment 1061 28 28 28.0 28
disapproval 1818 28 28 28.0 28
disgust 675 28 28 28.0 28
embarrassment 260 28 28 28.0 28
excitement 693 28 28 28.0 28
fear 570 28 28 28.0 28
gratitude 2279 28 28 28.0 28
grief 67 28 28 28.0 28
joy 1188 28 28 28.0 28
love 1808 28 28 28.0 28
nervousness 130 28 28 28.0 28
neutral 14417 28 28 28.0 28
optimism 1293 28 28 28.0 28
pride 81 28 28 28.0 28
realization 865 28 28 28.0 28
relief 126 28 28 28.0 28
remorse 474 28 28 28.0 28
sadness 1129 28 28 28.0 28
surprise 932 28 28 28.0 28

Source fields by label class (train)

Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.

Overall majority: 31.3%.

source field values purity top values → classes
target_kind 2 31.3% hard: neutral 35.6%, admiration 7.4%; soft: neutral 9.9%, admiration 9.6%
instructions 22 31.3% canonical: neutral 31.2%, admiration 7.8%; none: neutral 30.9%, admiration 7.4%; variant-18: neutral 31.7%, admiration 7.5%; variant-14: neutral 31.6%, admiration 9.5%; variant-12: neutral 32.5%, admiration 7.8%; variant-19: neutral 29.4%, admiration 8.3%
split 2 31.3% train: neutral 31.3%, admiration 7.8%; validation: neutral 31.6%, admiration 7.5%
copies 9 31.3% 1: neutral 31.3%, admiration 7.8%; 2: neutral 28.8%, gratitude 15.8%; 3: neutral 27.5%, gratitude 15.0%; 4: gratitude 27.3%, neutral 27.3%; 5: neutral 44.4%, surprise 11.1%; 6: neutral 50.0%, joy 25.0%

Formatting by label class (main text, share of rows)

feature lowest classes highest classes
ends with ? pride 0%, relief 1%, joy 1% curiosity 51%, confusion 32%, neutral 5% gap
ends with . curiosity 25%, amusement 34%, excitement 35% disapproval 62%, realization 58%, approval 57% gap
ends with ! grief 1%, remorse 3%, confusion 3% excitement 29%, gratitude 23%, joy 15% gap
no end punctuation curiosity 18%, confusion 25%, caring 27% amusement 54%, grief 43%, sadness 41% gap
starts lowercase relief 4%, desire 4%, realization 4% grief 9%, amusement 9%, nervousness 8%
all lowercase pride 1%, relief 2%, desire 2% embarrassment 7%, nervousness 6%, grief 6%
has a digit anger 4%, caring 5%, gratitude 5% disappointment 12%, surprise 12%, realization 12%
has newline admiration 0%, amusement 0%, anger 0% surprise 0%, sadness 0%, remorse 0%
has quotes relief 1%, grief 1%, optimism 2% annoyance 6%, amusement 5%, embarrassment 5%
has markup (HTML/markdown) caring 0%, desire 0%, disgust 0% anger 1%, relief 1%, joy 1%
has URL admiration 0%, amusement 0%, anger 0% surprise 0%, sadness 0%, remorse 0%
non-ASCII neutral 12%, anger 12%, desire 13% relief 21%, disapproval 21%, remorse 20%
non-Latin script admiration 0%, anger 0%, annoyance 0% embarrassment 0%, disapproval 0%, surprise 0%
emoji pride 0%, relief 0%, anger 1% grief 4%, love 4%, sadness 4%
ALL-CAPS word (4+) embarrassment 8%, nervousness 9%, gratitude 10% love 23%, anger 23%, neutral 21%
contains ' - ' or — grief 0%, pride 0%, sadness 0% nervousness 2%, gratitude 1%, embarrassment 1%

Same, by row kind

feature emotion
ends with ? 6%
ends with . 48%
ends with ! 9%
no end punctuation 34%
starts lowercase 6%
all lowercase 4%
has a digit 9%
has newline 0%
has quotes 4%
has markup (HTML/markdown) 0%
has URL 0%
non-ASCII 15%
non-Latin script 0%
emoji 2%
ALL-CAPS word (4+) 17%
contains ' - ' or — 1%

Over-represented words and phrases per label class (main text)

Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.

Words, admiration: great z=42 14.1% vs 0.8%; good z=32 14.1% vs 2.9%; awesome z=29 6.7% vs 0.2%; amazing z=28 6.0% vs 0.3%; nice z=24 5.4% vs 0.6%; beautiful z=22 4.1% vs 0.1%; best z=20 5.3% vs 1.0%; pretty z=20 5.4% vs 1.1%; cute z=18 2.6% vs 0.2%; appreciate z=16 2.2% vs 0.2%; looks z=14 3.7% vs 1.0%; cool z=13 2.6% vs 0.6%; fantastic z=12 1.2% vs 0.1%; job z=12 2.0% vs 0.4%; wow z=11 2.9% vs 1.0%; is z=11 23.0% vs 17.3%; wonderful z=11 1.0% vs 0.1%; incredible z=10 0.8% vs 0.0%; excellent z=10 0.8% vs 0.0%; interesting z=10 1.8% vs 0.5%

Words, amusement: lol z=64 42.6% vs 0.9%; haha z=33 11.0% vs 0.3%; funny z=30 9.1% vs 0.3%; lmao z=23 5.6% vs 0.3%; fun z=22 6.2% vs 0.6%; hilarious z=18 3.2% vs 0.1%; joke z=17 3.6% vs 0.3%; laugh z=17 2.9% vs 0.2%; hahaha z=16 2.8% vs 0.1%; laughed z=14 2.1% vs 0.0%; laughing z=12 1.6% vs 0.1%; ha z=8 1.0% vs 0.1%; loud z=7 0.8% vs 0.1%; kidding z=7 0.6% vs 0.0%; jokes z=7 0.7% vs 0.1%; joking z=7 0.5% vs 0.0%; hahahaha z=7 0.5% vs 0.0%; funnier z=7 0.5% vs 0.0%; entertaining z=6 0.5% vs 0.0%; actually z=6 3.2% vs 1.5%

Words, anger: fuck z=36 15.0% vs 0.5%; hate z=27 10.2% vs 0.6%; fucking z=26 8.8% vs 0.5%; angry z=15 2.8% vs 0.1%; stupid z=14 3.8% vs 0.5%; dare z=14 2.2% vs 0.1%; shut z=12 2.1% vs 0.1%; hell z=12 3.2% vs 0.5%; shit z=9 3.2% vs 0.7%; asshole z=9 1.0% vs 0.0%; fucked z=9 1.1% vs 0.1%; wtf z=9 1.1% vs 0.1%; bitch z=9 1.1% vs 0.1%; idiot z=9 1.4% vs 0.2%; bastard z=9 0.9% vs 0.0%; bullshit z=8 1.0% vs 0.1%; stop z=8 2.7% vs 0.7%; kill z=8 1.5% vs 0.3%; suck z=7 1.0% vs 0.1%; off z=7 3.2% vs 1.3%

Words, annoyance: stupid z=16 3.8% vs 0.4%; annoying z=13 1.9% vs 0.0%; fucking z=13 3.7% vs 0.6%; shit z=12 3.6% vs 0.7%; damn z=11 3.7% vs 0.9%; dumb z=10 1.7% vs 0.2%; idiot z=10 1.4% vs 0.1%; fuck z=9 3.1% vs 0.8%; sucks z=9 1.7% vs 0.3%; annoyed z=8 0.8% vs 0.0%; idiots z=8 0.9% vs 0.1%; weird z=8 2.2% vs 0.5%; stop z=8 2.5% vs 0.7%; pissed z=8 0.9% vs 0.1%; hate z=7 2.6% vs 0.8%; hell z=7 1.8% vs 0.5%; ridiculous z=7 0.9% vs 0.1%; asshole z=7 0.6% vs 0.0%; fool z=6 0.5% vs 0.0%; frustrating z=6 0.5% vs 0.0%

Words, approval: agree z=26 6.4% vs 0.2%; yes z=16 5.1% vs 0.9%; yeah z=13 5.7% vs 1.7%; right z=12 5.9% vs 1.9%; agreed z=12 1.4% vs 0.1%; true z=11 2.7% vs 0.6%; correct z=9 1.2% vs 0.1%; ok z=9 2.4% vs 0.6%; sure z=9 3.8% vs 1.4%; exactly z=8 1.8% vs 0.4%; yep z=7 0.9% vs 0.2%; definitely z=7 1.8% vs 0.6%; fine z=7 1.1% vs 0.3%; fair z=6 0.9% vs 0.2%; but z=6 12.0% vs 8.2%; free z=5 1.1% vs 0.4%; prefer z=5 0.5% vs 0.1%; it's z=5 6.5% vs 4.2%; with z=5 9.9% vs 6.9%; absolutely z=5 1.1% vs 0.4%

Words, caring: worry z=17 4.8% vs 0.2%; you z=16 47.3% vs 19.1%; yourself z=16 6.2% vs 0.5%; your z=14 19.1% vs 5.5%; stay z=14 4.5% vs 0.3%; safe z=13 3.2% vs 0.2%; luck z=13 4.8% vs 0.5%; help z=13 5.8% vs 0.8%; careful z=11 2.1% vs 0.0%; concerned z=11 1.8% vs 0.0%; care z=10 3.7% vs 0.5%; bless z=10 1.8% vs 0.1%; better z=9 5.9% vs 1.5%; get z=9 11.0% vs 4.0%; take z=9 4.8% vs 1.1%; need z=8 5.2% vs 1.3%; keep z=8 4.0% vs 0.9%; strong z=8 1.9% vs 0.2%; please z=7 3.2% vs 0.7%; praying z=7 0.9% vs 0.0%

Words, confusion: sure z=19 10.2% vs 1.3%; confused z=19 5.9% vs 0.0%; why z=17 11.5% vs 2.0%; or z=15 12.5% vs 3.0%; what z=14 17.6% vs 5.6%; understand z=14 4.6% vs 0.5%; don't z=12 10.9% vs 3.1%; know z=12 10.6% vs 3.0%; idea z=10 3.5% vs 0.5%; idk z=10 2.1% vs 0.2%; not z=10 17.7% vs 7.9%; how z=8 9.4% vs 3.7%; confusing z=8 0.9% vs 0.0%; maybe z=8 3.9% vs 1.0%; confusion z=7 0.8% vs 0.0%; don z=7 4.4% vs 1.4%; doubt z=7 1.2% vs 0.1%; clue z=6 0.7% vs 0.0%; do z=6 8.6% vs 4.2%; i z=6 45.8% vs 31.7%

Words, curiosity: what z=25 20.9% vs 5.2%; curious z=21 6.5% vs 0.1%; why z=17 9.0% vs 2.0%; how z=17 12.1% vs 3.4%; did z=16 7.8% vs 1.8%; do z=14 11.4% vs 4.0%; you z=13 33.0% vs 19.1%; does z=12 4.3% vs 1.0%; what's z=11 2.0% vs 0.2%; where z=11 4.0% vs 1.0%; are z=10 13.2% vs 6.8%; curiosity z=9 1.0% vs 0.0%; interesting z=8 2.2% vs 0.5%; anyone z=7 2.4% vs 0.7%; wonder z=7 1.4% vs 0.3%; explain z=7 0.9% vs 0.1%; wondering z=6 0.8% vs 0.1%; question z=6 1.2% vs 0.3%; any z=6 3.2% vs 1.4%; or z=6 5.7% vs 3.1%

Words, desire: wish z=44 39.7% vs 0.5%; want z=17 14.8% vs 1.7%; i z=15 74.5% vs 31.5%; could z=14 11.6% vs 1.5%; wanted z=12 4.8% vs 0.3%; need z=9 7.2% vs 1.3%; wanna z=7 2.2% vs 0.2%; i'd z=6 3.7% vs 0.7%; would z=6 10.3% vs 3.8%; hope z=6 5.9% vs 1.8%; more z=6 8.1% vs 3.0%; dream z=5 1.3% vs 0.1%; hadn't z=5 0.7% vs 0.0%; pray z=5 0.9% vs 0.1%; had z=5 5.9% vs 2.2%; to z=5 37.8% vs 24.7%; desire z=5 0.6% vs 0.0%; please z=4 2.6% vs 0.7%; expecting z=4 0.9% vs 0.1%; blunder z=4 0.4% vs 0.0%

Words, disappointment: disappointed z=15 3.3% vs 0.1%; bad z=14 9.0% vs 1.6%; disappointing z=9 2.4% vs 0.0%; upset z=9 1.6% vs 0.1%; depressing z=9 1.1% vs 0.0%; lost z=9 2.5% vs 0.3%; miss z=9 2.3% vs 0.3%; unfortunately z=8 1.8% vs 0.2%; disappointment z=8 1.2% vs 0.0%; missed z=8 1.7% vs 0.2%; tried z=7 1.8% vs 0.3%; hurt z=7 1.7% vs 0.2%; game z=7 4.2% vs 1.3%; upsetting z=6 0.7% vs 0.0%; lose z=6 1.5% vs 0.3%; poor z=6 1.9% vs 0.4%; disaster z=6 0.6% vs 0.0%; sucks z=5 1.4% vs 0.3%; worse z=5 1.6% vs 0.4%; boring z=5 0.8% vs 0.1%

Words, disapproval: not z=22 24.8% vs 7.5%; no z=18 14.4% vs 3.9%; don't z=16 11.1% vs 2.9%; t z=15 12.9% vs 4.0%; don z=13 5.9% vs 1.3%; disagree z=12 1.6% vs 0.1%; nope z=10 1.3% vs 0.1%; doesn't z=9 3.3% vs 0.9%; wrong z=9 3.1% vs 0.8%; can't z=9 3.9% vs 1.2%; nah z=8 1.5% vs 0.2%; isn't z=8 2.6% vs 0.7%; unpopular z=8 0.7% vs 0.0%; think z=8 7.0% vs 3.1%; dont z=8 2.1% vs 0.5%; doesn z=7 1.5% vs 0.3%; bad z=6 4.0% vs 1.7%; refuse z=6 0.5% vs 0.0%; illegal z=6 0.7% vs 0.1%; opinion z=6 1.1% vs 0.2%

Words, disgust: awful z=23 9.6% vs 0.2%; disgusting z=22 13.5% vs 0.0%; worst z=22 9.2% vs 0.3%; weird z=17 7.6% vs 0.5%; worse z=15 5.8% vs 0.3%; nasty z=11 2.2% vs 0.0%; ugly z=10 2.5% vs 0.1%; gross z=10 2.1% vs 0.1%; creepy z=10 2.2% vs 0.1%; terrible z=8 2.7% vs 0.4%; ugh z=7 1.6% vs 0.1%; is z=7 28.4% vs 17.6%; horrible z=7 1.9% vs 0.3%; hideous z=6 0.7% vs 0.0%; dirty z=6 0.9% vs 0.1%; disgusted z=6 1.2% vs 0.0%; disgust z=6 0.6% vs 0.0%; ever z=6 3.4% vs 1.0%; filthy z=6 0.6% vs 0.0%; bad z=5 4.7% vs 1.7%

Words, embarrassment: shame z=18 11.5% vs 0.2%; awkward z=18 10.4% vs 0.1%; embarrassing z=16 11.5% vs 0.0%; ashamed z=12 5.0% vs 0.0%; embarrassed z=12 4.6% vs 0.0%; embarrassment z=10 5.0% vs 0.0%; oops z=8 1.9% vs 0.0%; weird z=7 5.4% vs 0.6%; uncomfortable z=7 2.3% vs 0.1%; cringy z=7 1.5% vs 0.0%; shamed z=5 0.8% vs 0.0%; temporarily z=5 0.8% vs 0.0%; millionaires z=5 0.8% vs 0.0%; forgot z=5 2.3% vs 0.3%; translate z=4 0.8% vs 0.0%; shaming z=4 0.8% vs 0.0%; cringey z=4 0.8% vs 0.0%; feel z=4 5.8% vs 1.6%; socially z=4 0.8% vs 0.0%; abused z=4 0.8% vs 0.0%

Words, excitement: excited z=25 11.1% vs 0.1%; wait z=16 6.6% vs 0.6%; wow z=15 7.8% vs 1.0%; interesting z=14 5.8% vs 0.5%; happy z=13 7.4% vs 1.1%; cake z=11 2.6% vs 0.1%; birthday z=11 2.3% vs 0.1%; exciting z=10 1.6% vs 0.0%; yay z=9 1.4% vs 0.0%; new z=9 4.8% vs 1.0%; omg z=7 2.5% vs 0.4%; interested z=7 1.6% vs 0.2%; excitement z=7 0.9% vs 0.0%; day z=7 4.0% vs 1.2%; year z=6 3.8% vs 1.1%; crazy z=6 2.0% vs 0.4%; amazing z=6 2.7% vs 0.7%; stoked z=6 0.6% vs 0.0%; cheers z=6 1.3% vs 0.2%; cakeday z=6 0.7% vs 0.0%

Words, fear: afraid z=23 10.7% vs 0.1%; scared z=22 11.1% vs 0.1%; terrible z=21 10.0% vs 0.3%; horrible z=19 7.5% vs 0.2%; scary z=18 6.8% vs 0.0%; terrifying z=15 5.3% vs 0.0%; fear z=15 4.4% vs 0.1%; dangerous z=11 2.6% vs 0.1%; worried z=11 3.0% vs 0.1%; creepy z=10 2.8% vs 0.1%; cringe z=10 2.8% vs 0.1%; horrifying z=10 1.9% vs 0.0%; horror z=9 1.8% vs 0.0%; nightmare z=8 1.6% vs 0.1%; scares z=8 2.3% vs 0.0%; me z=8 15.6% vs 6.3%; horrific z=7 1.1% vs 0.0%; terrified z=6 1.6% vs 0.0%; horribly z=6 0.9% vs 0.0%; frightening z=6 0.7% vs 0.0%

Words, gratitude: thanks z=63 47.3% vs 0.5%; thank z=56 38.2% vs 0.4%; for z=37 39.7% vs 11.1%; you z=29 45.1% vs 18.3%; sharing z=17 2.9% vs 0.1%; advice z=16 2.5% vs 0.1%; appreciate z=15 2.7% vs 0.2%; i'll z=15 3.5% vs 0.5%; much z=13 6.6% vs 2.1%; welcome z=13 2.0% vs 0.2%; info z=12 1.5% vs 0.1%; congrats z=11 1.5% vs 0.2%; glad z=11 3.4% vs 1.0%; helpful z=10 1.1% vs 0.1%; very z=9 4.4% vs 1.7%; ll z=9 2.2% vs 0.5%; response z=9 1.2% vs 0.2%; luck z=9 2.0% vs 0.5%; tip z=9 0.8% vs 0.1%; reply z=8 1.1% vs 0.1%

Words, grief: died z=15 22.4% vs 0.1%; rip z=10 10.4% vs 0.1%; loss z=9 10.4% vs 0.2%; dead z=8 10.4% vs 0.2%; death z=7 9.0% vs 0.2%; heart z=6 7.5% vs 0.2%; deaths z=6 3.0% vs 0.0%; peace z=6 4.5% vs 0.1%; condolences z=6 3.0% vs 0.0%; sooner z=5 3.0% vs 0.0%; sorry z=5 14.9% vs 1.7%; disease z=5 3.0% vs 0.0%; passed z=4 3.0% vs 0.1%; lost z=4 6.0% vs 0.4%; lands z=4 1.5% vs 0.0%; depression z=4 3.0% vs 0.1%; 2013 z=4 1.5% vs 0.0%; grandfather z=4 1.5% vs 0.0%; windshield z=4 1.5% vs 0.0%; chills z=4 1.5% vs 0.0%

Words, joy: happy z=40 21.2% vs 0.7%; glad z=35 17.3% vs 0.6%; enjoy z=26 8.7% vs 0.2%; fun z=18 7.0% vs 0.7%; enjoyed z=14 2.7% vs 0.0%; cheers z=12 2.2% vs 0.1%; enjoying z=11 1.6% vs 0.0%; cake z=11 1.9% vs 0.1%; day z=10 4.9% vs 1.1%; joy z=10 1.3% vs 0.0%; i'm z=10 10.4% vs 4.0%; happiness z=8 1.0% vs 0.1%; cakeday z=7 0.8% vs 0.0%; so z=7 13.6% vs 7.6%; m z=7 5.4% vs 2.1%; smile z=7 1.1% vs 0.1%; happily z=7 0.6% vs 0.0%; made z=6 3.3% vs 1.1%; birthday z=6 0.9% vs 0.1%; gladly z=6 0.5% vs 0.0%

Words, love: love z=81 69.0% vs 1.6%; i z=28 66.3% vs 30.6%; loved z=22 5.2% vs 0.2%; favorite z=17 4.0% vs 0.4%; loves z=17 3.0% vs 0.1%; like z=12 14.5% vs 7.0%; loving z=12 1.4% vs 0.1%; my z=10 14.7% vs 7.8%; i'd z=9 2.7% vs 0.7%; name z=7 19.9% vs 14.0%; lovely z=7 0.8% vs 0.1%; favourite z=6 0.7% vs 0.1%; liked z=6 0.9% vs 0.2%; d z=6 2.0% vs 0.7%; this z=6 18.4% vs 13.9%; much z=6 4.2% vs 2.2%; song z=5 0.8% vs 0.2%; sweet z=5 1.0% vs 0.3%; it z=5 21.7% vs 17.2%; how z=5 6.0% vs 3.7%

Words, nervousness: worried z=15 13.8% vs 0.1%; nervous z=14 11.5% vs 0.0%; anxiety z=12 9.2% vs 0.1%; anxious z=11 6.9% vs 0.0%; worrying z=9 4.6% vs 0.0%; worry z=7 5.4% vs 0.2%; career z=6 3.1% vs 0.1%; stressed z=5 1.5% vs 0.0%; scary z=5 3.1% vs 0.1%; paranoid z=5 1.5% vs 0.0%; breathing z=5 1.5% vs 0.0%; nightmares z=5 1.5% vs 0.0%; about z=5 16.2% vs 4.3%; that'll z=4 1.5% vs 0.0%; panic z=4 1.5% vs 0.0%; me z=4 20.0% vs 6.4%; edgy z=4 1.5% vs 0.0%; pick z=4 3.1% vs 0.2%; i'm z=4 14.6% vs 4.2%; standing z=4 1.5% vs 0.0%

Words, neutral: they z=15 8.2% vs 5.0%; name z=15 17.3% vs 12.8%; he z=9 6.9% vs 5.0%; in z=8 14.1% vs 12.1%; on z=7 8.8% vs 7.4%; only z=7 2.8% vs 1.9%; by z=7 2.6% vs 1.8%; the z=7 33.8% vs 32.3%; she z=6 3.6% vs 2.7%; his z=6 3.3% vs 2.5%; or z=6 3.8% vs 3.0%; their z=6 2.5% vs 1.9%; left z=6 0.8% vs 0.4%; 1 z=5 0.9% vs 0.5%; r z=5 0.9% vs 0.5%; from z=5 3.9% vs 3.2%; who z=5 2.8% vs 2.2%; then z=5 2.4% vs 1.8%; go z=5 2.5% vs 1.9%; there's z=5 0.7% vs 0.4%

Words, optimism: hope z=51 36.0% vs 0.8%; luck z=21 7.3% vs 0.4%; hopefully z=21 6.8% vs 0.1%; hoping z=19 5.1% vs 0.1%; will z=13 10.2% vs 2.5%; good z=10 10.5% vs 3.6%; soon z=10 2.2% vs 0.2%; better z=9 5.6% vs 1.5%; wish z=7 3.4% vs 0.8%; can z=7 10.1% vs 4.2%; next z=7 2.7% vs 0.6%; i z=7 50.7% vs 31.5%; best z=7 4.3% vs 1.3%; win z=7 1.9% vs 0.3%; future z=7 1.5% vs 0.2%; goes z=7 1.6% vs 0.3%; probably z=6 3.3% vs 1.0%; get z=6 8.8% vs 4.0%; hopeful z=6 0.5% vs 0.0%; optimistic z=6 0.5% vs 0.0%

Words, pride: proud z=23 38.3% vs 0.1%; pride z=9 6.2% vs 0.0%; jersey z=5 3.7% vs 0.1%; winner z=5 2.5% vs 0.0%; myself z=5 6.2% vs 0.5%; am z=4 9.9% vs 1.3%; butterfly z=4 1.2% vs 0.0%; cubs z=4 1.2% vs 0.0%; european z=4 1.2% vs 0.0%; honored z=4 1.2% vs 0.0%; leaking z=4 1.2% vs 0.0%; integrity z=4 1.2% vs 0.0%; congrats z=4 3.7% vs 0.2%; corn z=4 1.2% vs 0.0%; playthrough z=4 1.2% vs 0.0%; accomplishment z=4 1.2% vs 0.0%; senator z=4 1.2% vs 0.0%; athletic z=4 1.2% vs 0.0%; 2006 z=4 1.2% vs 0.0%; glory z=4 1.2% vs 0.0%

Words, realization: realize z=18 5.8% vs 0.2%; realized z=16 4.2% vs 0.0%; thought z=10 6.1% vs 1.2%; forgot z=10 2.7% vs 0.2%; was z=9 19.3% vs 8.2%; reminds z=8 1.7% vs 0.1%; noticed z=8 1.8% vs 0.2%; until z=8 3.1% vs 0.6%; realizing z=7 0.9% vs 0.0%; realised z=7 0.9% vs 0.0%; figured z=7 1.0% vs 0.1%; ago z=7 2.7% vs 0.5%; realise z=7 0.8% vs 0.0%; didn't z=6 4.3% vs 1.2%; realization z=6 0.6% vs 0.0%; i z=6 48.7% vs 31.7%; aware z=5 0.9% vs 0.1%; wrong z=5 2.9% vs 0.8%; that z=5 28.6% vs 17.9%; ve z=5 2.7% vs 0.8%

Words, relief: glad z=14 22.2% vs 1.0%; relief z=9 4.8% vs 0.0%; god z=8 7.1% vs 0.3%; relieved z=8 3.2% vs 0.0%; least z=8 10.3% vs 0.8%; finally z=6 4.8% vs 0.2%; relax z=6 2.4% vs 0.0%; whew z=5 1.6% vs 0.0%; feel z=5 9.5% vs 1.6%; cool z=5 6.3% vs 0.8%; solved z=5 1.6% vs 0.0%; thank z=5 11.1% vs 2.3%; thankfully z=5 1.6% vs 0.0%; mac z=5 1.6% vs 0.0%; safe z=4 3.2% vs 0.2%; could've z=4 1.6% vs 0.0%; goodness z=4 1.6% vs 0.0%; m z=4 9.5% vs 2.2%; at z=4 15.1% vs 4.7%; oof z=4 1.6% vs 0.1%

Words, remorse: sorry z=59 75.9% vs 0.9%; regret z=16 5.5% vs 0.0%; apologies z=12 3.4% vs 0.0%; m z=11 11.4% vs 2.1%; i'm z=10 15.2% vs 4.1%; apologize z=9 1.9% vs 0.0%; guilty z=8 1.7% vs 0.0%; meant z=8 3.0% vs 0.2%; guilt z=7 1.3% vs 0.0%; i z=7 54.4% vs 31.8%; loss z=7 2.1% vs 0.2%; im z=6 3.4% vs 0.5%; am z=6 5.1% vs 1.3%; apology z=5 0.6% vs 0.0%; misunderstanding z=5 0.6% vs 0.0%; happened z=5 2.5% vs 0.5%; sincerely z=5 0.6% vs 0.0%; didn z=5 2.7% vs 0.6%; misread z=5 0.6% vs 0.0%; should've z=5 0.6% vs 0.0%

Words, sadness: sad z=35 17.3% vs 0.2%; sorry z=23 12.8% vs 1.4%; sadly z=18 5.0% vs 0.0%; miss z=14 3.5% vs 0.3%; hurts z=13 2.4% vs 0.1%; painful z=13 2.3% vs 0.0%; poor z=13 3.5% vs 0.3%; feel z=12 6.9% vs 1.5%; crying z=12 2.2% vs 0.1%; pain z=11 2.3% vs 0.2%; bad z=11 6.8% vs 1.6%; cry z=11 1.9% vs 0.1%; loss z=9 1.8% vs 0.1%; lonely z=8 1.0% vs 0.0%; my z=8 15.1% vs 7.9%; hurt z=8 1.7% vs 0.2%; i'm z=8 9.2% vs 4.1%; lost z=7 1.9% vs 0.3%; hard z=7 3.2% vs 0.8%; sick z=7 1.3% vs 0.2%

Words, surprise: wow z=31 17.0% vs 0.8%; surprised z=30 14.2% vs 0.1%; wonder z=22 7.6% vs 0.2%; omg z=18 6.1% vs 0.3%; wondering z=14 3.2% vs 0.1%; shocked z=14 3.2% vs 0.0%; oh z=14 9.4% vs 1.9%; surprise z=14 2.8% vs 0.1%; believe z=13 5.2% vs 0.7%; surprising z=10 1.4% vs 0.0%; unexpected z=8 1.1% vs 0.0%; wondered z=8 1.1% vs 0.0%; strange z=8 1.3% vs 0.1%; was z=8 16.1% vs 8.2%; amazed z=8 0.9% vs 0.0%; unbelievable z=8 0.9% vs 0.0%; god z=7 2.0% vs 0.3%; shocking z=7 0.8% vs 0.0%; how z=7 8.5% vs 3.7%; holy z=7 1.5% vs 0.2%

2–4-word phrases, admiration: a great z=22 3.9% vs 0.2%; the best z=19 3.8% vs 0.5%; a good z=15 2.8% vs 0.5%; i appreciate z=13 1.4% vs 0.1%; this is z=13 6.0% vs 2.5%; is amazing z=12 1.2% vs 0.0%; is awesome z=12 1.3% vs 0.0%; is great z=12 1.1% vs 0.1%; what a z=11 2.1% vs 0.5%; is the best z=11 0.9% vs 0.1%; good job z=10 0.9% vs 0.0%; such a z=10 1.6% vs 0.4%; so good z=10 0.8% vs 0.1%; a beautiful z=10 0.8% vs 0.0%; pretty good z=9 0.7% vs 0.0%; is a great z=9 0.7% vs 0.0%; so cute z=9 0.6% vs 0.0%; an amazing z=9 0.6% vs 0.0%; a pretty z=8 0.7% vs 0.1%; a nice z=8 0.8% vs 0.1%

2–4-word phrases, amusement: lol i z=16 2.7% vs 0.1%; haha i z=12 1.4% vs 0.0%; me laugh z=11 1.3% vs 0.0%; a joke z=10 1.3% vs 0.1%; made me laugh z=10 1.0% vs 0.0%; i laughed z=10 1.2% vs 0.0%; it lol z=9 0.9% vs 0.0%; fun to z=8 0.7% vs 0.0%; so hard z=8 1.0% vs 0.1%; is hilarious z=8 0.8% vs 0.0%; name lol z=8 0.7% vs 0.0%; lol you z=8 0.7% vs 0.0%; made me z=8 1.4% vs 0.3%; out loud z=8 0.6% vs 0.0%; it's funny z=8 0.6% vs 0.0%; funny how z=8 0.6% vs 0.0%; at this z=7 1.2% vs 0.2%; lol the z=7 0.6% vs 0.0%; lol this z=7 0.7% vs 0.0%; the joke z=7 0.6% vs 0.1%

2–4-word phrases, anger: i hate z=22 6.0% vs 0.3%; the fuck z=19 4.2% vs 0.1%; how dare z=13 2.0% vs 0.0%; what the z=12 2.6% vs 0.3%; what the fuck z=11 1.4% vs 0.0%; the hell z=11 1.5% vs 0.1%; dare you z=11 1.3% vs 0.0%; fuck you z=10 1.4% vs 0.0%; fuck is z=10 1.2% vs 0.0%; fuck off z=10 1.3% vs 0.0%; shut up z=10 1.1% vs 0.0%; how dare you z=10 1.3% vs 0.0%; the fuck is z=10 1.1% vs 0.0%; fuck the z=9 1.2% vs 0.0%; the fucking z=9 0.9% vs 0.0%; a fucking z=9 1.0% vs 0.1%; hate the z=8 0.9% vs 0.0%; you fucking z=8 0.7% vs 0.0%; hate that z=8 0.7% vs 0.0%; what the hell z=8 0.7% vs 0.0%

2–4-word phrases, annoyance: an idiot z=9 0.9% vs 0.0%; a fucking z=7 0.7% vs 0.1%; that sucks z=7 0.5% vs 0.0%; the hell z=6 0.8% vs 0.1%; an asshole z=6 0.4% vs 0.0%; a weird z=6 0.6% vs 0.1%; i hate z=6 1.3% vs 0.4%; can't even z=6 0.4% vs 0.0%; stupid to z=6 0.3% vs 0.0%; what the hell z=5 0.4% vs 0.0%; shut up z=5 0.4% vs 0.1%; is stupid z=5 0.3% vs 0.0%; fuck that z=5 0.3% vs 0.0%; a fuck z=5 0.3% vs 0.0%; waste of z=5 0.3% vs 0.0%; makes no z=5 0.3% vs 0.0%; as hell z=5 0.5% vs 0.1%; tired of z=5 0.3% vs 0.0%; a god z=5 0.3% vs 0.0%; a stupid z=5 0.3% vs 0.0%

2–4-word phrases, approval: i agree z=18 3.8% vs 0.1%; agree with z=14 1.8% vs 0.1%; i agree with z=10 1.0% vs 0.0%; yeah i z=8 1.5% vs 0.3%; agree with you z=8 0.6% vs 0.0%; you're right z=8 0.7% vs 0.1%; pretty sure z=7 0.7% vs 0.1%; you re right z=7 0.6% vs 0.0%; right i z=7 0.7% vs 0.1%; re right z=7 0.6% vs 0.0%; agree that z=7 0.5% vs 0.0%; with you z=7 1.1% vs 0.2%; i think z=6 3.3% vs 1.5%; yes i z=6 0.8% vs 0.2%; agree but z=6 0.4% vs 0.0%; for sure z=6 0.7% vs 0.1%; agree i z=6 0.5% vs 0.0%; yes that z=6 0.4% vs 0.0%; i believe z=6 0.6% vs 0.1%; i completely z=6 0.4% vs 0.0%

2–4-word phrases, caring: don't worry z=12 2.5% vs 0.0%; for you z=12 4.7% vs 0.6%; good luck z=11 3.6% vs 0.3%; you need z=11 2.9% vs 0.2%; take care z=10 1.6% vs 0.0%; be careful z=10 1.5% vs 0.0%; stay strong z=9 1.5% vs 0.0%; stay safe z=9 1.8% vs 0.0%; name bless z=8 1.2% vs 0.0%; care of z=8 1.3% vs 0.1%; keep your z=8 1.1% vs 0.0%; you have z=8 3.8% vs 0.8%; if you z=8 4.9% vs 1.2%; t worry z=8 1.0% vs 0.0%; don t worry z=8 1.0% vs 0.0%; don't be z=7 1.0% vs 0.0%; take care of z=7 0.9% vs 0.0%; you feel z=7 1.4% vs 0.1%; get some z=7 1.1% vs 0.1%; no worries z=7 0.9% vs 0.0%

2–4-word phrases, confusion: not sure z=23 7.8% vs 0.2%; i don't z=15 7.7% vs 1.0%; don't know z=14 4.1% vs 0.3%; i don't know z=13 3.1% vs 0.2%; no idea z=13 2.9% vs 0.1%; sure what z=12 2.3% vs 0.0%; i'm not sure z=11 1.9% vs 0.0%; don t know z=11 2.1% vs 0.1%; not sure what z=11 1.7% vs 0.0%; have no idea z=11 1.8% vs 0.1%; i don t know z=10 1.7% vs 0.1%; i have no idea z=10 1.6% vs 0.0%; i have no z=10 1.8% vs 0.1%; t know z=10 2.4% vs 0.2%; m not sure z=10 1.4% vs 0.0%; i m not sure z=10 1.4% vs 0.0%; sure if z=10 1.4% vs 0.0%; i don z=9 3.5% vs 0.5%; i don t z=9 3.5% vs 0.5%; know what z=9 2.6% vs 0.3%

2–4-word phrases, curiosity: do you z=22 6.5% vs 0.5%; are you z=18 4.9% vs 0.5%; did you z=16 3.2% vs 0.2%; is it z=14 2.7% vs 0.2%; is this z=13 2.7% vs 0.3%; can you z=11 1.7% vs 0.1%; what are z=11 1.4% vs 0.1%; is that z=10 2.8% vs 0.5%; do you have z=10 1.2% vs 0.1%; just curious z=10 1.4% vs 0.0%; have you z=10 1.3% vs 0.1%; how do z=9 1.1% vs 0.1%; you think z=9 1.8% vs 0.3%; how did z=9 1.0% vs 0.0%; what do z=9 1.1% vs 0.1%; why is z=9 1.1% vs 0.1%; what are you z=9 0.9% vs 0.0%; how is z=9 0.9% vs 0.0%; is there z=9 1.2% vs 0.1%; do you think z=9 1.0% vs 0.1%

2–4-word phrases, desire: i wish z=36 29.9% vs 0.2%; wish i z=24 13.1% vs 0.1%; i wish i z=22 10.9% vs 0.1%; wish i could z=17 6.6% vs 0.1%; i want z=17 8.5% vs 0.4%; i could z=16 7.7% vs 0.3%; i wish i could z=16 5.5% vs 0.1%; i want to z=12 4.2% vs 0.2%; want to z=10 6.6% vs 0.8%; wish we z=10 2.6% vs 0.0%; i just want z=9 2.2% vs 0.1%; wish i had z=9 2.0% vs 0.0%; i wanna z=9 2.2% vs 0.1%; wish the z=9 2.4% vs 0.0%; i need z=9 3.0% vs 0.2%; just want z=9 2.2% vs 0.1%; wish they z=8 1.7% vs 0.0%; wish it z=8 1.7% vs 0.0%; wish name z=8 1.3% vs 0.0%; i wish name z=8 1.3% vs 0.0%

2–4-word phrases, disappointment: my bad z=8 1.4% vs 0.1%; i miss z=8 1.5% vs 0.1%; i tried z=7 1.0% vs 0.1%; too bad z=7 1.1% vs 0.1%; disappointed in z=7 0.8% vs 0.0%; bad for z=6 1.0% vs 0.1%; missed it z=6 0.6% vs 0.0%; i missed z=6 0.8% vs 0.1%; so bad z=6 0.9% vs 0.1%; i tried to z=6 0.6% vs 0.0%; a disaster z=6 0.5% vs 0.0%; i lost z=5 0.6% vs 0.0%; oh no z=5 0.8% vs 0.1%; it didn z=5 0.5% vs 0.0%; it didn t z=5 0.5% vs 0.0%; tried to z=5 0.8% vs 0.1%; but no z=5 0.6% vs 0.0%; so disappointed z=5 0.4% vs 0.0%; filled with z=5 0.5% vs 0.0%; this game z=5 0.9% vs 0.2%

2–4-word phrases, disapproval: i don't z=14 5.6% vs 1.0%; don t z=13 5.9% vs 1.3%; don't think z=12 2.3% vs 0.2%; i don z=12 3.2% vs 0.5%; i don t z=12 3.2% vs 0.5%; that's not z=12 1.8% vs 0.1%; is not z=12 2.7% vs 0.4%; i don't think z=11 2.0% vs 0.2%; not a z=10 2.5% vs 0.4%; it's not z=10 2.3% vs 0.4%; i dont z=9 1.5% vs 0.2%; don t think z=8 1.1% vs 0.1%; no it z=8 0.8% vs 0.0%; t think z=8 1.2% vs 0.1%; don t want z=8 0.7% vs 0.0%; disagree with z=8 0.7% vs 0.0%; don't like z=8 0.9% vs 0.1%; i don t think z=7 0.9% vs 0.1%; not true z=7 0.6% vs 0.0%; no way z=7 0.9% vs 0.1%

2–4-word phrases, disgust: the worst z=18 6.2% vs 0.2%; is the worst z=10 1.6% vs 0.0%; is disgusting z=9 1.9% vs 0.0%; is awful z=8 1.2% vs 0.0%; worse than z=8 1.5% vs 0.1%; i've ever z=8 1.5% vs 0.1%; an awful z=7 1.0% vs 0.0%; even worse z=7 1.0% vs 0.0%; a weird z=7 1.2% vs 0.1%; i've ever seen z=7 1.0% vs 0.0%; weird to z=7 0.9% vs 0.0%; this is the worst z=7 0.7% vs 0.0%; it's weird z=6 0.7% vs 0.0%; ever seen z=6 1.2% vs 0.1%; i hate z=6 2.2% vs 0.4%; disgusting and z=6 0.7% vs 0.0%; that's awful z=6 0.6% vs 0.0%; be worse z=6 0.6% vs 0.0%; absolute worst z=6 0.6% vs 0.0%; is weird z=6 0.6% vs 0.0%

2–4-word phrases, embarrassment: a shame z=8 2.7% vs 0.1%; be ashamed z=7 1.5% vs 0.0%; my bad z=6 2.3% vs 0.1%; s a shame z=6 1.2% vs 0.0%; i wear z=6 1.2% vs 0.0%; ashamed of z=6 1.2% vs 0.0%; ashamed to z=6 1.2% vs 0.0%; of shame z=6 1.2% vs 0.0%; it s a shame z=6 1.2% vs 0.0%; an embarrassment z=5 1.5% vs 0.0%; my body z=5 1.2% vs 0.0%; sorry i z=5 2.3% vs 0.2%; second hand z=5 0.8% vs 0.0%; even know that z=5 0.8% vs 0.0%; it was weird z=5 0.8% vs 0.0%; body and z=5 0.8% vs 0.0%; was weird z=5 0.8% vs 0.0%; to much z=5 0.8% vs 0.0%; an easy z=5 0.8% vs 0.0%; be ashamed of z=5 0.8% vs 0.0%

2–4-word phrases, excitement: wait to z=14 3.2% vs 0.0%; can't wait z=13 3.0% vs 0.1%; t wait z=12 2.6% vs 0.0%; can t wait z=12 2.6% vs 0.0%; excited to z=12 2.7% vs 0.0%; wait for z=11 2.3% vs 0.1%; excited for z=11 3.3% vs 0.0%; happy new z=10 2.0% vs 0.1%; cake day z=10 2.0% vs 0.1%; so excited z=10 1.9% vs 0.0%; happy new year z=10 1.7% vs 0.1%; excited to see z=9 1.6% vs 0.0%; happy birthday z=9 1.4% vs 0.0%; happy cake day z=9 1.6% vs 0.1%; happy cake z=9 1.6% vs 0.1%; can t wait to z=9 1.4% vs 0.0%; t wait to z=9 1.4% vs 0.0%; wait to see z=9 1.3% vs 0.0%; i can't wait z=9 1.3% vs 0.0%; new year z=9 1.7% vs 0.1%

2–4-word phrases, fear: a terrible z=13 3.3% vs 0.1%; afraid of z=13 3.3% vs 0.0%; afraid to z=11 2.3% vs 0.0%; scared of z=10 2.5% vs 0.0%; i'm afraid z=10 2.6% vs 0.0%; a horrible z=9 1.9% vs 0.1%; scared to z=8 2.8% vs 0.0%; is terrible z=8 1.4% vs 0.0%; worried about z=8 1.6% vs 0.1%; fear of z=8 1.2% vs 0.0%; i'm scared z=7 1.8% vs 0.0%; terrible idea z=7 1.2% vs 0.0%; a terrible idea z=7 1.2% vs 0.0%; was afraid z=7 1.2% vs 0.0%; out of me z=7 0.9% vs 0.0%; be afraid z=6 0.9% vs 0.0%; afraid i z=6 1.1% vs 0.0%; is terrifying z=6 1.1% vs 0.0%; is scary z=6 0.9% vs 0.0%; more worried about z=6 0.7% vs 0.0%

2–4-word phrases, gratitude: thank you z=50 35.2% vs 0.3%; thanks for z=37 18.9% vs 0.2%; you for z=30 10.0% vs 0.2%; for the z=29 12.9% vs 1.4%; thank you for z=27 10.0% vs 0.1%; thanks for the z=24 8.9% vs 0.1%; you i z=18 3.6% vs 0.2%; for your z=17 3.7% vs 0.3%; thanks i z=17 3.1% vs 0.1%; you so z=16 2.9% vs 0.1%; thank you i z=16 3.5% vs 0.0%; for sharing z=15 2.7% vs 0.1%; you so much z=15 2.7% vs 0.0%; thank you so z=14 2.9% vs 0.0%; thank you so much z=14 2.7% vs 0.0%; so much z=13 3.7% vs 0.6%; thanks for sharing z=12 1.6% vs 0.0%; you for the z=12 2.2% vs 0.0%; you for your z=12 1.8% vs 0.0%; i will z=12 2.5% vs 0.3%

2–4-word phrases, grief: your loss z=8 7.5% vs 0.1%; sorry about your z=7 4.5% vs 0.0%; sorry for z=7 9.0% vs 0.2%; sorry about z=6 4.5% vs 0.0%; passed away z=6 3.0% vs 0.0%; my condolences z=6 3.0% vs 0.0%; your friend z=6 4.5% vs 0.1%; for your loss z=5 4.5% vs 0.1%; sorry for your loss z=5 4.5% vs 0.1%; he died z=5 3.0% vs 0.0%; very sorry z=5 3.0% vs 0.0%; my only z=5 3.0% vs 0.0%; sorry for your z=5 4.5% vs 0.1%; this may z=5 3.0% vs 0.0%; about your z=5 4.5% vs 0.1%; to death z=5 3.0% vs 0.0%; two years z=5 3.0% vs 0.0%; after he z=5 3.0% vs 0.0%; to die z=4 3.0% vs 0.1%; he could z=4 3.0% vs 0.1%

2–4-word phrases, joy: glad you z=17 4.0% vs 0.1%; so happy z=16 3.6% vs 0.1%; glad to z=15 3.1% vs 0.1%; happy to z=14 2.9% vs 0.1%; i'm glad z=14 3.1% vs 0.1%; happy for z=13 2.5% vs 0.1%; so glad z=13 2.3% vs 0.1%; glad i z=12 2.1% vs 0.1%; i enjoy z=11 1.9% vs 0.0%; cake day z=11 1.7% vs 0.1%; enjoy it z=11 1.6% vs 0.0%; happy cake day z=11 1.6% vs 0.0%; happy cake z=11 1.6% vs 0.0%; me happy z=10 1.5% vs 0.0%; glad to see z=10 1.5% vs 0.0%; i enjoyed z=10 1.7% vs 0.0%; be happy z=10 1.5% vs 0.1%; enjoy the z=10 1.3% vs 0.0%; happy for you z=10 1.3% vs 0.0%; m glad z=10 1.4% vs 0.1%

2–4-word phrases, love: i love z=54 36.9% vs 0.5%; love it z=26 8.0% vs 0.2%; i like z=25 8.2% vs 0.5%; love the z=23 6.2% vs 0.2%; love this z=22 6.1% vs 0.1%; love to z=21 5.1% vs 0.1%; i love it z=19 4.6% vs 0.1%; would love z=17 3.4% vs 0.1%; i love this z=17 4.1% vs 0.0%; i love the z=17 3.4% vs 0.1%; love that z=16 3.4% vs 0.0%; love you z=16 3.3% vs 0.0%; love how z=16 3.5% vs 0.0%; love name z=15 2.8% vs 0.1%; my favorite z=15 3.3% vs 0.2%; i love how z=14 2.9% vs 0.0%; would love to z=14 2.2% vs 0.1%; i'd love z=14 2.2% vs 0.0%; i loved z=14 2.2% vs 0.1%; i would love z=13 2.0% vs 0.0%

2–4-word phrases, nervousness: i'm worried z=7 3.1% vs 0.0%; worried that z=7 3.1% vs 0.0%; worried about z=7 3.8% vs 0.1%; worried about the z=6 2.3% vs 0.0%; i'm worried that z=6 2.3% vs 0.0%; worried for z=6 3.1% vs 0.0%; to pick z=6 2.3% vs 0.0%; very worried z=5 1.5% vs 0.0%; get older z=5 1.5% vs 0.0%; so don't z=5 1.5% vs 0.0%; be worried z=5 1.5% vs 0.0%; so worried z=5 1.5% vs 0.0%; i can feel z=5 1.5% vs 0.0%; more worried about z=5 1.5% vs 0.0%; i have this z=5 1.5% vs 0.0%; more worried z=5 1.5% vs 0.0%; me off z=5 1.5% vs 0.0%; worrying about z=5 1.5% vs 0.0%; can feel z=4 1.5% vs 0.0%; makes me z=4 4.6% vs 0.5%

2–4-word phrases, neutral: name and z=8 1.4% vs 0.8%; in the z=7 3.7% vs 3.0%; and name z=7 1.2% vs 0.7%; name is z=6 1.8% vs 1.3%; name name z=6 0.6% vs 0.2%; they are z=6 1.0% vs 0.6%; name and name z=6 0.8% vs 0.5%; if they z=6 0.6% vs 0.3%; with name z=6 0.6% vs 0.3%; because they z=6 0.4% vs 0.2%; they don't z=6 0.3% vs 0.1%; on the z=5 1.9% vs 1.5%; name was z=5 0.7% vs 0.4%; name got z=5 0.2% vs 0.0%; he was z=5 0.9% vs 0.6%; looks like z=5 0.8% vs 0.5%; they can z=5 0.3% vs 0.1%; because she z=5 0.2% vs 0.0%; does not z=5 0.2% vs 0.1%; on name z=5 0.2% vs 0.1%

2–4-word phrases, optimism: i hope z=36 18.7% vs 0.4%; hope you z=22 7.1% vs 0.2%; good luck z=18 5.5% vs 0.3%; i hope you z=18 4.7% vs 0.1%; really hope z=12 2.2% vs 0.0%; hope it z=12 2.1% vs 0.0%; hope that z=12 2.1% vs 0.0%; i really hope z=12 2.1% vs 0.0%; hope he z=12 2.0% vs 0.0%; hope this z=11 1.9% vs 0.0%; hope the z=11 1.8% vs 0.0%; hope they z=11 1.7% vs 0.0%; hope she z=9 1.3% vs 0.0%; hope we z=9 1.2% vs 0.0%; hope you're z=9 1.2% vs 0.0%; was hoping z=9 1.2% vs 0.0%; hope for z=9 1.2% vs 0.0%; hope your z=9 1.2% vs 0.0%; i hope it z=9 1.2% vs 0.0%; just hope z=9 1.2% vs 0.0%

2–4-word phrases, pride: proud of z=16 21.0% vs 0.1%; so proud z=13 13.6% vs 0.0%; so proud of z=10 8.6% vs 0.0%; proud of you z=10 8.6% vs 0.1%; be proud z=8 4.9% vs 0.0%; of you z=7 8.6% vs 0.2%; i'm so proud z=7 4.9% vs 0.0%; so proud of you z=7 3.7% vs 0.0%; should be proud z=6 3.7% vs 0.0%; i'm so proud of z=6 3.7% vs 0.0%; of you and z=6 2.5% vs 0.0%; m proud z=6 2.5% vs 0.0%; i m proud z=6 2.5% vs 0.0%; proud to z=6 2.5% vs 0.0%; i'm proud z=5 2.5% vs 0.0%; be proud of z=5 2.5% vs 0.0%; this community z=5 2.5% vs 0.0%; i am just z=5 2.5% vs 0.0%; been so z=5 2.5% vs 0.0%; am just z=5 2.5% vs 0.0%

2–4-word phrases, realization: i realized z=10 2.2% vs 0.0%; realize that z=9 1.6% vs 0.1%; i forgot z=9 1.6% vs 0.1%; i thought z=8 3.7% vs 0.6%; i didn't z=8 2.9% vs 0.4%; didn't realize z=8 1.2% vs 0.0%; reminds me z=8 1.6% vs 0.1%; you realize z=8 1.0% vs 0.0%; i didn't realize z=8 1.0% vs 0.0%; just realized z=8 1.0% vs 0.0%; reminds me of z=8 1.5% vs 0.1%; me of z=8 1.7% vs 0.1%; i was z=7 6.4% vs 1.8%; i thought it z=7 1.5% vs 0.1%; it took z=7 0.9% vs 0.0%; realized that z=6 0.8% vs 0.0%; realize it z=6 0.7% vs 0.0%; thought it z=6 1.6% vs 0.2%; to realize z=6 0.7% vs 0.0%; but then i z=6 0.7% vs 0.0%

2–4-word phrases, relief: thank god z=10 5.6% vs 0.0%; m glad z=9 5.6% vs 0.1%; i m glad z=9 5.6% vs 0.1%; glad i z=8 5.6% vs 0.1%; at least z=7 10.3% vs 0.7%; m glad i z=7 3.2% vs 0.0%; i m glad i z=7 3.2% vs 0.0%; have to worry z=6 2.4% vs 0.0%; have to worry about z=6 2.4% vs 0.0%; glad i'm not z=6 2.4% vs 0.0%; makes me feel z=6 3.2% vs 0.1%; glad i'm z=6 2.4% vs 0.0%; someone i z=6 2.4% vs 0.0%; this makes me feel z=6 2.4% vs 0.0%; thank name z=6 3.2% vs 0.1%; so glad z=6 4.0% vs 0.1%; better about z=6 2.4% vs 0.0%; the only one z=6 4.0% vs 0.1%; don't have to z=6 2.4% vs 0.0%; to worry about z=6 2.4% vs 0.0%

2–4-word phrases, remorse: sorry i z=20 10.1% vs 0.1%; so sorry z=20 10.1% vs 0.1%; i'm sorry z=19 9.3% vs 0.1%; sorry for z=19 8.6% vs 0.1%; i m sorry z=16 6.1% vs 0.1%; m sorry z=16 6.1% vs 0.1%; sorry but z=15 5.5% vs 0.1%; sorry you z=14 5.1% vs 0.1%; i m so sorry z=13 4.0% vs 0.0%; m so sorry z=13 4.0% vs 0.0%; sorry to z=12 3.8% vs 0.1%; m so z=12 4.2% vs 0.1%; i m so z=12 4.2% vs 0.1%; sorry that z=11 3.0% vs 0.1%; i'm so sorry z=10 2.7% vs 0.0%; sorry for the z=10 2.5% vs 0.0%; sorry for your z=10 2.5% vs 0.1%; i m z=10 11.4% vs 2.1%; i regret z=9 2.3% vs 0.0%; i meant z=9 2.5% vs 0.1%

2–4-word phrases, sadness: so sad z=12 2.1% vs 0.0%; i feel z=12 4.3% vs 0.6%; bad for z=12 2.0% vs 0.1%; i miss z=11 2.0% vs 0.1%; sorry for z=11 2.2% vs 0.2%; feel bad z=11 1.7% vs 0.1%; i'm sorry z=10 2.2% vs 0.2%; feel bad for z=10 1.5% vs 0.0%; sad that z=10 1.7% vs 0.0%; sorry to z=9 1.4% vs 0.1%; so sorry z=9 1.9% vs 0.2%; me sad z=9 1.3% vs 0.0%; i m sorry z=9 1.6% vs 0.1%; m sorry z=9 1.6% vs 0.1%; sorry for your z=9 1.2% vs 0.1%; sorry that z=9 1.2% vs 0.1%; i feel bad z=8 1.0% vs 0.0%; for your loss z=8 1.1% vs 0.0%; sorry for your loss z=8 1.1% vs 0.0%; your loss z=8 1.1% vs 0.1%

2–4-word phrases, surprise: i wonder z=18 4.7% vs 0.1%; wow i z=13 2.9% vs 0.1%; wonder if z=13 2.5% vs 0.0%; can't believe z=12 2.1% vs 0.1%; wonder how z=12 2.0% vs 0.0%; i wonder if z=11 1.8% vs 0.0%; oh my z=11 2.4% vs 0.2%; t believe z=11 1.7% vs 0.0%; was wondering z=10 1.7% vs 0.0%; be surprised z=10 1.8% vs 0.1%; can t believe z=9 1.4% vs 0.0%; wonder why z=9 1.3% vs 0.0%; wow you z=9 1.3% vs 0.0%; i wonder how z=9 1.2% vs 0.0%; surprised if z=9 1.2% vs 0.0%; a surprise z=9 1.2% vs 0.0%; i'm surprised z=8 2.5% vs 0.0%; i was wondering z=8 1.5% vs 0.0%; wonder what z=8 1.1% vs 0.0%; i can't believe z=8 1.1% vs 0.0%

Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):

  • remorse: sorry 75.9% vs 0.9%
  • love: love 69.0% vs 1.6%
  • gratitude: thanks 47.3% vs 0.5%
  • amusement: lol 42.6% vs 0.9%
  • desire: wish 39.7% vs 0.5%
  • pride: proud 38.3% vs 0.1%
  • gratitude: thank 38.2% vs 0.4%
  • love: i love 36.9% vs 0.5%
  • optimism: hope 36.0% vs 0.8%
  • gratitude: thank you 35.2% vs 0.3%
  • desire: i wish 29.9% vs 0.2%
  • grief: died 22.4% vs 0.1%
  • relief: glad 22.2% vs 1.0%
  • joy: happy 21.2% vs 0.7%
  • pride: proud of 21.0% vs 0.1%
  • gratitude: thanks for 18.9% vs 0.2%
  • optimism: i hope 18.7% vs 0.4%
  • sadness: sad 17.3% vs 0.2%
  • joy: glad 17.3% vs 0.6%
  • surprise: wow 17.0% vs 0.8%
  • anger: fuck 15.0% vs 0.5%
  • grief: sorry 14.9% vs 1.7%
  • desire: want 14.8% vs 1.7%
  • surprise: surprised 14.2% vs 0.1%
  • admiration: good 14.1% vs 2.9%
  • admiration: great 14.1% vs 0.8%
  • nervousness: worried 13.8% vs 0.1%
  • pride: so proud 13.6% vs 0.0%
  • disgust: disgusting 13.5% vs 0.0%
  • desire: wish i 13.1% vs 0.1%
  • gratitude: for the 12.9% vs 1.4%
  • sadness: sorry 12.8% vs 1.4%
  • confusion: or 12.5% vs 3.0%
  • desire: could 11.6% vs 1.5%
  • embarrassment: shame 11.5% vs 0.2%
  • embarrassment: embarrassing 11.5% vs 0.0%
  • nervousness: nervous 11.5% vs 0.0%
  • confusion: why 11.5% vs 2.0%
  • remorse: m 11.4% vs 2.1%
  • remorse: i m 11.4% vs 2.1%

Shortcut models

Predicting the label class on test (5,408 rows). Chance 3.6%, majority class ('neutral') 31.1%; balanced chance 3.6%.

model (logistic regression, trained on the train sample) test accuracy balanced accuracy (mean recall)
bag of words, whole state (words and word pairs) 53.2% 32.0%
bag of words, main text only (comment) 41.7% 20.5%
surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) 31.1% 5.0%
surface features of the main text only 31.0% 4.8%

Strongest single surface features (logistic regression on one feature, balanced accuracy on test):

feature accuracy balanced accuracy
ends_? 31.3% 5.2%
count_? 31.0% 3.8%
count_! 30.8% 3.6%
chars(log) 31.1% 3.6%
words(log) 31.1% 3.6%
upper_ratio 31.1% 3.6%
digit_ratio 31.1% 3.6%
nonlatin 31.1% 3.6%
emoji 31.1% 3.6%
html_tag 31.1% 3.6%

Other state fields alone (predicting the label class on test from one field, without the main text):

field treated as accuracy balanced accuracy
source categorical, 1 values 31.1% 3.6%

No-meaning option picker: a logistic ranker scores each option from its position, length, key type, fixed-option identity and shape (commas, brackets, capitals), never reading the state or the option's words, and picks the top option per row.

  • Test accuracy 31.1% against uniform chance 3.6% (this includes the fixed options, whose share is a class prior).

2. Duplicates and split separation

Families shared between splits

splits shared families examples
train ∩ dev 0
train ∩ calibration 0
train ∩ test 0
dev ∩ calibration 0
dev ∩ test 0
calibration ∩ test 0
  • Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
  • Train rows identical in the whole prompt (state, options, instructions): 0.
  • Identical whole prompt, different answer: 0 groups (0 rows).
  • Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)

Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):

split rows examples
dev 0 (0.0%)
calibration 0 (0.0%)
test 0 (0.0%)

Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)

  • Train: 39 near-duplicate pairs; 73 rows (0.2%) sit in 35 clusters; largest cluster 4; excess rows (cluster size − 1) 38 (0.1%).
  • Clusters with more than one label class: 15 (32 rows).
    • ×4: "I'm sorry for your loss" → caring 1, grief 1, remorse 1, amusement 1
    • ×2: "He's hot" → neutral 1, admiration 1
    • ×2: "[NAME]. What a time to be alive." → neutral 1, approval 1
    • ×2: "because it’s funny" → amusement 1, confusion 1
    • ×2: "Oh that’s terrifying" → surprise 1, fear 1
  • Held-out rows with a near duplicate in train: dev 2 (0.2%), calibration 0 (0.0%), test 6 (0.1%)
    • train "Weird flex but okay" ~ test "Weird flex but ok" (J=0.87)
    • train "What? That doesn’t answer my question." ~ test "That doesn't answer my question." (J=0.80)
    • train "Love the username" ~ test "Love the username <3" (J=0.87)
    • train "Thanks a bunch!" ~ test "Thanks a bunch <3" (J=0.83)
    • train "Yes YES #YES" ~ test "* yes * yes * yes * yes" (J=1.00)

Largest train clusters:

  • ×4: "I'm sorry for your loss"
  • ×3: "HAHAHAHA!!!!!"
  • ×2: "He's hot"
  • ×2: "RemindMe! 7 days"
  • ×2: "RemindMe! 3 Days"

5. Junk

split empty main text main text under 10 characters
train 0 240
dev 0 5
calibration 0 6
test 0 40

Very short train examples: "He's hot" (neutral); "by [NAME]" (neutral); "Hey now!" (approval); "$390 CAD." (neutral); "[NAME] 3!" (neutral); "Go you!" (neutral); "Shame." (neutral); "Wow. Yes" (approval); "funny as." (amusement); "Or both!" (neutral); "Lol, wut" (amusement); "My bad." (neutral)

Pattern scan of train main texts (count, then the share of each class's rows):

pattern rows by class
placeholder [NAME]-style 6,405 (13.9%) admiration 14.5%, amusement 13.0%, anger 13.9%, annoyance 12.3%, approval 11.4%, caring 9.2%, confusion 12.9%, curiosity 13.1%, desire 17.0%, disappointment 13.9%, disapproval 10.6%, disgust 11.4%, embarrassment 7.3%, excitement 13.0%, fear 13.5%, gratitude 7.9%, grief 14.9%, joy 9.3%, love 19.4%, nervousness 7.7%, neutral 17.0%, optimism 12.9%, pride 11.1%, realization 10.3%, relief 10.3%, remorse 9.7%, sadness 10.2%, surprise 15.9%
lorem ipsum 0 (0.0%)
TODO/TBD/FIXME 0 (0.0%)
'As an AI' / refusal 1 (0.0%) annoyance 0.0%
chat preamble ('Here is/are...', 'Sure!') 39 (0.1%) admiration 0.1%, approval 0.4%, curiosity 0.1%, disappointment 0.1%, disapproval 0.1%, excitement 0.3%, gratitude 0.0%, joy 0.1%, neutral 0.1%, optimism 0.1%, sadness 0.2%
meta words (example/variation/message:) 0 (0.0%)
model thinking tags 0 (0.0%)
JSON/code-fence leftovers 0 (0.0%)
encoding garbage (mojibake/replacement char) 0 (0.0%)
HTML tag 3 (0.0%) approval 0.0%, neutral 0.0%
HTML entity 0 (0.0%)
base64-like run (40+ chars) 11 (0.0%) admiration 0.0%, amusement 0.0%, anger 0.1%, confusion 0.1%, curiosity 0.1%, neutral 0.0%, surprise 0.1%
URL 0 (0.0%)
  • placeholder [NAME]-style: emotion:go_emotions:train:efeo8w6:emotion (admiration): Unreal. [NAME]- Expert level | emotion:go_emotions:train:eeezcc7:emotion (disapproval): [NAME] would've had to get lucky on that one. Can't give up that opportunity. | emotion:go_emotions:train:eeajm3e:emotion (approval): I bet he’s not going to visit their hall of fame in order to avoid meeting w/ [NAME]

  • 'As an AI' / refusal: emotion:go_emotions:train:ef9ebro:emotion (annoyance): I cannot help but get a sense of schadenfreude when a seemingly shitty person ruins their own career right before my eyes. ☕️

  • chat preamble ('Here is/are...', 'Sure!'): emotion:go_emotions:train:eebqp3y:emotion (neutral): Sure, there are edge cases where this is less clear, but in the vast majority of conflict, this is true. | emotion:go_emotions:train:ee7fxah:emotion (neutral): Sure, its just that it was completely irrelevant to my reply | emotion:go_emotions:train:ed1m9qs:emotion (approval): Sure, they can help nurse [NAME] back to health.

  • HTML tag: emotion:go_emotions:train:ee5v5uy:emotion (neutral): <<creepy loui face.jpg>> | emotion:go_emotions:train:edfd4z0:emotion (approval): Yes, definitely classy. <rolls eyes> | emotion:go_emotions:train:ef8b1z0:emotion (neutral): If by "better", you mean "more fun", then I'd go with the second one. <smile>

  • base64-like run (40+ chars): emotion:go_emotions:train:ef9cbbq:emotion (surprise): Oh no, my secret identity as a neurosurgeon/entomologist/radiologist/anesthetist/dentist has been discovered! I knew I should have given le… | emotion:go_emotions:train:edrtk8u:emotion (amusement): I did anyway hahahahahahahahahahahahahahahahahahahahahahahahahahhahahahhahahaa fucking rekt XD | emotion:go_emotions:train:ef59tsi:emotion (neutral): Omg ur like soooooooooooooooooooooooooooooooooooooooooooooooooo quirky!

  • Possibly cut off: 1 of 4 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: admiration 100.0%, neutral 0.0%

    • emotion:go_emotions:train:edy8hly:emotion: …765434567654323454323456543345678987654323456789876565656565656565656565656565454545654565454323456765432345678765456 IQ

Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):

  • none

6. Samples

20 random train rows per kind: emotion-samples.txt. Reading notes are in the findings above.