Skip to content
JeffHub

QA report: triage

The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.

The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.

QA: triage

Verdict: PASS WITH NOTES

Triage (READY-triage 2026-09-30 13:33, train sha256 3aba76c5…3b16c0; train 66,532 rows = 16,633 messages × 4 questions). This is the first full check; the earlier triage reports checked withdrawn files. Each question type has its own report: triage-route.md, triage-urgency.md, triage-sentiment.md and triage-needs_human.md. Cue rates come from triage_extra.py.

Earlier cues, now fixed (checked)

Each rate below is the share of that label's rows containing the cue, compared across labels:

cue before now
"Subject: URGENT" 22.6% of urgency-4 against 3.0% of the rest 3.0–4.4% at every urgency level
"just wanted to" 20.4% of urgency-0 6.0–6.3% at every level, and 5.9% / 6.2% for needs_human no / yes
"by tomorrow" 14.9% of urgency-2 1.8–4.1%
"formal complaint" 7.9% of person-needed against 0.1% 2.8% / 3.8%
"Hi there" 17% of sentiment-3 4.1–5.6%
ends with "!" 39% of very positive 22–29% at every sentiment level
  • "manual", "real person" and "no reply needed" appear in 0 rows.

  • Surface-only models:

    • sentiment: 31.4% balanced accuracy (was 47.8%; chance 20%);
    • urgency: 33.4% (chance 20%);
    • needs_human: 60.5% (chance 50%);
    • route: 9.1% (chance 4.2%).

    No single surface feature beats chance by more than 1.3 points.

  • Route options: the correct team is the longest option 10.6% of the time against 9.8% chance, positions are uniform, and the option picker is at chance.

  • Duplicates: there are no near duplicates in train or across splits, and no identical prompt carries two different answers.

  • The edited messages read naturally. In my sample, urgency-4 messages now contain "just wanted to" ("I just wanted to say I need my boxes from unit 12 today!"), low-urgency messages carry "Subject: Urgent" (a spam investment offer; a confirmation email), and very negative messages end with "!".

Notes (meaning; for the model card)

  1. The remaining surface signal is meaning:
    • sentiment: unmatched ")" versus "(" counts, which are emoticons like :) and :( — and message length;
    • urgency: fewer "?" in urgent messages (questions are rarely emergencies);
    • needs_human: weak length and punctuation mixtures, with no single feature above chance.
  2. The standard phrase flags are meaning words:
    • urgency 4: "right now", "fix this", "deadline", "tonight", "court", "lawyer", "breach";
    • urgency 1: "no rush at all", the customer's own statement (kept by agreement);
    • urgency 0: "confirm that", "say thank you";
    • sentiment: "thrilled", "amazing", "warmest regards", "unacceptable", "nobody";
    • needs_human no: "a quick note", "to let you know that", "confirm that the", "perfectly";
    • route "other": "my thesis", "journalism student", "sponsoring" (student and sponsorship requests belong to no team).
  3. The label mix changed with the re-rating: urgency level 4 is now 34.7% of rows, and sentiment level 0 (very negative) is 35.6%. Dev and calibration are small: 1,368 and 1,208 rows, which is 342 and 302 messages (5 organisations each), so calibrate with care.
  4. About 51% of messages had small cue edits by qwen3.8-flash, recorded in source.decorrelated_by. Score targets are the mean of three raters (the plan, qwen3.8-max and deepseek-v4-flash).

QA: triage-route

Checked 2026-09-30 13:37 by adapters/qa/qa.py (READY file READY-triage, ).

Verdict: PASS WITH NOTES (full notes in triage-needs_human.md and triage.md)

Automatic flags (for the reviewer to judge; not all are problems)

  • 16 strong phrase flags (see list): review whether they are meaning or leakage
  • 7 standard phrase flags (≥2% of a class, mostly that class): review

Data checked

split rows families file
train 16,633 260 train.jsonl
dev 342 5 dev.jsonl
calibration 302 5 calibration.jsonl
test 1,814 30 test.jsonl
  • Train sha256: 3aba76c53cd16a1c2822d1b5e4090d183cc77fd94a233898c53501f25c3b16c0 (READY file gives no checksum)
  • Main text field (the text the phrase and length checks use): state.message.
  • Label classes: <listed option>, other (<listed option> = one of the per-row listed options such as t3 or o12). Row kinds (source.kind): route.

3. Balance

Label class share per split

lclass train dev calibration test train rows
<listed option> 93.2% 94.4% 95.4% 94.0% 15,509
other 6.8% 5.6% 4.6% 6.0% 1,124

Row kind share per split

kind train dev calibration test train rows
route 100.0% 100.0% 100.0% 100.0% 16,633

4. Format

split row-level format problems
train none
dev none
calibration none
test none

Options per choice row

split min median p99 max
train 5 13 38 39
dev 12 16 27 27
calibration 5 11 36 36
test 6 15 36 36

Prompt length in tokens

split measure median p99 max > 8192
train estimate: characters / 3 (upper bound for English) 841 2039 2524 0
dev estimate: characters / 3 (upper bound for English) 893 1614 1967 0
calibration estimate: characters / 3 (upper bound for English) 729 2327 2433 0
test estimate: characters / 3 (upper bound for English) 820 1958 2394 0

State key sets (train)

keys rows
channel, company, message 16,633 (100.0%)

Instructions (train)

  • Canonical (the most common text) 70.1%, reworded 26.7% (70 distinct rewordings), none 3.2%. Target about 70 / 27 / 3.
  • Canonical text: "Which team should handle this message? If it covers several issues, choose the team for the most important one."
  • source instruction tag: canonical 70.1%, none 3.2%, variant-42 0.5%, variant-21 0.5%, variant-68 0.5%, variant-30 0.5%
class canonical none
<listed option> 70.0% 3.2%
other 70.6% 2.9%

1. Shortcuts

Phrase statistics and models use a label-stratified sample of 16,633 train rows; models are scored on the full test file (1,814 rows).

Text length by label class (main text, characters)

split class rows p10 median p90 mean
train <listed option> 15509 136 434 1388 627
train other 1124 127 425 1431 637
test <listed option> 1705 124 423 1415 626
test other 109 110 410 1297 592

By row kind (train): main-text length, length of the rest of the state, options

kind rows median chars mean chars median other-state chars median options
route 16633 433 628 145 13

Correct option: longest / shortest / position / key

For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):

split rows correct is longest correct is shortest chance (1/listed) mean relative position (0 first, 1 last; 0.5 expected) position fifths
train 15509 10.6% 10.9% 9.8% 0.497 22% / 18% / 18% / 18% / 24%
dev 323 9.0% 10.5% 6.9% 0.484 18% / 22% / 21% / 17% / 22%
calibration 288 13.9% 9.0% 12.0% 0.520 19% / 19% / 18% / 22% / 22%
test 1705 11.1% 11.3% 10.1% 0.500 21% / 18% / 20% / 18% / 24%

Correct key and position by option count

split options rows mean options top correct keys most common position (0-based)
train 2-5 775 5.0 k4 23.9%, k3 23.1%, k2 22.8%, k1 21.0%, other 9.2% 4 (23.9%)
train 6-10 5116 7.8 k1 14.7%, k2 14.3%, k5 13.9%, k3 13.8%, k4 13.7% 1 (14.7%)
train 11-30 9463 17.7 k1 6.6%, k8 6.5%, k3 6.4%, k7 6.4%, k5 6.4% 1 (6.6%)
train 31-80 1279 35.4 other 4.7%, k27 3.7%, k29 3.6%, k13 3.6%, k1 3.4% 0 (4.7%)
test 6-10 709 7.3 k4 18.1%, k2 16.9%, k3 15.4%, k5 14.5%, k1 13.7% 4 (18.1%)
test 11-30 982 18.4 k3 7.3%, k8 6.7%, k2 6.6%, k1 6.3%, k9 6.2% 3 (7.3%)
test 31-80 123 33.6 k22 6.5%, k26 4.9%, k9 4.9%, k15 4.9%, k12 4.1% 22 (6.5%)
  • Train rows whose correct listed key is the first listed key (t1/o1): 10.2%, chance 9.8%.

Option count by label class (train)

class rows min median mean max
<listed option> 15509 5 13 15.6 39
other 1124 5 12 13.5 39

Source fields by label class (train)

Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.

Overall majority: 93.2%.

source field values purity top values → classes
other_kind 9 100.0% null: <listed option> 100.0%; a job application or question about working there: other 100.0%; a message clearly meant for a different organisation: other 100.0%; a sales pitch from another business offering its services: other 100.0%; a request to sponsor or donate to a local event: other 100.0%; a journalist or student asking for an interview or information for a project: other 100.0%
channel 6 93.2% email: <listed option> 93.3%, other 6.7%; web form: <listed option> 93.2%, other 6.8%; chat: <listed option> 93.0%, other 7.0%; phone transcript: <listed option> 92.7%, other 7.3%; app review: <listed option> 94.3%, other 5.7%; social media reply: <listed option> 92.9%, other 7.1%
language 11 93.2% English: <listed option> 93.2%, other 6.8%; Danish: <listed option> 92.9%, other 7.1%; Czech: <listed option> 93.8%, other 6.2%; French: <listed option> 95.3%, other 4.7%; Swedish: <listed option> 92.0%, other 8.0%; Dutch: <listed option> 91.1%, other 8.9%
length_target 17 93.2% 60: <listed option> 93.6%, other 6.4%; 120: <listed option> 93.2%, other 6.8%; 25: <listed option> 93.5%, other 6.5%; 220: <listed option> 93.0%, other 7.0%; 50: <listed option> 92.2%, other 7.8%; 10: <listed option> 92.5%, other 7.5%
tone 25 93.2% politely formal: <listed option> 90.2%, other 9.8%; cheerful: <listed option> 92.5%, other 7.5%; polite and friendly: <listed option> 89.0%, other 11.0%; warm: <listed option> 90.5%, other 9.5%; disappointed: <listed option> 95.8%, other 4.2%; appreciative: <listed option> 90.6%, other 9.4%
prompt_version 3 93.2% 3: <listed option> 93.4%, other 6.6%; 2: <listed option> 92.9%, other 7.1%; 1: <listed option> 93.1%, other 6.9%
human_reason 7 93.2% manual: <listed option> 98.9%, other 1.1%; legal_safety: <listed option> 100.0%; routine: <listed option> 88.3%, other 11.7%; acknowledge: <listed option> 81.5%, other 18.5%; escalation: <listed option> 100.0%; no_reply: <listed option> 80.6%, other 19.4%
decorrelated_by 2 93.2% qwen3.8-flash: <listed option> 92.3%, other 7.7%; null: <listed option> 94.3%, other 5.7%
situation 17 93.2% null: <listed option> 92.9%, other 7.1%; the writer reports or confirms something that needs no decision, such as that a delivery arrived, a payment was made, a problem sorted itself out, or a form was sent: <listed option> 80.9%, other 19.1%; the writer is about to take it further and says what they will do next, for example cancel, go to a regulator or ombudsman, post a public review, contact t…
regenerated 2 93.2% false: <listed option> 93.0%, other 7.0%; true: <listed option> 94.0%, other 6.0%

Formatting by label class (main text, share of rows)

feature <listed option> other
ends with ? 4.6% 2.0%
ends with . 38.0% 35.2%
ends with ! 13.3% 14.5%
no end punctuation 43.9% 48.0%
starts lowercase 6.4% 7.0%
all lowercase 4.2% 6.4%
has a digit 88.8% 76.7%
has newline 77.3% 73.4%
has quotes 6.5% 1.5%
has markup (HTML/markdown) 3.1% 3.3%
has URL 0.0% 0.4%
non-ASCII 51.7% 44.0%
non-Latin script 0.0% 0.1%
emoji 5.4% 5.0%
ALL-CAPS word (4+) 16.1% 17.2%
contains ' - ' or — 17.8% 14.0%

Same, by row kind

feature route
ends with ? 4%
ends with . 38%
ends with ! 13%
no end punctuation 44%
starts lowercase 6%
all lowercase 4%
has a digit 88%
has newline 77%
has quotes 6%
has markup (HTML/markdown) 3%
has URL 0%
non-ASCII 51%
non-Latin script 0%
emoji 5%
ALL-CAPS word (4+) 16%
contains ' - ' or — 18%

Over-represented words and phrases per label class (main text)

Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.

Words, <listed option>: today z=11 28.9% vs 9.4%; before z=11 21.5% vs 5.3%; now z=10 32.3% vs 14.1%; account z=8 25.7% vs 12.3%; also z=8 20.5% vs 9.2%; cannot z=7 10.0% vs 2.3%; fix z=7 10.5% vs 0.4%; my z=7 63.3% vs 44.8%; tomorrow z=7 13.3% vs 5.1%; deadline z=7 9.2% vs 2.6%; right z=6 14.8% vs 6.9%; legal z=6 6.7% vs 1.1%; but z=6 29.1% vs 18.3%; immediately z=6 12.4% vs 5.4%; need z=6 17.6% vs 9.3%; card z=6 6.5% vs 1.7%; into z=6 9.8% vs 4.1%; t z=5 10.1% vs 4.4%; charge z=5 5.6% vs 1.3%; says z=5 5.1% vs 1.0%

Words, other: interview z=16 7.5% vs 0.8%; role z=16 6.3% vs 0.5%; university z=15 6.4% vs 0.7%; opportunity z=14 4.5% vs 0.2%; thesis z=14 4.7% vs 0.1%; delivery z=13 10.7% vs 2.7%; cv z=13 4.3% vs 0.2%; research z=13 4.8% vs 0.4%; hiring z=13 4.3% vs 0.3%; solutions z=13 6.1% vs 1.0%; proposal z=13 4.8% vs 0.5%; sponsorship z=12 4.4% vs 0.5%; sponsor z=12 3.6% vs 0.1%; company z=12 14.7% vs 5.4%; offer z=12 5.2% vs 0.8%; student z=12 6.2% vs 1.2%; position z=11 5.0% vs 0.8%; testing z=11 3.2% vs 0.2%; journalism z=11 3.0% vs 0.1%; best z=11 14.4% vs 5.6%

2–4-word phrases, <listed option>: right now z=8 11.0% vs 2.2%; i need z=6 8.5% vs 2.3%; my account z=6 7.7% vs 1.9%; i cannot z=5 5.3% vs 1.2%; before the z=5 5.3% vs 1.5%; today i z=5 4.4% vs 0.8%; need to z=5 7.5% vs 3.1%; is not z=5 5.4% vs 1.8%; if i z=5 4.6% vs 1.2%; to the z=5 12.4% vs 7.0%; look into z=4 3.4% vs 0.4%; instead of z=4 3.2% vs 0.5%; to my z=4 6.2% vs 2.8%; but the z=4 3.9% vs 1.2%; when i z=4 6.1% vs 2.8%; tomorrow morning z=4 3.3% vs 0.2%; my bank z=4 3.5% vs 1.0%; i can z=4 5.3% vs 2.3%; could you z=4 9.1% vs 5.2%; you please z=4 4.1% vs 1.6%

2–4-word phrases, other: your company z=12 6.9% vs 1.4%; to your z=11 9.5% vs 3.1%; to say z=11 13.3% vs 5.3%; my cv z=11 2.7% vs 0.1%; my thesis z=10 2.8% vs 0.1%; my application z=10 4.5% vs 0.9%; working with z=10 3.6% vs 0.5%; hr 2024 z=10 2.4% vs 0.1%; dear hiring z=10 2.2% vs 0.1%; reaching out z=10 4.8% vs 1.1%; we are z=10 17.4% vs 8.6%; the delivery z=9 3.3% vs 0.5%; interview request z=9 2.0% vs 0.1%; application for z=9 3.6% vs 0.7%; i wanted to z=9 8.8% vs 3.3%; i wanted z=9 8.8% vs 3.4%; hiring team z=9 2.0% vs 0.1%; my application for the z=9 2.1% vs 0.2%; my application for z=9 2.3% vs 0.3%; dear hiring team z=9 1.8% vs 0.1%

Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):

  • <listed option>: before 21.5% vs 5.3%
  • <listed option>: right now 11.0% vs 2.2%
  • <listed option>: fix 10.5% vs 0.4%
  • <listed option>: cannot 10.0% vs 2.3%
  • <listed option>: my account 7.7% vs 1.9%
  • other: interview 7.5% vs 0.8%
  • other: your company 6.9% vs 1.4%
  • <listed option>: legal 6.7% vs 1.1%
  • other: university 6.4% vs 0.7%
  • other: role 6.3% vs 0.5%
  • other: student 6.2% vs 1.2%
  • other: solutions 6.1% vs 1.0%
  • <listed option>: charge 5.6% vs 1.3%
  • <listed option>: i cannot 5.3% vs 1.2%
  • other: offer 5.2% vs 0.8%
  • <listed option>: says 5.1% vs 1.0%

Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):

  • all rows (16,633 rows): other: thesis 4.7% of class, 75% of its 71 rows; other: sponsor 3.6% of class, 71% of its 56 rows; other: student at 3.4% of class, 95% of its 40 rows; other: my thesis 2.8% of class, 80% of its 40 rows; other: journalism student 2.7% of class, 100% of its 30 rows; other: careers 2.5% of class, 80% of its 35 rows; other: sponsoring 2.4% of class, 77% of its 35 rows
  • state.channel = email (6,611 rows): other: thesis 6.4% of class, 72% of its 39 rows; other: student at 4.8% of class, 100% of its 21 rows
  • state.channel = web form (3,009 rows): none
  • state.channel = chat (2,529 rows): none
  • state.channel = phone transcript (1,846 rows): none
  • state.channel = app review (1,430 rows): none
  • state.channel = social media reply (1,208 rows): none

2. Duplicates and split separation

Families shared between splits

splits shared families examples
train ∩ dev 0
train ∩ calibration 0
train ∩ test 0
dev ∩ calibration 0
dev ∩ test 0
calibration ∩ test 0
  • Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
  • Train rows identical in the whole prompt (state, options, instructions): 0.
  • Identical whole prompt, different answer: 0 groups (0 rows).
  • Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)

Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):

split rows examples
dev 0 (0.0%)
calibration 0 (0.0%)
test 0 (0.0%)

5. Junk

split empty main text main text under 10 characters
train 0 0
dev 0 0
calibration 0 0
test 0 0

Pattern scan of train main texts (count, then the share of each class's rows):

pattern rows by class
placeholder [NAME]-style 0 (0.0%)
lorem ipsum 0 (0.0%)
TODO/TBD/FIXME 0 (0.0%)
'As an AI' / refusal 1 (0.0%) <listed option> 0.0%
chat preamble ('Here is/are...', 'Sure!') 0 (0.0%)
meta words (example/variation/message:) 0 (0.0%)
model thinking tags 0 (0.0%)
JSON/code-fence leftovers 0 (0.0%)
encoding garbage (mojibake/replacement char) 0 (0.0%)
HTML tag 511 (3.1%) <listed option> 3.1%, other 3.2%
HTML entity 0 (0.0%)
base64-like run (40+ chars) 1 (0.0%) other 0.1%
URL 8 (0.0%) <listed option> 0.0%, other 0.4%
  • 'As an AI' / refusal: triage-c290-m140-route (<listed option>): Subject: Concern regarding stonework at Hrubý Rohozec ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I hope this message finds you well, though I must confes…

  • HTML tag: triage-c227-m067-route (<listed option>): ---------- Forwarded message --------- ⏎ From: Berlin Dental Group noreply@berlindental.eu ⏎ Date: 24 May 2024 at 08:15 ⏎ Subject: RE: Ne… | triage-c196-m133-route (<listed option>): Subject: Fwd: RE: Redemption statement completion - 12 Maple Road ⏎ ⏎ From: Helen Croft helen.croft@smithsolicitors.co.uk ⏎ Sent: 21 Nov… | triage-c151-m155-route (<listed option>): ---------- Forwarded message --------- ⏎ From: Billing Dept billing@secureguard.cz ⏎ Date: Mon, Oct 14, 2024 at 8:00 AM ⏎ Subject: Your i…

  • base64-like run (40+ chars): triage-c056-m162-route (other): Subject: URGENT: Your store domain will expire!!! ⏎ ⏎ Dear Beloved Manager, ⏎ ⏎ I am writing to you today with a heavy heart because your…

  • URL: triage-c133-m096-route (other): Subject: Strategic Partnership Proposal - Global Digital Infrastructure ⏎ ⏎ Dear Sir/Madam, ⏎ ⏎ I hope this email finds you well. My name… | triage-c073-m068-route (other): Name: Rajesh Kumer ⏎ Email: rajesh.kumer82@yahho.co.in ⏎ Topic: Claims Processing ⏎ Message: ⏎ Dear Sir or Madam, ⏎ ⏎ I am writting to you… | triage-c074-m101-route (other): Subject: We offer printing services for your charity ⏎ ⏎ Hello friends, ⏎ ⏎ My name is Laszlo and I am working for GreenPrint Solutions K…

  • Possibly cut off: 5,505 of 11,135 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: <listed option> 49.1%, other 54.6%

    • triage-c195-m064-route: …t pop by next month for a routine checkup. ⏎ ⏎ Thank you again for everyting. ⏎ ⏎ With much appriciation, ⏎ Katarzyna Wisniewska
    • triage-c142-m013-route: …ay. It truly makes a difference to families like ours. ⏎ ⏎ Warmest regards, ⏎ ⏎ Kristin Hagen ⏎ Parent / Guardian ⏎ +47 912 34 …
    • triage-c206-m086-route: …l double charge on my last invoice, probably just a glitch. No rush at all. ⏎ Warmly, ⏎ Sarah Jenkins ⏎ Homeowner ⏎ 082 555 1234

Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):

  • ×175: "---------- Forwarded message ---------" (<listed option> 163, other 12)
  • ×166: "Any unauthorized review, use, disclosure, or distribution is strictly prohibited." (<listed option> 155, other 11)
  • ×159: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and intended solely for the use of the individu…" (<listed option> 146, other 13)
  • ×148: "----- Forwarded message -----" (<listed option> 142, other 6)
  • ×132: "I hope this message finds you well." (<listed option> 121, other 11)
  • ×119: "-----Original Message-----" (<listed option> 111, other 8)
  • ×118: "If you have received this email in error, please notify the sender immediately and delete this email from your system." (<listed option> 112, other 6)
  • ×116: "If you are not the intended recipient, please delete all copies and notify the sender immediately." (<listed option> 112, other 4)
  • ×107: "If you have received this email in error, please notify the sender immediately and delete this message from your system." (<listed option> 97, other 10)
  • ×107: "Topic: Billing and Invoices" (<listed option> 99, other 8)
  • ×91: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and may also be privileged." (<listed option> 87, other 4)
  • ×81: "CONFIDENTIALITY NOTICE: This email and any attachments are strictly confidential and intended solely for the addressee." (<listed option> 71, other 10)
  • ×66: "I hope this email finds you well." (<listed option> 59, other 7)
  • ×63: "----- End forwarded message -----" (<listed option> 60, other 3)
  • ×60: "--- Forwarded message ---" (<listed option> 56, other 4)

6. Samples

20 random train rows per kind: triage-route-samples.txt. Reading notes are in the findings above.

QA: triage-urgency

Checked 2026-09-30 13:36 by adapters/qa/qa.py (READY file READY-triage, ).

Verdict: PASS WITH NOTES (full notes in triage-needs_human.md and triage.md)

Automatic flags (for the reviewer to judge; not all are problems)

  • 56 strong phrase flags (see list): review whether they are meaning or leakage
  • 60 standard phrase flags (≥2% of a class, mostly that class): review

Data checked

split rows families file
train 16,633 260 train.jsonl
dev 342 5 dev.jsonl
calibration 302 5 calibration.jsonl
test 1,814 30 test.jsonl
  • Train sha256: 3aba76c53cd16a1c2822d1b5e4090d183cc77fd94a233898c53501f25c3b16c0 (READY file gives no checksum)
  • Main text field (the text the phrase and length checks use): state.message.
  • Label classes: 0, 1, 2, 3, 4 (<listed option> = one of the per-row listed options such as t3 or o12). Row kinds (source.kind): urgency.

3. Balance

Label class share per split

lclass train dev calibration test train rows
0 17.9% 18.4% 18.2% 18.2% 2,976
1 11.5% 10.2% 10.6% 12.1% 1,914
2 22.4% 25.4% 24.2% 21.3% 3,724
3 13.5% 13.7% 13.2% 13.2% 2,253
4 34.7% 32.2% 33.8% 35.1% 5,766

Row kind share per split

kind train dev calibration test train rows
urgency 100.0% 100.0% 100.0% 100.0% 16,633

4. Format

split row-level format problems
train none
dev none
calibration none
test none

Prompt length in tokens

split measure median p99 max > 8192
train estimate: characters / 3 (upper bound for English) 205 881 1334 0
dev estimate: characters / 3 (upper bound for English) 221 880 1088 0
calibration estimate: characters / 3 (upper bound for English) 215 975 1152 0
test estimate: characters / 3 (upper bound for English) 203 842 1165 0

State key sets (train)

keys rows
channel, company, message 16,633 (100.0%)

Instructions (train)

  • Canonical (the most common text) 69.7%, reworded 27.1% (70 distinct rewordings), none 3.2%. Target about 70 / 27 / 3.
  • Canonical text: "How urgently does this need a response?"
  • source instruction tag: canonical 69.7%, none 3.2%, variant-62 0.5%, variant-25 0.5%, variant-16 0.5%, variant-6 0.5%
class canonical none
0 69.4% 2.9%
1 69.7% 3.2%
2 70.5% 3.2%
3 70.6% 2.8%
4 69.0% 3.6%

1. Shortcuts

Phrase statistics and models use a label-stratified sample of 16,633 train rows; models are scored on the full test file (1,814 rows).

Text length by label class (main text, characters)

split class rows p10 median p90 mean
train 0 2976 121 439 1395 628
train 1 1914 170 622 1532 743
train 2 3724 126 384 1306 566
train 3 2253 120 405 1338 582
train 4 5766 143 450 1413 647
test 0 331 117 403 1576 629
test 1 220 150 667 1542 745
test 2 386 114 396 1277 546
test 3 240 100 423 1401 608
test 4 637 143 435 1370 633

By row kind (train): main-text length, length of the rest of the state, options

kind rows median chars mean chars median other-state chars median options
urgency 16633 433 628 145 0

Source fields by label class (train)

Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.

Overall majority: 34.7%.

source field values purity top values → classes
human_reason 7 60.5% manual: 4 46.0%, 2 29.0%; legal_safety: 4 66.0%, 2 17.2%; routine: 2 53.0%, 1 35.5%; acknowledge: 0 63.3%, 1 23.0%; escalation: 4 71.2%, 3 21.8%; no_reply: 0 86.4%, 1 13.6%
situation 17 53.7% null: 4 33.0%, 2 26.6%; the writer reports or confirms something that needs no decision, such as that a delivery arrived, a payment was made, a problem sorted itself out, or a form was sent: 0 73.9%, 1 19.1%; the writer is about to take it further and says what they will do next, for example cancel, go to a regulator or ombudsman, post a public review, contact the press, or make a formal complain…
tone 25 49.5% politely formal: 2 35.2%, 0 28.8%; cheerful: 2 31.9%, 0 21.3%; polite and friendly: 2 28.3%, 0 23.8%; warm: 2 30.5%, 0 23.5%; disappointed: 4 53.8%, 2 18.1%; appreciative: 2 29.4%, 0 26.8%
other_kind 9 38.8% null: 4 37.2%, 2 23.3%; a job application or question about working there: 1 45.0%, 0 41.1%; a message clearly meant for a different organisation: 0 77.9%, 1 13.8%; a sales pitch from another business offering its services: 0 67.8%, 1 28.3%; a request to sponsor or donate to a local event: 1 45.9%, 0 38.5%; a journalist or student asking for an interview or information for a project: 0 42.4%, 1 4…
channel 6 34.7% email: 4 33.6%, 2 21.4%; web form: 4 34.6%, 2 22.4%; chat: 4 38.6%, 2 26.3%; phone transcript: 4 36.5%, 0 20.9%; app review: 4 32.9%, 0 22.2%; social media reply: 4 32.0%, 2 27.5%
language 11 34.7% English: 4 34.9%, 2 22.5%; Danish: 4 31.7%, 0 25.1%; Czech: 4 28.8%, 2 23.2%; French: 4 34.1%, 2 24.1%; Swedish: 4 27.6%, 2 24.5%; Dutch: 4 34.2%, 0 21.5%
length_target 17 34.7% 60: 4 34.2%, 2 22.8%; 120: 4 34.6%, 2 20.5%; 25: 4 33.3%, 2 28.1%; 220: 4 33.6%, 2 19.4%; 50: 4 36.6%, 2 26.5%; 10: 4 32.8%, 2 23.6%
prompt_version 3 34.7% 3: 4 35.4%, 2 20.6%; 2: 4 32.7%, 2 26.7%; 1: 4 45.7%, 2 21.6%
decorrelated_by 2 34.7% qwen3.8-flash: 4 32.1%, 2 23.1%; null: 4 37.4%, 2 21.6%
regenerated 2 34.7% false: 4 33.9%, 2 23.5%; true: 4 36.7%, 0 20.6%
target_kind 2 34.7% soft: 4 31.1%, 2 26.2%; hard: 4 41.5%, 0 28.5%

Formatting by label class (main text, share of rows)

feature 0 1 2 3 4
ends with ? 0.1% 2.9% 8.3% 6.5% 3.8%
ends with . 38.9% 33.9% 33.4% 36.5% 41.9%
ends with ! 14.1% 12.9% 13.3% 12.9% 13.3%
no end punctuation 46.4% 49.9% 44.6% 44.0% 40.9%
starts lowercase 5.4% 4.0% 7.1% 6.5% 7.4%
all lowercase 5.2% 3.3% 4.7% 4.9% 3.7%
has a digit 84.0% 84.7% 87.1% 89.7% 91.0%
has newline 76.9% 82.1% 76.5% 76.7% 75.8%
has quotes 2.1% 5.1% 5.2% 8.4% 8.4%
has markup (HTML/markdown) 4.8% 4.0% 2.0% 2.4% 3.1%
has URL 0.2% 0.1% 0.0% 0.0% 0.0%
non-ASCII 46.1% 51.2% 50.6% 53.8% 53.3%
non-Latin script 0.0% 0.0% 0.0% 0.0% 0.0%
emoji 4.9% 4.1% 6.6% 5.9% 5.2%
ALL-CAPS word (4+) 16.4% 16.1% 15.8% 15.8% 16.5%
contains ' - ' or — 14.9% 18.0% 15.0% 18.2% 20.1%

Same, by row kind

feature urgency
ends with ? 4%
ends with . 38%
ends with ! 13%
no end punctuation 44%
starts lowercase 6%
all lowercase 4%
has a digit 88%
has newline 77%
has quotes 6%
has markup (HTML/markdown) 3%
has URL 0%
non-ASCII 51%
non-Latin script 0%
emoji 5%
ALL-CAPS word (4+) 16%
contains ' - ' or — 18%

Over-represented words and phrases per label class (main text)

Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.

Words, 0: thank z=25 43.8% vs 19.5%; perfectly z=25 13.9% vs 2.2%; everything z=22 26.1% vs 10.2%; arrived z=21 12.1% vs 2.7%; say z=20 19.8% vs 7.1%; went z=20 11.5% vs 2.7%; bye z=17 6.7% vs 1.0%; perfect z=17 6.0% vs 0.8%; wanted z=17 20.3% vs 9.0%; successfully z=17 5.6% vs 0.6%; completed z=16 5.7% vs 1.0%; smoothly z=15 6.1% vs 1.2%; was z=15 43.0% vs 27.1%; note z=15 11.0% vs 3.9%; all z=15 29.9% vs 17.0%; made z=15 9.4% vs 3.0%; finally z=15 6.4% vs 1.5%; sorted z=14 8.9% vs 3.1%; such z=13 11.0% vs 4.6%; settled z=13 4.3% vs 0.8%

Words, 1: rush z=27 19.5% vs 2.0%; whenever z=24 15.4% vs 1.6%; question z=17 15.4% vs 3.8%; wondering z=16 8.5% vs 1.2%; moment z=14 6.9% vs 1.2%; small z=13 13.9% vs 4.6%; thought z=12 7.8% vs 1.9%; would z=12 20.3% vs 8.5%; such z=11 13.2% vs 4.8%; curious z=11 3.0% vs 0.2%; maybe z=11 10.5% vs 3.5%; much z=11 33.2% vs 17.2%; fine z=11 9.7% vs 3.2%; all z=10 32.9% vs 17.6%; warmest z=10 10.7% vs 3.9%; anyway z=10 14.1% vs 5.8%; quick z=10 11.2% vs 4.2%; recently z=10 5.1% vs 1.2%; how z=10 27.5% vs 14.3%; thank z=10 38.9% vs 21.9%

Words, 2: could z=27 31.4% vs 12.6%; question z=16 10.1% vs 3.8%; kindly z=14 6.8% vs 2.2%; would z=13 14.7% vs 8.4%; can z=13 23.2% vs 15.4%; thanks z=12 18.1% vs 11.4%; love z=12 10.0% vs 5.3%; hope z=11 9.6% vs 5.1%; standard z=11 7.1% vs 3.3%; hello z=11 22.6% vs 16.0%; possible z=11 4.3% vs 1.5%; ask z=11 6.9% vs 3.3%; plan z=11 7.2% vs 3.6%; much z=10 24.0% vs 17.6%; advise z=10 3.7% vs 1.2%; next z=10 15.8% vs 10.6%; hi z=10 23.8% vs 17.7%; want z=10 8.6% vs 4.9%; arrange z=10 3.1% vs 1.0%; please z=10 31.9% vs 25.5%

Words, 3: before z=17 34.2% vs 18.3%; today z=16 42.1% vs 25.3%; need z=12 26.1% vs 15.6%; tomorrow z=11 19.7% vs 11.6%; closes z=8 3.4% vs 1.2%; afternoon z=8 9.7% vs 5.7%; friday z=8 7.0% vs 3.7%; thursday z=8 5.5% vs 2.7%; could z=7 21.8% vs 16.0%; tomorow z=7 2.6% vs 0.8%; can z=7 22.0% vs 16.4%; 17 z=7 5.5% vs 2.8%; still z=7 12.5% vs 8.6%; evening z=6 4.8% vs 2.6%; trustpilot z=6 1.4% vs 0.4%; day z=6 10.4% vs 7.3%; be z=6 26.5% vs 21.8%; also z=6 23.5% vs 19.1%; approved z=6 3.9% vs 2.1%; please z=6 31.0% vs 26.3%

Words, 4: today z=36 49.9% vs 15.7%; right z=30 27.7% vs 7.1%; now z=29 50.1% vs 20.9%; fix z=29 20.8% vs 3.9%; immediately z=26 22.7% vs 6.1%; deadline z=26 18.0% vs 3.8%; tomorrow z=23 22.2% vs 7.7%; cannot z=22 17.4% vs 5.3%; tonight z=21 10.1% vs 1.3%; legal z=20 12.4% vs 3.1%; nobody z=18 8.2% vs 1.4%; filing z=18 9.6% vs 2.3%; court z=18 6.9% vs 0.8%; before z=17 29.6% vs 15.6%; lose z=17 6.4% vs 0.7%; if z=17 43.8% vs 26.1%; hour z=17 6.1% vs 0.6%; completely z=17 12.3% vs 4.4%; breach z=16 5.8% vs 0.5%; will z=16 25.1% vs 13.0%

2–4-word phrases, 0: to say z=26 17.6% vs 3.3%; thank you z=23 42.2% vs 19.1%; confirm that z=21 8.9% vs 0.7%; to confirm z=21 10.2% vs 1.5%; that the z=21 14.4% vs 3.6%; to confirm that z=21 8.2% vs 0.6%; i wanted z=20 10.8% vs 2.2%; i wanted to z=20 10.7% vs 2.2%; everything is z=19 8.7% vs 1.4%; let you z=18 6.7% vs 0.7%; you know z=18 8.0% vs 1.4%; to let z=18 6.6% vs 0.9%; let you know z=17 6.1% vs 0.6%; you for z=17 17.0% vs 6.5%; thank you for z=17 16.4% vs 6.2%; to let you z=17 6.0% vs 0.6%; to let you know z=17 5.7% vs 0.6%; this morning z=17 16.1% vs 6.2%; wanted to z=16 19.9% vs 8.9%; confirm that the z=16 5.1% vs 0.2%

2–4-word phrases, 1: no rush z=24 15.4% vs 1.3%; at all z=21 16.2% vs 2.7%; rush at z=21 10.3% vs 0.5%; rush at all z=21 10.3% vs 0.4%; no rush at z=19 9.1% vs 0.4%; no rush at all z=19 9.1% vs 0.4%; whenever you z=18 9.0% vs 0.8%; you get z=13 4.8% vs 0.5%; whenever you get z=13 4.2% vs 0.3%; question about z=13 7.9% vs 1.7%; you get a z=12 4.1% vs 0.4%; whenever you have z=12 4.0% vs 0.4%; you have a z=12 5.0% vs 0.7%; whenever you get a z=12 3.6% vs 0.2%; one small z=12 4.4% vs 0.6%; quick question z=12 6.1% vs 1.2%; perfectly fine z=11 3.3% vs 0.2%; just wondering z=11 3.7% vs 0.4%; wondering if z=11 3.9% vs 0.6%; whenever you have a z=11 3.1% vs 0.3%

2–4-word phrases, 2: could you z=23 18.6% vs 6.0%; could you please z=14 7.0% vs 2.2%; you please z=14 7.7% vs 2.8%; let me know z=12 5.4% vs 1.8%; me know z=12 5.4% vs 1.8%; do i z=12 4.9% vs 1.5%; question about z=12 5.0% vs 1.7%; i would z=11 6.4% vs 2.9%; you kindly z=11 3.7% vs 1.2%; look into z=11 5.7% vs 2.4%; let me know what z=11 2.4% vs 0.4%; me know what z=11 2.4% vs 0.4%; within the next day z=10 2.4% vs 0.5%; so much z=10 19.5% vs 13.4%; like to z=10 2.8% vs 0.7%; the next day z=10 2.5% vs 0.5%; next day z=10 2.5% vs 0.5%; how do i z=10 2.1% vs 0.4%; how do z=10 2.4% vs 0.5%; quick question z=10 3.7% vs 1.3%

2–4-word phrases, 3: end of z=14 5.6% vs 1.2%; before the end of z=14 4.0% vs 0.5%; before the end z=14 4.0% vs 0.5%; the end of z=13 4.8% vs 1.0%; the end z=12 4.8% vs 1.1%; of thursday z=12 2.9% vs 0.3%; end of thursday z=12 2.9% vs 0.3%; the end of thursday z=12 2.9% vs 0.3%; before the z=11 9.5% vs 4.4%; need to z=10 12.2% vs 6.4%; i need z=10 12.9% vs 7.3%; today before z=10 2.8% vs 0.6%; next day z=9 2.8% vs 0.7%; the next day z=9 2.8% vs 0.7%; within the next day z=9 2.7% vs 0.6%; before then z=8 2.3% vs 0.5%; we need z=8 4.3% vs 1.8%; i need to z=8 6.2% vs 3.1%; need to know z=8 3.4% vs 1.3%; could someone z=7 3.5% vs 1.4%

2–4-word phrases, 4: right now z=34 24.3% vs 3.1%; fix this z=25 13.4% vs 1.3%; this is z=22 19.3% vs 6.9%; tomorrow morning z=18 7.2% vs 0.9%; today i z=18 8.5% vs 1.9%; if this z=17 9.7% vs 2.9%; i cannot z=17 9.3% vs 2.8%; or i z=16 5.4% vs 0.8%; 1 5 z=15 4.9% vs 0.5%; is not z=15 8.9% vs 3.1%; 1 5 stars z=15 4.6% vs 0.5%; 17 00 z=14 5.1% vs 1.0%; my lawyer z=14 4.3% vs 0.5%; 00 today z=14 4.6% vs 0.8%; or we z=14 4.4% vs 0.2%; right now i z=14 4.1% vs 0.5%; filing a z=14 5.1% vs 1.2%; complaint with z=14 4.1% vs 0.6%; this to z=14 4.9% vs 1.1%; is completely z=14 4.0% vs 0.6%

Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):

  • 4: right now 24.3% vs 3.1%
  • 4: fix 20.8% vs 3.9%
  • 1: rush 19.5% vs 2.0%
  • 4: deadline 18.0% vs 3.8%
  • 0: to say 17.6% vs 3.3%
  • 1: at all 16.2% vs 2.7%
  • 1: whenever 15.4% vs 1.6%
  • 1: question 15.4% vs 3.8%
  • 1: no rush 15.4% vs 1.3%
  • 0: perfectly 13.9% vs 2.2%
  • 4: fix this 13.4% vs 1.3%
  • 0: arrived 12.1% vs 2.7%
  • 0: went 11.5% vs 2.7%
  • 0: i wanted 10.8% vs 2.2%
  • 0: i wanted to 10.7% vs 2.2%
  • 1: rush at 10.3% vs 0.5%
  • 1: rush at all 10.3% vs 0.4%
  • 0: to confirm 10.2% vs 1.5%
  • 4: tonight 10.1% vs 1.3%
  • 4: filing 9.6% vs 2.3%
  • 1: no rush at 9.1% vs 0.4%
  • 1: no rush at all 9.1% vs 0.4%
  • 1: whenever you 9.0% vs 0.8%
  • 0: confirm that 8.9% vs 0.7%
  • 0: everything is 8.7% vs 1.4%
  • 1: wondering 8.5% vs 1.2%
  • 4: today i 8.5% vs 1.9%
  • 0: to confirm that 8.2% vs 0.6%
  • 4: nobody 8.2% vs 1.4%
  • 0: you know 8.0% vs 1.4%
  • 1: question about 7.9% vs 1.7%
  • 1: thought 7.8% vs 1.9%
  • 4: tomorrow morning 7.2% vs 0.9%
  • 4: court 6.9% vs 0.8%
  • 1: moment 6.9% vs 1.2%
  • 0: bye 6.7% vs 1.0%
  • 0: let you 6.7% vs 0.7%
  • 0: to let 6.6% vs 0.9%
  • 4: lose 6.4% vs 0.7%
  • 0: finally 6.4% vs 1.5%

Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):

  • all rows (16,633 rows): 4: right now 24.3% of class, 81% of its 1732 rows; 4: fix 20.8% of class, 74% of its 1626 rows; 4: deadline 18.0% of class, 71% of its 1451 rows; 4: fix this 13.4% of class, 85% of its 911 rows; 1: rush at all 10.3% of class, 75% of its 264 rows; 4: tonight 10.1% of class, 80% of its 727 rows; 1: no rush at all 9.1% of class, 74% of its 234 rows; 0: confirm that 8.9% of class, 73% of its 363 rows; 4: today i 8.5% of class, 71% of its 690 rows; 0: to confirm that 8.2% of class, 76% of its 321 rows; 4: nobody 8.2% of class, 76% of its 619 rows; 4: tomorrow morning 7.2% of class, 81% of its 513 rows; 4: court 6.7% of class, 81% of its 476 rows; 4: emergency 6.7% of class, 74% of its 525 rows; 4: lose 6.4% of class, 83% of its 445 rows; 4: lawyer 5.9% of class, 78% of its 441 rows; 4: hour 5.8% of class, 87% of its 384 rows; 4: breach 5.6% of class, 86% of its 375 rows; 4: or i 5.4% of class, 77% of its 400 rows; 4: noon 5.2% of class, 84% of its 360 rows; 4: filing a 5.1% of class, 70% of its 421 rows; 4: 17 00 5.1% of class, 73% of its 403 rows; 0: confirm that the 5.1% of class, 82% of its 185 rows; 4: unacceptable 5.1% of class, 76% of its 383 rows; 4: midnight 5.0% of class, 83% of its 344 rows; 4: consumer 4.9% of class, 74% of its 381 rows; 4: immediate 4.9% of class, 78% of its 358 rows; 4: shaking 4.8% of class, 95% of its 293 rows; 0: say thank you 4.6% of class, 72% of its 191 rows; 4: 1 5 stars 4.6% of class, 84% of its 315 rows; 0: to confirm that the 4.6% of class, 83% of its 163 rows; 0: let you know that 4.5% of class, 72% of its 188 rows; 4: 00 today 4.5% of class, 76% of its 345 rows; 4: 5pm 4.5% of class, 81% of its 318 rows; 4: posting 4.5% of class, 74% of its 349 rows; 4: or we 4.4% of class, 91% of its 280 rows; 4: my lawyer 4.3% of class, 81% of its 308 rows; 4: this now 4.3% of class, 94% of its 263 rows; 0: am writing to confirm 4.3% of class, 81% of its 156 rows; 4: begging 4.3% of class, 95% of its 260 rows
  • state.channel = email (6,611 rows): 4: right now 26.7% of class, 81% of its 736 rows; 4: fix 22.6% of class, 76% of its 664 rows; 4: fix this 15.3% of class, 85% of its 397 rows; 0: confirm that 15.2% of class, 73% of its 252 rows; 0: to confirm that 14.2% of class, 76% of its 226 rows; 4: tonight 12.2% of class, 80% of its 339 rows; 4: emergency 10.0% of class, 75% of its 299 rows; 1: rush at all 9.2% of class, 75% of its 104 rows; 4: tomorrow morning 9.2% of class, 78% of its 260 rows; 0: subject thank you 9.0% of class, 76% of its 142 rows; 4: nobody 9.0% of class, 75% of its 264 rows; 4: immediate 8.9% of class, 77% of its 256 rows; 4: breach 8.8% of class, 86% of its 226 rows; 4: 17 00 8.2% of class, 70% of its 260 rows; 4: hour 8.2% of class, 88% of its 208 rows; 0: to confirm that the 8.2% of class, 85% of its 116 rows; 1: no rush at all 8.1% of class, 75% of its 92 rows; 0: am writing to confirm 8.0% of class, 81% of its 118 rows; 4: lose 7.9% of class, 78% of its 224 rows; 0: writing to confirm that 7.7% of class, 82% of its 112 rows; 0: you know that 7.7% of class, 70% of its 131 rows; 4: 00 today 7.2% of class, 73% of its 220 rows; 4: court 7.1% of class, 76% of its 208 rows; 4: unacceptable 7.0% of class, 73% of its 213 rows; 4: noon 6.6% of class, 83% of its 178 rows; 4: consumer 6.5% of class, 70% of its 206 rows; 4: lawyer 6.2% of class, 74% of its 187 rows; 4: shaking 6.0% of class, 96% of its 140 rows; 4: filing a 6.0% of class, 71% of its 187 rows; 0: subject thank you for 5.9% of class, 77% of its 92 rows; 4: midnight 5.7% of class, 79% of its 161 rows; 4: this immediately 5.7% of class, 92% of its 138 rows; 4: is completely 5.5% of class, 77% of its 160 rows; 4: complaint with 5.5% of class, 82% of its 148 rows; 4: begging 5.4% of class, 93% of its 129 rows; 4: this is a 5.4% of class, 86% of its 138 rows; 4: solicitor 5.3% of class, 70% of its 167 rows; 4: risk 5.2% of class, 77% of its 151 rows; 4: pm today 5.2% of class, 80% of its 143 rows; 4: calling 5.1% of class, 80% of its 141 rows
  • state.channel = web form (3,009 rows): 4: immediately 23.8% of class, 89% of its 280 rows; 4: right now 20.3% of class, 78% of its 271 rows; 4: this is 19.6% of class, 71% of its 289 rows; 4: fix 19.5% of class, 70% of its 288 rows; 4: deadline 19.0% of class, 76% of its 261 rows; 4: fix this 13.3% of class, 78% of its 177 rows; 4: legal 10.8% of class, 73% of its 153 rows; 0: you know 9.6% of class, 73% of its 66 rows; 1: rush at all 9.2% of class, 70% of its 50 rows; 4: nobody 9.1% of class, 79% of its 120 rows; 4: tonight 9.0% of class, 78% of its 121 rows; 0: to confirm that 8.2% of class, 77% of its 53 rows; 0: to let you know 8.0% of class, 71% of its 56 rows; 0: successfully 7.8% of class, 81% of its 48 rows; 4: tomorrow morning 7.5% of class, 85% of its 92 rows; 4: pm 7.4% of class, 75% of its 103 rows; 4: 17 7.3% of class, 78% of its 98 rows; 4: court 7.3% of class, 85% of its 89 rows; 4: 00 today 6.9% of class, 83% of its 87 rows; 4: noon 6.9% of class, 85% of its 85 rows; 4: hospital 6.8% of class, 71% of its 100 rows; 4: 17 00 6.6% of class, 81% of its 85 rows; 0: a quick note 6.6% of class, 75% of its 44 rows; 0: note to 6.6% of class, 85% of its 39 rows; 0: let you know that 6.4% of class, 76% of its 42 rows; 4: or i 6.1% of class, 72% of its 87 rows; 0: to drop 6.0% of class, 73% of its 41 rows; 4: lawyer 5.7% of class, 79% of its 75 rows; 4: this immediately 5.3% of class, 96% of its 57 rows; 0: to drop a 5.2% of class, 79% of its 33 rows; 4: 5pm 5.2% of class, 86% of its 63 rows; 4: lose 5.1% of class, 90% of its 59 rows; 4: unacceptable 5.1% of class, 79% of its 67 rows; 0: a quick note to 5.0% of class, 83% of its 30 rows; 0: wanted to drop 5.0% of class, 86% of its 29 rows; 4: calling 4.9% of class, 89% of its 57 rows; 4: tomorrow at 4.9% of class, 85% of its 60 rows; 4: consumer 4.8% of class, 77% of its 65 rows; 0: am writing to confirm 4.8% of class, 77% of its 31 rows; 0: note to say 4.8% of class, 80% of its 30 rows
  • state.channel = chat (2,529 rows): 4: fix 19.1% of class, 78% of its 237 rows; 4: right 16.7% of class, 84% of its 194 rows; 4: right now 15.8% of class, 90% of its 172 rows; 4: fix this 9.6% of class, 90% of its 104 rows; 0: perfectly 9.1% of class, 73% of its 48 rows; 4: filing 8.4% of class, 81% of its 101 rows; 4: immediately 8.0% of class, 91% of its 86 rows; 4: or i 7.2% of class, 91% of its 77 rows; 4: twice 7.2% of class, 78% of its 90 rows; 4: now or 6.5% of class, 95% of its 66 rows; 4: tonight 6.5% of class, 85% of its 74 rows; 4: this now 6.2% of class, 100% of its 60 rows; 4: court 6.1% of class, 88% of its 67 rows; 4: oh 6.1% of class, 83% of its 71 rows; 4: posting 5.7% of class, 85% of its 66 rows; 2: question 5.3% of class, 71% of its 49 rows; 4: now i 5.2% of class, 77% of its 66 rows; 4: cannot 5.0% of class, 79% of its 62 rows; 4: fix this now 5.0% of class, 100% of its 49 rows; 0: say the 4.9% of class, 79% of its 24 rows; 4: calling 4.7% of class, 87% of its 53 rows; 4: filing a 4.6% of class, 85% of its 53 rows; 4: lose 4.4% of class, 96% of its 45 rows; 4: if this 4.2% of class, 77% of its 53 rows; 4: going 4.1% of class, 82% of its 49 rows; 4: or we 3.9% of class, 95% of its 40 rows; 4: deadline is 3.8% of class, 92% of its 40 rows; 4: now or i 3.8% of class, 97% of its 38 rows; 0: i wanted to say 3.6% of class, 70% of its 20 rows; 0: perfect 3.6% of class, 70% of its 20 rows; 4: midnight 3.6% of class, 92% of its 38 rows; 4: reporting 3.6% of class, 90% of its 39 rows; 4: down 3.5% of class, 83% of its 41 rows; 4: lawyer 3.5% of class, 77% of its 44 rows; 2: send me 3.5% of class, 77% of its 30 rows; 2: copy 3.3% of class, 85% of its 26 rows; 4: everywhere 3.3% of class, 89% of its 36 rows; 4: hour 3.3% of class, 97% of its 33 rows; 4: it now 3.3% of class, 97% of its 33 rows; 4: nobody 3.3% of class, 78% of its 41 rows
  • state.channel = phone transcript (1,846 rows): 4: right now 52.1% of class, 78% of its 449 rows; 1: whenever 30.2% of class, 70% of its 105 rows; 4: fix 19.3% of class, 74% of its 176 rows; 0: to leave 18.1% of class, 80% of its 87 rows; 1: rush at all 18.0% of class, 83% of its 53 rows; 1: whenever you 18.0% of class, 72% of its 61 rows; 0: say thank you 17.1% of class, 80% of its 82 rows; 4: deadline 16.0% of class, 81% of its 133 rows; 4: hello is 15.9% of class, 78% of its 138 rows; 0: wanted to leave 15.8% of class, 90% of its 68 rows; 4: look i 15.7% of class, 76% of its 140 rows; 1: no rush at all 15.5% of class, 84% of its 45 rows; 4: anyone 15.4% of class, 76% of its 137 rows; 0: leave a 15.3% of class, 71% of its 83 rows; 4: nobody 15.3% of class, 76% of its 136 rows; 0: calling to 14.8% of class, 74% of its 77 rows; 0: to say thank you 14.8% of class, 81% of its 70 rows; 4: listen 13.6% of class, 87% of its 106 rows; 4: tonight 13.6% of class, 84% of its 110 rows; 4: is anyone 13.1% of class, 94% of its 94 rows; 4: o'clock 13.1% of class, 79% of its 112 rows; 0: message to 13.0% of class, 96% of its 52 rows; 4: fix this 12.9% of class, 84% of its 103 rows; 4: please i 12.8% of class, 92% of its 93 rows; 0: to leave a 12.7% of class, 88% of its 56 rows; 4: by five 11.6% of class, 91% of its 86 rows; 4: tomorrow morning 11.4% of class, 81% of its 95 rows; 4: right now i 11.3% of class, 75% of its 101 rows; 0: perfect 11.1% of class, 81% of its 53 rows; 0: wanted to leave a 11.1% of class, 93% of its 46 rows; 4: literally 11.1% of class, 75% of its 100 rows; 4: clock 11.0% of class, 84% of its 88 rows; 0: s all 10.9% of class, 78% of its 54 rows; 4: today i 10.7% of class, 82% of its 88 rows; 0: i wanted to say 10.6% of class, 77% of its 53 rows; 4: shaking 10.5% of class, 92% of its 77 rows; 4: immediately 10.2% of class, 93% of its 74 rows; 4: to me 10.2% of class, 77% of its 90 rows; 0: message to say 10.1% of class, 100% of its 39 rows; 4: you have to 10.1% of class, 97% of its 70 rows
  • state.channel = app review (1,430 rows): 4: 1 5 stars 56.6% of class, 84% of its 315 rows; 4: right 17.7% of class, 78% of its 107 rows; 4: right now 15.1% of class, 88% of its 81 rows; 4: fix this 14.3% of class, 88% of its 76 rows; 4: immediately 13.8% of class, 88% of its 74 rows; 0: perfectly 12.6% of class, 73% of its 55 rows; 4: this is 12.3% of class, 83% of its 70 rows; 1: rush 12.2% of class, 73% of its 33 rows; 1: no rush 9.6% of class, 76% of its 25 rows; 4: tonight 8.7% of class, 79% of its 52 rows; 4: lawyer 7.7% of class, 88% of its 41 rows; 4: legal 7.7% of class, 73% of its 49 rows; 4: nobody 7.7% of class, 75% of its 48 rows; 1: whenever 7.6% of class, 75% of its 20 rows; 4: times 7.2% of class, 76% of its 45 rows; 4: this to 6.4% of class, 79% of its 38 rows; 4: unacceptable 6.4% of class, 88% of its 34 rows; 4: today i 6.2% of class, 81% of its 36 rows; 4: calling 6.0% of class, 85% of its 33 rows; 4: never 6.0% of class, 70% of its 40 rows; 4: my lawyer 5.7% of class, 90% of its 30 rows; 0: everything is 5.7% of class, 82% of its 22 rows; 4: ago 5.5% of class, 76% of its 34 rows; 4: broken 5.5% of class, 70% of its 37 rows; 4: emergency 5.5% of class, 81% of its 32 rows; 0: perfect 5.4% of class, 81% of its 21 rows; 4: tomorrow morning 5.3% of class, 83% of its 30 rows; 4: midnight 5.1% of class, 83% of its 29 rows; 4: he 4.9% of class, 72% of its 32 rows; 4: lose 4.9% of class, 82% of its 28 rows; 4: absolute 4.7% of class, 76% of its 29 rows; 4: consumer 4.7% of class, 79% of its 28 rows; 4: dont 4.7% of class, 76% of its 29 rows; 4: or we 4.7% of class, 92% of its 24 rows; 4: reporting 4.7% of class, 76% of its 29 rows; 4: this now 4.7% of class, 96% of its 23 rows; 4: court 4.5% of class, 75% of its 28 rows; 4: this to the 4.5% of class, 91% of its 23 rows; 4: expires 4.3% of class, 87% of its 23 rows; 4: 5 stars unacceptable 4.0% of class, 83% of its 23 rows
  • state.channel = social media reply (1,208 rows): 4: fix this 11.9% of class, 85% of its 54 rows; 4: deadline 11.7% of class, 73% of its 62 rows; 4: twice 11.1% of class, 75% of its 57 rows; 0: perfectly 10.7% of class, 95% of its 21 rows; 2: do i 10.5% of class, 74% of its 47 rows; 4: this now 8.0% of class, 97% of its 32 rows; 2: loving 7.5% of class, 74% of its 34 rows; 4: immediately 7.3% of class, 90% of its 31 rows; 4: fix this now 6.5% of class, 96% of its 26 rows; 2: loving the 6.3% of class, 81% of its 26 rows; 4: 5pm 6.2% of class, 86% of its 28 rows; 4: lawyer 6.0% of class, 79% of its 29 rows; 4: or i 6.0% of class, 72% of its 32 rows; 4: emailed twice 5.7% of class, 79% of its 28 rows; 4: now or 5.7% of class, 88% of its 25 rows; 4: right 5.7% of class, 73% of its 30 rows; 4: tonight 5.7% of class, 71% of its 31 rows; 2: hi i 5.4% of class, 82% of its 22 rows; 2: how do i 5.1% of class, 74% of its 23 rows; 4: court 4.9% of class, 90% of its 21 rows; 4: locked 4.9% of class, 73% of its 26 rows; 4: right now 4.9% of class, 83% of its 23 rows; 4: or we 4.7% of class, 90% of its 20 rows; 4: this is 4.7% of class, 75% of its 24 rows; 4: 00 3.9% of class, 75% of its 20 rows

Shortcut models

Predicting the label class on test (1,814 rows). Chance 20.0%, majority class ('4') 35.1%; balanced chance 20.0%.

model (logistic regression, trained on the train sample) test accuracy balanced accuracy (mean recall)
bag of words, whole state (words and word pairs) 70.2% 60.0%
bag of words, main text only (message) 72.3% 62.6%
surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) 43.7% 33.4%
surface features of the main text only 42.8% 32.7%

Strongest single surface features (logistic regression on one feature, balanced accuracy on test):

feature accuracy balanced accuracy
ends_? 35.9% 21.3%
count_? 34.5% 20.7%
chars(log) 35.2% 20.1%
words(log) 35.1% 20.0%
upper_ratio 35.1% 20.0%
digit_ratio 35.1% 20.0%
nonascii_ratio 35.1% 20.0%
emoji 35.1% 20.0%
newlines 35.1% 20.0%
html_tag 35.1% 20.0%

Other state fields alone (predicting the label class on test from one field, without the main text):

field treated as accuracy balanced accuracy
company text: bag of words / length+empty 35.1% / 35.1% 20.0% / 20.0%
channel categorical, 6 values 35.1% 20.0%

2. Duplicates and split separation

Families shared between splits

splits shared families examples
train ∩ dev 0
train ∩ calibration 0
train ∩ test 0
dev ∩ calibration 0
dev ∩ test 0
calibration ∩ test 0
  • Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
  • Train rows identical in the whole prompt (state, options, instructions): 0.
  • Identical whole prompt, different answer: 0 groups (0 rows).
  • Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)

Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):

split rows examples
dev 0 (0.0%)
calibration 0 (0.0%)
test 0 (0.0%)

Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)

  • Train: 0 near-duplicate pairs; 0 rows (0.0%) sit in 0 clusters; largest cluster 0; excess rows (cluster size − 1) 0 (0.0%).
  • Clusters with more than one label class: 0 (0 rows).
  • Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 0 (0.0%), test 0 (0.0%)

5. Junk

split empty main text main text under 10 characters
train 0 0
dev 0 0
calibration 0 0
test 0 0

Pattern scan of train main texts (count, then the share of each class's rows):

pattern rows by class
placeholder [NAME]-style 0 (0.0%)
lorem ipsum 0 (0.0%)
TODO/TBD/FIXME 0 (0.0%)
'As an AI' / refusal 1 (0.0%) 2 0.0%
chat preamble ('Here is/are...', 'Sure!') 0 (0.0%)
meta words (example/variation/message:) 0 (0.0%)
model thinking tags 0 (0.0%)
JSON/code-fence leftovers 0 (0.0%)
encoding garbage (mojibake/replacement char) 0 (0.0%)
HTML tag 511 (3.1%) 0 4.6%, 1 3.9%, 2 1.9%, 3 2.4%, 4 3.1%
HTML entity 0 (0.0%)
base64-like run (40+ chars) 1 (0.0%) 0 0.0%
URL 8 (0.0%) 0 0.2%, 1 0.1%, 2 0.0%, 4 0.0%
  • 'As an AI' / refusal: triage-c290-m140-urgency (2): Subject: Concern regarding stonework at Hrubý Rohozec ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I hope this message finds you well, though I must confes…

  • HTML tag: triage-c105-m100-urgency (0): Subject: URGENT - Confirmation of received replacement card - Account 4401 22XX XXXX 8812 ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I am writing to conf… | triage-c218-m004-urgency (0): Subject: Re: Corporate Fleet Inquiry ⏎ ⏎ Noted! As mentioned when I emailed last week, we've decided against the bulk employee plan for no… | triage-c185-m219-urgency (1): Subject: URGENT - Fwd: Missing artist catalog request ⏎ ⏎ ----- Forwarded message ----- ⏎ From: Janusz Lewandowski <j.lewandowski@interia.…

  • base64-like run (40+ chars): triage-c056-m162-urgency (0): Subject: URGENT: Your store domain will expire!!! ⏎ ⏎ Dear Beloved Manager, ⏎ ⏎ I am writing to you today with a heavy heart because your…

  • URL: triage-c133-m096-urgency (0): Subject: Strategic Partnership Proposal - Global Digital Infrastructure ⏎ ⏎ Dear Sir/Madam, ⏎ ⏎ I hope this email finds you well. My name… | triage-c146-m064-urgency (2): Oggetto: R: Richiesta informazioni piano anticipato ⏎ ⏎ ---------- Messaggio inoltrato ---------- ⏎ Da: Assistenza Clienti <assistenza@ser… | triage-c238-m151-urgency (0): Subject: Meine Bougainvillea blüht wunderschön. ⏎ ⏎ Liebes Team, ⏎ ⏎ ich wollte Ihnen nur mitteilen, dass Ihre Tipps zum Zurückschneiden …

  • Possibly cut off: 5,505 of 11,135 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: 0 52.1%, 1 54.9%, 2 53.8%, 3 50.5%, 4 43.3%

    • triage-c156-m218-urgency: …ine is maintained. ⏎ ⏎ Awaiting your direction. ⏎ ⏎ Marcus Teo ⏎ Director of Facilities, Meridian Holdings Pte Ltd ⏎ +65 6823 9…
    • triage-c089-m048-urgency: …nita, your request to revise margins for Q4 (ref PO-2291) has been approved. ⏎ ⏎ Anita Desai ⏎ Store Manager ⏎ +91 20 2555 0198
    • triage-c088-m025-urgency: …plan befor next month starts so we can add more users. Plese tell me how. ⏎ Marco Bianchi, Sales Director, +39 06 492 8173

Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):

  • ×175: "---------- Forwarded message ---------" (4 65, 0 34, 1 27)
  • ×166: "Any unauthorized review, use, disclosure, or distribution is strictly prohibited." (4 65, 0 35, 1 25)
  • ×159: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and intended solely for the use of the individu…" (2 39, 4 34, 1 34)
  • ×148: "----- Forwarded message -----" (4 85, 0 22, 1 16)
  • ×132: "I hope this message finds you well." (2 55, 1 30, 3 22)
  • ×119: "-----Original Message-----" (4 55, 0 21, 2 16)
  • ×118: "If you have received this email in error, please notify the sender immediately and delete this email from your system." (4 37, 2 26, 1 21)
  • ×116: "If you are not the intended recipient, please delete all copies and notify the sender immediately." (4 58, 0 19, 2 17)
  • ×107: "Topic: Billing and Invoices" (4 32, 0 21, 3 21)
  • ×107: "If you have received this email in error, please notify the sender immediately and delete this message from your system." (0 25, 4 25, 2 22)
  • ×91: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and may also be privileged." (4 33, 0 20, 2 18)
  • ×81: "CONFIDENTIALITY NOTICE: This email and any attachments are strictly confidential and intended solely for the addressee." (4 24, 2 19, 0 18)
  • ×66: "I hope this email finds you well." (2 27, 1 24, 3 8)
  • ×63: "----- End forwarded message -----" (4 38, 0 10, 3 5)
  • ×60: "--- Forwarded message ---" (4 31, 2 10, 0 10)

6. Samples

20 random train rows per kind: triage-urgency-samples.txt. Reading notes are in the findings above.

QA: triage-sentiment

Checked 2026-09-30 13:36 by adapters/qa/qa.py (READY file READY-triage, ).

Verdict: PASS WITH NOTES (full notes in triage-needs_human.md and triage.md)

Automatic flags (for the reviewer to judge; not all are problems)

  • lclass '2' share varies across splits by more than 5 points: train 13.7%, dev 10.2%, calibration 17.2%, test 13.0%
  • 74 strong phrase flags (see list): review whether they are meaning or leakage
  • 60 standard phrase flags (≥2% of a class, mostly that class): review

Data checked

split rows families file
train 16,633 260 train.jsonl
dev 342 5 dev.jsonl
calibration 302 5 calibration.jsonl
test 1,814 30 test.jsonl
  • Train sha256: 3aba76c53cd16a1c2822d1b5e4090d183cc77fd94a233898c53501f25c3b16c0 (READY file gives no checksum)
  • Main text field (the text the phrase and length checks use): state.message.
  • Label classes: 0, 1, 2, 3, 4 (<listed option> = one of the per-row listed options such as t3 or o12). Row kinds (source.kind): sentiment.

3. Balance

Label class share per split

lclass train dev calibration test train rows
0 35.6% 33.0% 32.8% 36.5% 5,924
1 9.8% 10.5% 7.6% 10.3% 1,636
2 13.7% 10.2% 17.2% 13.0% 2,285
3 20.8% 24.0% 21.9% 20.0% 3,452
4 20.1% 22.2% 20.5% 20.3% 3,336

Row kind share per split

kind train dev calibration test train rows
sentiment 100.0% 100.0% 100.0% 100.0% 16,633

4. Format

split row-level format problems
train none
dev none
calibration none
test none

Prompt length in tokens

split measure median p99 max > 8192
train estimate: characters / 3 (upper bound for English) 203 877 1330 0
dev estimate: characters / 3 (upper bound for English) 223 878 1086 0
calibration estimate: characters / 3 (upper bound for English) 213 971 1148 0
test estimate: characters / 3 (upper bound for English) 200 842 1161 0

State key sets (train)

keys rows
channel, company, message 16,633 (100.0%)

Instructions (train)

  • Canonical (the most common text) 70.3%, reworded 26.8% (62 distinct rewordings), none 3.0%. Target about 70 / 27 / 3.
  • Canonical text: "How does the customer feel?"
  • source instruction tag: canonical 70.3%, none 3.0%, variant-17 0.6%, variant-23 0.6%, variant-21 0.5%, variant-40 0.5%
class canonical none
0 70.4% 2.8%
1 70.7% 3.2%
2 70.3% 3.1%
3 70.5% 2.9%
4 69.5% 3.2%

1. Shortcuts

Phrase statistics and models use a label-stratified sample of 16,633 train rows; models are scored on the full test file (1,814 rows).

Text length by label class (main text, characters)

split class rows p10 median p90 mean
train 0 5924 140 421 1357 612
train 1 1636 131 415 1349 608
train 2 2285 113 377 1271 553
train 3 3452 132 433 1369 616
train 4 3336 146 561 1518 728
test 0 662 141 417 1364 614
test 1 186 104 394 1409 570
test 2 235 103 396 1315 577
test 3 363 111 414 1389 590
test 4 368 153 521 1523 733

By row kind (train): main-text length, length of the rest of the state, options

kind rows median chars mean chars median other-state chars median options
sentiment 16633 433 628 145 0

Source fields by label class (train)

Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.

Overall majority: 35.6%.

source field values purity top values → classes
tone 25 77.6% politely formal: 3 67.4%, 4 20.2%; cheerful: 3 65.8%, 4 30.8%; polite and friendly: 3 70.2%, 4 23.6%; warm: 3 67.0%, 4 30.0%; disappointed: 0 63.4%, 1 33.9%; appreciative: 3 66.4%, 4 29.7%
human_reason 7 51.7% manual: 0 45.4%, 3 22.3%; legal_safety: 0 56.4%, 3 23.5%; routine: 4 34.2%, 2 28.1%; acknowledge: 4 40.6%, 3 26.7%; escalation: 0 88.0%, 1 10.8%; no_reply: 4 52.3%, 3 24.1%
situation 17 47.7% null: 0 34.0%, 3 21.4%; the writer reports or confirms something that needs no decision, such as that a delivery arrived, a payment was made, a problem sorted itself out, or a form was sent: 4 43.2%, 3 26.2%; the writer is about to take it further and says what they will do next, for example cancel, go to a regulator or ombudsman, post a public review, contact the press, or make a formal complain…
other_kind 9 38.2% null: 0 38.1%, 3 20.4%; a job application or question about working there: 4 35.6%, 3 29.2%; a message clearly meant for a different organisation: 4 51.9%, 2 16.6%; a sales pitch from another business offering its services: 2 36.2%, 3 23.7%; a request to sponsor or donate to a local event: 4 40.7%, 3 25.9%; a journalist or student asking for an interview or information for a project: 4 37.9%, 3 3…
channel 6 35.6% email: 0 34.6%, 4 22.0%; web form: 0 37.6%, 3 20.5%; chat: 0 36.1%, 3 22.1%; phone transcript: 0 33.0%, 3 19.7%; app review: 0 37.1%, 3 21.7%; social media reply: 0 37.4%, 3 21.4%
language 11 35.6% English: 0 35.7%, 3 20.9%; Danish: 0 33.3%, 4 22.4%; Czech: 0 35.0%, 3 21.5%; French: 0 37.1%, 4 22.9%; Swedish: 0 31.3%, 3 23.9%; Dutch: 0 36.7%, 3 19.6%
length_target 17 35.6% 60: 0 37.7%, 3 20.9%; 120: 0 34.8%, 4 23.6%; 25: 0 37.2%, 3 21.1%; 220: 0 32.0%, 4 27.5%; 50: 0 34.2%, 3 21.0%; 10: 0 36.2%, 2 19.4%
prompt_version 3 35.6% 3: 0 36.3%, 4 21.5%; 2: 0 33.8%, 3 21.5%; 1: 0 42.2%, 3 17.2%
decorrelated_by 2 35.6% qwen3.8-flash: 0 31.9%, 4 25.3%; null: 0 39.6%, 3 19.5%
regenerated 2 35.6% false: 0 35.1%, 3 20.7%; true: 0 37.2%, 4 21.0%
target_kind 2 35.6% soft: 0 33.4%, 3 24.1%; hard: 0 38.4%, 4 21.5%

Formatting by label class (main text, share of rows)

feature 0 1 2 3 4
ends with ? 3.6% 7.3% 7.7% 4.0% 2.6%
ends with . 43.3% 37.5% 36.1% 33.4% 34.0%
ends with ! 13.2% 14.1% 14.8% 13.0% 12.7%
no end punctuation 39.8% 41.1% 41.3% 49.1% 50.2%
starts lowercase 7.0% 7.9% 7.6% 5.6% 4.9%
all lowercase 4.0% 5.0% 5.2% 4.3% 4.0%
has a digit 91.0% 86.6% 87.6% 87.1% 84.7%
has newline 75.9% 74.6% 75.4% 77.7% 80.4%
has quotes 9.9% 5.3% 3.9% 4.3% 3.4%
has markup (HTML/markdown) 2.7% 2.1% 2.7% 3.1% 4.8%
has URL 0.0% 0.1% 0.1% 0.0% 0.1%
non-ASCII 53.0% 47.6% 42.5% 53.0% 53.9%
non-Latin script 0.0% 0.0% 0.0% 0.0% 0.0%
emoji 4.9% 4.9% 3.2% 7.0% 6.5%
ALL-CAPS word (4+) 16.3% 16.6% 15.6% 15.8% 16.5%
contains ' - ' or — 18.8% 13.4% 12.4% 18.0% 20.3%

Same, by row kind

feature sentiment
ends with ? 4%
ends with . 38%
ends with ! 13%
no end punctuation 44%
starts lowercase 6%
all lowercase 4%
has a digit 88%
has newline 77%
has quotes 6%
has markup (HTML/markdown) 3%
has URL 0%
non-ASCII 51%
non-Latin script 0%
emoji 5%
ALL-CAPS word (4+) 16%
contains ' - ' or — 18%

Over-represented words and phrases per label class (main text)

Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.

Words, 0: fix z=33 22.4% vs 2.8%; now z=32 49.4% vs 20.9%; today z=30 44.2% vs 18.4%; immediately z=27 22.1% vs 6.2%; right z=25 24.5% vs 8.6%; cannot z=25 17.8% vs 4.9%; tomorrow z=22 21.0% vs 8.1%; nobody z=21 8.8% vs 0.9%; if z=20 43.5% vs 26.0%; will z=20 25.6% vs 12.5%; filing z=19 9.6% vs 2.2%; this z=19 65.1% vs 44.4%; money z=19 9.6% vs 2.4%; twice z=19 8.6% vs 1.8%; tonight z=18 8.7% vs 2.0%; completely z=17 12.0% vs 4.4%; oh z=17 10.8% vs 3.7%; lawyer z=17 5.9% vs 0.9%; lose z=17 5.9% vs 0.9%; consumer z=16 5.5% vs 0.6%

Words, 1: why z=15 14.1% vs 4.5%; disappointed z=13 4.8% vs 0.7%; frustrating z=13 3.7% vs 0.3%; still z=13 18.1% vs 8.2%; disappointing z=12 3.1% vs 0.2%; quite z=12 7.3% vs 2.1%; anyway z=12 13.9% vs 5.9%; annoying z=11 2.5% vs 0.1%; but z=11 41.0% vs 27.0%; confusing z=10 2.3% vs 0.2%; took z=10 5.4% vs 1.6%; yeah z=9 9.8% vs 4.4%; confused z=9 3.5% vs 0.8%; long z=9 4.5% vs 1.4%; not z=8 30.5% vs 20.7%; honestly z=8 7.0% vs 3.0%; need z=8 24.8% vs 16.2%; annoyed z=8 1.5% vs 0.0%; weeks z=8 6.8% vs 3.0%; finally z=8 5.3% vs 2.1%

Words, 2: advise z=14 5.3% vs 1.2%; confirm z=13 12.3% vs 5.8%; hey z=12 4.5% vs 1.2%; current z=12 6.7% vs 2.5%; standard z=11 8.0% vs 3.6%; regarding z=11 15.8% vs 9.7%; noting z=11 2.1% vs 0.2%; require z=10 4.3% vs 1.4%; clarify z=10 2.2% vs 0.4%; noted z=10 2.9% vs 0.7%; form z=9 7.4% vs 3.7%; required z=9 4.6% vs 1.8%; documentation z=9 3.4% vs 1.1%; sure z=9 7.7% vs 4.0%; 3 z=9 12.5% vs 8.0%; madam z=9 6.2% vs 3.0%; fine z=9 6.8% vs 3.5%; additionally z=9 2.7% vs 0.8%; sir z=9 6.3% vs 3.2%; further z=8 4.2% vs 1.8%

Words, 3: could z=30 36.5% vs 11.6%; hope z=30 19.0% vs 2.7%; much z=26 37.1% vs 14.3%; thanks z=24 26.9% vs 9.3%; appreciate z=22 13.6% vs 3.0%; kindly z=22 10.3% vs 1.4%; love z=20 14.5% vs 4.2%; lovely z=20 10.5% vs 2.3%; warm z=19 9.5% vs 1.9%; team z=18 28.9% vs 14.3%; great z=18 10.2% vs 2.6%; having z=17 11.0% vs 3.2%; thank z=17 36.6% vs 20.5%; good z=17 11.1% vs 3.5%; so z=17 51.7% vs 32.5%; always z=16 10.2% vs 3.1%; would z=16 17.7% vs 7.8%; well z=16 13.0% vs 4.9%; hello z=15 27.1% vs 15.0%; finds z=15 5.2% vs 0.8%

Words, 4: thank z=37 58.2% vs 15.2%; much z=36 48.6% vs 11.6%; absolutely z=34 29.5% vs 4.0%; wonderful z=30 27.7% vs 5.1%; thrilled z=30 19.9% vs 1.0%; such z=30 20.3% vs 2.2%; happy z=29 20.6% vs 2.8%; so z=29 69.8% vs 28.1%; perfectly z=28 16.6% vs 1.2%; warmest z=27 16.9% vs 1.6%; amazing z=26 15.4% vs 0.8%; team z=25 37.2% vs 12.4%; everything z=24 29.5% vs 8.9%; best z=24 17.5% vs 3.4%; say z=22 21.8% vs 6.2%; incredibly z=21 10.0% vs 0.8%; made z=20 12.1% vs 2.1%; quick z=20 13.4% vs 2.9%; guys z=19 14.4% vs 3.8%; grateful z=18 11.1% vs 2.3%

2–4-word phrases, 0: right now z=33 22.4% vs 3.8%; this is z=26 20.1% vs 6.3%; fix this z=25 14.3% vs 0.6%; if this z=23 11.0% vs 2.1%; i cannot z=20 9.8% vs 2.4%; today i z=20 8.5% vs 1.8%; i will z=19 11.4% vs 3.7%; is not z=19 9.7% vs 2.6%; tomorrow morning z=18 6.6% vs 1.2%; this to z=17 5.8% vs 0.6%; filing a z=17 5.8% vs 0.7%; or i z=16 6.4% vs 0.2%; my lawyer z=15 4.4% vs 0.5%; complaint with z=15 4.7% vs 0.2%; is completely z=15 4.2% vs 0.4%; a formal z=14 8.0% vs 3.2%; or we z=14 3.9% vs 0.5%; if i z=14 7.2% vs 2.7%; i will be z=14 3.9% vs 0.6%; filing a formal z=14 4.0% vs 0.7%

2–4-word phrases, 1: 2 5 z=11 6.5% vs 1.8%; it took z=11 2.9% vs 0.1%; 2 5 stars z=11 5.8% vs 1.6%; understand why z=10 2.8% vs 0.5%; hi yeah z=9 3.2% vs 0.7%; need to know z=9 4.4% vs 1.3%; to know if z=9 3.4% vs 0.8%; need to know if z=9 2.6% vs 0.5%; am quite z=9 1.8% vs 0.0%; i am quite z=9 1.8% vs 0.0%; caller um hi yeah z=9 2.6% vs 0.5%; um hi yeah z=9 2.6% vs 0.5%; anyway the z=9 3.1% vs 0.7%; but the z=8 7.4% vs 3.3%; need to z=8 12.1% vs 6.6%; caller um z=8 7.1% vs 3.2%; caller um hi z=8 4.0% vs 1.3%; can you z=8 5.0% vs 1.9%; um hi z=8 4.2% vs 1.4%; i need to z=8 6.9% vs 3.2%

2–4-word phrases, 2: not sure z=16 5.0% vs 0.7%; 3 5 stars z=16 4.9% vs 0.7%; 3 5 z=15 5.3% vs 1.0%; sure if z=13 3.0% vs 0.4%; please advise z=12 3.2% vs 0.6%; no further z=12 2.6% vs 0.3%; need to z=11 11.3% vs 6.5%; not sure if z=11 2.1% vs 0.3%; please confirm z=11 3.2% vs 0.8%; you for your time z=11 2.0% vs 0.2%; dear sir z=10 6.3% vs 2.9%; am not sure z=10 1.9% vs 0.2%; i am not sure z=10 1.9% vs 0.2%; do i z=10 4.7% vs 1.9%; to confirm z=10 5.6% vs 2.6%; i need to z=10 6.1% vs 3.1%; or if z=9 3.9% vs 1.6%; advise on z=9 1.6% vs 0.2%; madam i z=9 5.4% vs 2.7%; i am not z=9 3.9% vs 1.6%

2–4-word phrases, 3: so much z=25 30.5% vs 10.7%; thanks so much z=24 11.8% vs 1.3%; thanks so z=24 11.8% vs 1.3%; i hope z=23 13.1% vs 2.1%; could you z=22 19.5% vs 6.0%; hope you z=22 10.1% vs 1.3%; much for your z=21 9.3% vs 1.0%; having a z=21 9.3% vs 1.1%; for your z=20 18.2% vs 6.0%; much for z=19 16.7% vs 5.5%; warm regards z=18 8.1% vs 1.3%; so much for your z=18 6.9% vs 0.8%; i hope you z=17 6.8% vs 1.1%; so much for z=17 13.7% vs 4.8%; you kindly z=16 5.7% vs 0.7%; love the z=16 5.4% vs 0.7%; thank you z=16 35.8% vs 19.9%; hope you are z=16 5.8% vs 0.9%; 4 5 z=16 5.6% vs 0.8%; team i hope z=16 5.4% vs 0.7%

2–4-word phrases, 4: thank you z=33 56.3% vs 15.0%; so much z=32 40.3% vs 8.4%; you so z=31 28.4% vs 4.2%; thank you so z=31 27.9% vs 3.9%; you so much z=29 25.6% vs 3.9%; thank you so much z=29 25.3% vs 3.8%; to say z=25 18.6% vs 2.6%; such a z=24 13.8% vs 1.3%; much for z=22 20.8% vs 4.6%; so much for z=22 18.4% vs 3.7%; again for z=21 11.2% vs 1.0%; you so much for z=21 14.6% vs 2.5%; team i z=21 16.4% vs 3.3%; warmest regards z=20 10.3% vs 1.0%; so happy z=19 9.1% vs 0.6%; absolutely thrilled z=19 9.5% vs 0.3%; i wanted to z=19 11.2% vs 1.8%; i wanted z=19 11.3% vs 1.8%; team i am z=19 9.5% vs 1.1%; guys are z=18 8.7% vs 0.9%

Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):

  • 4: much 48.6% vs 11.6%
  • 4: so much 40.3% vs 8.4%
  • 4: absolutely 29.5% vs 4.0%
  • 4: you so 28.4% vs 4.2%
  • 4: thank you so 27.9% vs 3.9%
  • 4: wonderful 27.7% vs 5.1%
  • 4: you so much 25.6% vs 3.9%
  • 4: thank you so much 25.3% vs 3.8%
  • 0: fix 22.4% vs 2.8%
  • 0: right now 22.4% vs 3.8%
  • 4: much for 20.8% vs 4.6%
  • 4: happy 20.6% vs 2.8%
  • 4: such 20.3% vs 2.2%
  • 4: thrilled 19.9% vs 1.0%
  • 3: hope 19.0% vs 2.7%
  • 4: to say 18.6% vs 2.6%
  • 4: so much for 18.4% vs 3.7%
  • 4: best 17.5% vs 3.4%
  • 4: warmest 16.9% vs 1.6%
  • 4: perfectly 16.6% vs 1.2%
  • 4: team i 16.4% vs 3.3%
  • 4: amazing 15.4% vs 0.8%
  • 4: you so much for 14.6% vs 2.5%
  • 0: fix this 14.3% vs 0.6%
  • 4: such a 13.8% vs 1.3%
  • 3: appreciate 13.6% vs 3.0%
  • 4: quick 13.4% vs 2.9%
  • 3: i hope 13.1% vs 2.1%
  • 4: made 12.1% vs 2.1%
  • 3: thanks so 11.8% vs 1.3%
  • 3: thanks so much 11.8% vs 1.3%
  • 4: i wanted 11.3% vs 1.8%
  • 4: i wanted to 11.2% vs 1.8%
  • 4: again for 11.2% vs 1.0%
  • 4: grateful 11.1% vs 2.3%
  • 0: if this 11.0% vs 2.1%
  • 3: lovely 10.5% vs 2.3%
  • 4: warmest regards 10.3% vs 1.0%
  • 3: kindly 10.3% vs 1.4%
  • 3: hope you 10.1% vs 1.3%

Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):

  • all rows (16,633 rows): 0: fix 22.4% of class, 82% of its 1626 rows; 0: right now 22.4% of class, 77% of its 1732 rows; 4: such 20.3% of class, 70% of its 963 rows; 4: thrilled 19.9% of class, 83% of its 799 rows; 4: warmest 16.9% of class, 73% of its 772 rows; 4: perfectly 16.6% of class, 77% of its 717 rows; 4: amazing 15.4% of class, 84% of its 615 rows; 0: fix this 14.3% of class, 93% of its 911 rows; 4: such a 13.8% of class, 73% of its 628 rows; 3: thanks so much 11.8% of class, 70% of its 579 rows; 4: again for 11.2% of class, 75% of its 500 rows; 0: if this 11.0% of class, 74% of its 872 rows; 4: warmest regards 10.3% of class, 71% of its 482 rows; 4: incredibly 10.0% of class, 75% of its 447 rows; 4: absolutely thrilled 9.5% of class, 87% of its 362 rows; 3: much for your 9.3% of class, 71% of its 454 rows; 4: so happy 9.1% of class, 80% of its 379 rows; 0: nobody 8.8% of class, 85% of its 619 rows; 0: twice 8.6% of class, 73% of its 698 rows; 4: you guys are 8.5% of class, 71% of its 400 rows; 0: today i 8.5% of class, 73% of its 690 rows; 4: you again 8.2% of class, 74% of its 367 rows; 4: thank you again 7.8% of class, 82% of its 316 rows; 4: incredible 6.9% of class, 94% of its 245 rows; 4: perfect 6.9% of class, 80% of its 287 rows; 4: gratitude 6.7% of class, 92% of its 242 rows; 4: the best 6.7% of class, 82% of its 273 rows; 0: tomorrow morning 6.6% of class, 76% of its 513 rows; 0: unacceptable 6.4% of class, 99% of its 383 rows; 0: or i 6.4% of class, 94% of its 400 rows; 4: thank you again for 6.2% of class, 81% of its 257 rows; 4: thrilled with 5.9% of class, 86% of its 230 rows; 0: lawyer 5.9% of class, 79% of its 441 rows; 0: lose 5.9% of class, 78% of its 445 rows; 0: filing a 5.8% of class, 81% of its 421 rows; 0: this to 5.8% of class, 84% of its 407 rows; 0: consumer 5.5% of class, 85% of its 381 rows; 4: for making 5.5% of class, 85% of its 215 rows; 0: hour 5.4% of class, 83% of its 384 rows; 0: 1 5 stars 5.3% of class, 99% of its 315 rows
  • state.channel = email (6,611 rows): 4: warmest 34.2% of class, 73% of its 680 rows; 4: such 30.4% of class, 70% of its 628 rows; 4: thrilled 26.7% of class, 82% of its 474 rows; 0: right now 25.0% of class, 78% of its 736 rows; 0: this is 24.5% of class, 72% of its 780 rows; 0: fix 24.4% of class, 84% of its 664 rows; 4: warmest regards 21.2% of class, 71% of its 434 rows; 4: such a 20.0% of class, 72% of its 402 rows; 4: perfectly 19.7% of class, 76% of its 374 rows; 4: again for 19.2% of class, 75% of its 374 rows; 3: much for your 16.7% of class, 72% of its 310 rows; 0: fix this 16.3% of class, 94% of its 397 rows; 4: incredibly 15.7% of class, 73% of its 314 rows; 4: absolutely thrilled 14.9% of class, 86% of its 251 rows; 0: if this 14.6% of class, 72% of its 464 rows; 4: you again 13.3% of class, 71% of its 270 rows; 4: gratitude 13.1% of class, 94% of its 202 rows; 4: thank you again 12.6% of class, 79% of its 232 rows; 4: amazing 12.4% of class, 85% of its 213 rows; 0: today i 12.2% of class, 71% of its 393 rows; 0: oh 11.9% of class, 83% of its 328 rows; 4: thank you again for 11.1% of class, 77% of its 208 rows; 0: tonight 10.5% of class, 71% of its 339 rows; 3: thanks so much 10.3% of class, 73% of its 190 rows; 4: i am so 10.1% of class, 72% of its 204 rows; 0: nobody 9.8% of class, 85% of its 264 rows; 4: dear team i am 9.3% of class, 73% of its 186 rows; 4: subject thank you 9.2% of class, 94% of its 142 rows; 0: unacceptable 9.2% of class, 99% of its 213 rows; 4: incredible 8.8% of class, 93% of its 138 rows; 4: a quick 8.5% of class, 72% of its 171 rows; 0: breach 8.4% of class, 85% of its 226 rows; 4: so happy 8.3% of class, 80% of its 151 rows; 4: to share 8.3% of class, 77% of its 158 rows; 0: tomorrow morning 8.3% of class, 73% of its 260 rows; 4: delighted 8.2% of class, 75% of its 161 rows; 0: expect 8.2% of class, 81% of its 231 rows; 0: this to 8.1% of class, 81% of its 230 rows; 4: i am absolutely thrilled 8.1% of class, 91% of its 129 rows; 4: for making 8.0% of class, 88% of its 133 rows
  • state.channel = web form (3,009 rows): 4: such 22.6% of class, 74% of its 184 rows; 4: thrilled 22.1% of class, 85% of its 157 rows; 0: immediately 22.0% of class, 89% of its 280 rows; 0: this is 21.0% of class, 82% of its 289 rows; 3: hope 20.8% of class, 76% of its 168 rows; 0: fix 20.4% of class, 80% of its 288 rows; 4: to say 19.7% of class, 77% of its 155 rows; 0: right now 18.8% of class, 78% of its 271 rows; 4: perfectly 16.9% of class, 79% of its 129 rows; 4: such a 16.1% of class, 80% of its 121 rows; 3: i hope 15.3% of class, 75% of its 126 rows; 0: fix this 14.0% of class, 89% of its 177 rows; 0: if this 12.6% of class, 79% of its 180 rows; 3: hope you 12.2% of class, 74% of its 102 rows; 4: amazing 12.1% of class, 74% of its 98 rows; 3: having a 11.7% of class, 77% of its 94 rows; 4: again for 10.8% of class, 74% of its 88 rows; 4: absolutely thrilled 10.6% of class, 91% of its 70 rows; 4: warmest 10.3% of class, 73% of its 85 rows; 0: nobody 9.5% of class, 89% of its 120 rows; 0: message my 9.2% of class, 83% of its 125 rows; 0: oh 8.9% of class, 87% of its 116 rows; 4: i am so 8.6% of class, 76% of its 68 rows; 0: tonight 8.1% of class, 76% of its 121 rows; 0: message oh 7.9% of class, 89% of its 100 rows; 4: thrilled with 7.6% of class, 94% of its 49 rows; 4: the best 7.5% of class, 80% of its 56 rows; 0: or i 7.4% of class, 97% of its 87 rows; 4: you again 7.3% of class, 88% of its 50 rows; 0: tomorrow morning 6.8% of class, 84% of its 92 rows; 4: thank you again 6.8% of class, 93% of its 44 rows; 4: to let you know 6.8% of class, 73% of its 56 rows; 4: so happy 6.6% of class, 73% of its 55 rows; 4: team i am 6.6% of class, 77% of its 52 rows; 4: thrilled to 6.3% of class, 73% of its 52 rows; 0: fixed 6.2% of class, 76% of its 92 rows; 3: your help 6.2% of class, 72% of its 53 rows; 4: i am absolutely thrilled 6.1% of class, 90% of its 41 rows; 4: you guys are 6.1% of class, 77% of its 48 rows; 0: filing a 5.9% of class, 77% of its 87 rows
  • state.channel = chat (2,529 rows): 0: now 53.8% of class, 73% of its 676 rows; 0: fix 21.6% of class, 83% of its 237 rows; 4: amazing 20.3% of class, 84% of its 102 rows; 0: right 16.4% of class, 77% of its 194 rows; 4: absolutely 16.1% of class, 86% of its 79 rows; 0: right now 15.9% of class, 84% of its 172 rows; 4: you guys are 14.7% of class, 74% of its 84 rows; 4: best 14.4% of class, 72% of its 85 rows; 0: this is 11.6% of class, 78% of its 136 rows; 4: so happy 11.6% of class, 78% of its 63 rows; 0: fix this 10.6% of class, 93% of its 104 rows; 3: kindly 9.3% of class, 71% of its 73 rows; 4: perfectly 8.7% of class, 77% of its 48 rows; 3: love the 8.6% of class, 71% of its 68 rows; 0: filing 8.5% of class, 77% of its 101 rows; 0: immediately 8.5% of class, 91% of its 86 rows; 0: twice 8.5% of class, 87% of its 90 rows; 0: or i 8.0% of class, 95% of its 77 rows; 3: hope 7.3% of class, 79% of its 52 rows; 0: now or 7.1% of class, 98% of its 66 rows; 3: you kindly 7.0% of class, 72% of its 54 rows; 0: oh 6.9% of class, 89% of its 71 rows; 0: posting 6.7% of class, 92% of its 66 rows; 0: this now 6.5% of class, 98% of its 60 rows; 4: i wanted to 6.4% of class, 71% of its 38 rows; 4: i am so 6.1% of class, 74% of its 35 rows; 4: you guys are amazing 6.1% of class, 90% of its 29 rows; 4: the best 5.9% of class, 81% of its 31 rows; 0: cannot 5.5% of class, 81% of its 62 rows; 4: thrilled 5.4% of class, 88% of its 26 rows; 3: having 5.4% of class, 73% of its 41 rows; 0: fix this now 5.4% of class, 100% of its 49 rows; 0: now i 5.4% of class, 74% of its 66 rows; 0: calling 5.0% of class, 87% of its 53 rows; 3: having a 4.8% of class, 90% of its 30 rows; 0: filing a 4.8% of class, 83% of its 53 rows; 0: i cant 4.7% of class, 77% of its 56 rows; 0: if this 4.7% of class, 81% of its 53 rows; 4: so happy with 4.5% of class, 83% of its 23 rows; 0: lose 4.4% of class, 89% of its 45 rows
  • state.channel = phone transcript (1,846 rows): 4: amazing 27.9% of class, 86% of its 115 rows; 4: oh my 27.0% of class, 96% of its 100 rows; 3: thanks so much 22.6% of class, 73% of its 113 rows; 0: fix 22.5% of class, 78% of its 176 rows; 4: gosh 21.7% of class, 92% of its 84 rows; 4: thrilled 21.4% of class, 84% of its 91 rows; 4: say thank you 20.6% of class, 89% of its 82 rows; 0: cannot 19.2% of class, 77% of its 151 rows; 4: perfectly 18.6% of class, 73% of its 90 rows; 0: nobody 18.4% of class, 82% of its 136 rows; 0: hello is 18.0% of class, 80% of its 138 rows; 0: look i 18.0% of class, 79% of its 140 rows; 4: best 17.7% of class, 80% of its 79 rows; 4: to say thank you 17.2% of class, 87% of its 70 rows; 0: anyone 17.0% of class, 76% of its 137 rows; 4: oh my gosh 16.9% of class, 98% of its 61 rows; 4: so so 16.6% of class, 97% of its 61 rows; 0: fix this 15.4% of class, 91% of its 103 rows; 4: just so 15.2% of class, 75% of its 72 rows; 0: if this 14.6% of class, 73% of its 122 rows; 0: is anyone 14.6% of class, 95% of its 94 rows; 0: please i 14.3% of class, 94% of its 93 rows; 0: listen 13.3% of class, 76% of its 106 rows; 4: perfect 13.2% of class, 89% of its 53 rows; 4: i am just 12.4% of class, 80% of its 55 rows; 0: right now i 12.1% of class, 73% of its 101 rows; 0: shaking 12.1% of class, 96% of its 77 rows; 4: i wanted to say 12.1% of class, 81% of its 53 rows; 0: today i 12.0% of class, 83% of its 88 rows; 4: say thank you so 11.5% of class, 91% of its 45 rows; 4: so happy 11.5% of class, 79% of its 52 rows; 4: incredibly 11.3% of class, 83% of its 48 rows; 4: thank you so so 11.3% of class, 98% of its 41 rows; 4: the best 11.3% of class, 80% of its 50 rows; 4: you so so much 11.3% of class, 98% of its 41 rows; 0: to me 11.1% of class, 76% of its 90 rows; 0: you have to 11.1% of class, 97% of its 70 rows; 0: by five 10.8% of class, 77% of its 86 rows; 0: i cannot 10.5% of class, 77% of its 83 rows; 3: ever so 10.5% of class, 84% of its 45 rows
  • state.channel = app review (1,430 rows): 0: 1 5 stars 58.8% of class, 99% of its 315 rows; 4: 5 5 stars 51.5% of class, 90% of its 168 rows; 4: absolutely 23.7% of class, 71% of its 98 rows; 0: fix 22.0% of class, 77% of its 151 rows; 3: 5 stars great 16.1% of class, 76% of its 66 rows; 4: best 15.3% of class, 88% of its 51 rows; 4: perfectly 14.9% of class, 80% of its 55 rows; 4: amazing 14.2% of class, 88% of its 48 rows; 0: fix this 13.9% of class, 97% of its 76 rows; 4: 5 stars absolutely 13.6% of class, 82% of its 49 rows; 0: right now 12.8% of class, 84% of its 81 rows; 3: 5 stars lovely 12.5% of class, 72% of its 54 rows; 0: immediately 12.4% of class, 89% of its 74 rows; 4: so happy 11.9% of class, 83% of its 42 rows; 3: 4 5 stars great 11.6% of class, 80% of its 45 rows; 3: could you 11.6% of class, 73% of its 49 rows; 3: kindly 11.6% of class, 82% of its 44 rows; 4: 5 5 stars absolutely 11.2% of class, 100% of its 33 rows; 0: this is 10.7% of class, 81% of its 70 rows; 3: overall 10.0% of class, 72% of its 43 rows; 4: thrilled 9.2% of class, 93% of its 29 rows; 4: ever 8.5% of class, 78% of its 32 rows; 0: filing 8.5% of class, 80% of its 56 rows; 4: such a 8.5% of class, 76% of its 33 rows; 3: 4 5 stars lovely 8.4% of class, 74% of its 35 rows; 0: money 8.1% of class, 75% of its 57 rows; 0: oh 7.7% of class, 93% of its 44 rows; 0: nobody 7.5% of class, 83% of its 48 rows; 4: 5 stars absolutely brilliant 7.1% of class, 95% of its 22 rows; 0: this to 6.8% of class, 95% of its 38 rows; 0: times 6.8% of class, 80% of its 45 rows; 4: 5 stars amazing 6.4% of class, 90% of its 21 rows; 0: lawyer 6.4% of class, 83% of its 41 rows; 0: unacceptable 6.4% of class, 100% of its 34 rows; 3: thanks so much 6.1% of class, 79% of its 24 rows; 4: with how 6.1% of class, 82% of its 22 rows; 0: broken 5.8% of class, 84% of its 37 rows; 0: if this 5.8% of class, 79% of its 39 rows; 0: 5 stars brilliant 5.6% of class, 86% of its 35 rows; 0: never 5.6% of class, 75% of its 40 rows
  • state.channel = social media reply (1,208 rows): 0: fix 19.9% of class, 82% of its 110 rows; 0: or 18.8% of class, 80% of its 106 rows; 4: amazing 17.1% of class, 90% of its 39 rows; 4: absolutely 11.7% of class, 83% of its 29 rows; 4: happy 11.7% of class, 83% of its 29 rows; 4: thank you so much 11.7% of class, 80% of its 30 rows; 0: fix this 11.3% of class, 94% of its 54 rows; 4: you guys 11.2% of class, 82% of its 28 rows; 0: twice 11.1% of class, 88% of its 57 rows; 3: loving 10.8% of class, 82% of its 34 rows; 4: best 10.7% of class, 96% of its 23 rows; 4: much for 9.8% of class, 95% of its 21 rows; 4: you guys are 9.8% of class, 100% of its 20 rows; 4: perfectly 9.3% of class, 90% of its 21 rows; 3: loving the 8.9% of class, 88% of its 26 rows; 0: filing 8.8% of class, 83% of its 48 rows; 4: thrilled 7.8% of class, 73% of its 22 rows; 3: love the 6.9% of class, 75% of its 24 rows; 3: loved 6.9% of class, 82% of its 22 rows; 0: charged 6.9% of class, 89% of its 35 rows; 0: or i 6.9% of class, 97% of its 32 rows; 0: this now 6.9% of class, 97% of its 32 rows; 0: immediately 6.6% of class, 97% of its 31 rows; 0: oh 6.2% of class, 90% of its 31 rows; 0: fix this now 5.8% of class, 100% of its 26 rows; 0: emailed twice 5.5% of class, 89% of its 28 rows; 0: now or 5.5% of class, 100% of its 25 rows; 0: oh brilliant 5.1% of class, 100% of its 23 rows; 0: this is 5.1% of class, 96% of its 24 rows; 0: lawyer 4.9% of class, 76% of its 29 rows; 0: locked 4.9% of class, 85% of its 26 rows; 0: or we 4.4% of class, 100% of its 20 rows; 0: right now 4.0% of class, 78% of its 23 rows; 0: court 3.8% of class, 81% of its 21 rows; 0: filing a 3.5% of class, 80% of its 20 rows

Shortcut models

Predicting the label class on test (1,814 rows). Chance 20.0%, majority class ('0') 36.5%; balanced chance 20.0%.

model (logistic regression, trained on the train sample) test accuracy balanced accuracy (mean recall)
bag of words, whole state (words and word pairs) 76.0% 66.1%
bag of words, main text only (message) 77.2% 68.3%
surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) 44.5% 31.4%
surface features of the main text only 44.3% 31.2%

Strongest single surface features (logistic regression on one feature, balanced accuracy on test):

feature accuracy balanced accuracy
count_, 36.7% 20.7%
count_; 36.5% 20.1%
digit_ratio 36.5% 20.1%
url 36.5% 20.1%
count_( 36.5% 20.1%
count_) 36.5% 20.1%
newlines 36.5% 20.0%
chars(log) 36.5% 20.0%
words(log) 36.5% 20.0%
upper_ratio 36.5% 20.0%

Other state fields alone (predicting the label class on test from one field, without the main text):

field treated as accuracy balanced accuracy
company text: bag of words / length+empty 36.5% / 36.5% 20.0% / 20.0%
channel categorical, 6 values 36.5% 20.0%

2. Duplicates and split separation

Families shared between splits

splits shared families examples
train ∩ dev 0
train ∩ calibration 0
train ∩ test 0
dev ∩ calibration 0
dev ∩ test 0
calibration ∩ test 0
  • Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
  • Train rows identical in the whole prompt (state, options, instructions): 0.
  • Identical whole prompt, different answer: 0 groups (0 rows).
  • Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)

Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):

split rows examples
dev 0 (0.0%)
calibration 0 (0.0%)
test 0 (0.0%)

Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)

  • Train: 0 near-duplicate pairs; 0 rows (0.0%) sit in 0 clusters; largest cluster 0; excess rows (cluster size − 1) 0 (0.0%).
  • Clusters with more than one label class: 0 (0 rows).
  • Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 0 (0.0%), test 0 (0.0%)

5. Junk

split empty main text main text under 10 characters
train 0 0
dev 0 0
calibration 0 0
test 0 0

Pattern scan of train main texts (count, then the share of each class's rows):

pattern rows by class
placeholder [NAME]-style 0 (0.0%)
lorem ipsum 0 (0.0%)
TODO/TBD/FIXME 0 (0.0%)
'As an AI' / refusal 1 (0.0%) 1 0.1%
chat preamble ('Here is/are...', 'Sure!') 0 (0.0%)
meta words (example/variation/message:) 0 (0.0%)
model thinking tags 0 (0.0%)
JSON/code-fence leftovers 0 (0.0%)
encoding garbage (mojibake/replacement char) 0 (0.0%)
HTML tag 511 (3.1%) 0 2.7%, 1 2.0%, 2 2.6%, 3 3.0%, 4 4.7%
HTML entity 0 (0.0%)
base64-like run (40+ chars) 1 (0.0%) 1 0.1%
URL 8 (0.0%) 0 0.0%, 1 0.1%, 2 0.1%, 4 0.1%
  • 'As an AI' / refusal: triage-c290-m140-sentiment (1): Subject: Concern regarding stonework at Hrubý Rohozec ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I hope this message finds you well, though I must confes…

  • HTML tag: triage-c008-m090-sentiment (4): Subject: Fwd: Content Request - Please add Ilo Ilo to your library! ⏎ ⏎ From: Sarah Lim sarah.lim82@gmail.com ⏎ Date: 14 November 2023 a… | triage-c205-m124-sentiment (3): Subject: Fwd: Výsledek zkoušky – BUS402 ⏎ ⏎ ---------- Přeposlaná zpráva ---------- ⏎ Od: Zkouškové oddělení zkousky@edulearn.co.za ⏎ Da… | triage-c016-m054-sentiment (0): Subject: Incorrect Invoice #4409281 - Kronoberg VA - Legal action deadline tomorrow ⏎ ⏎ I am writing to you again because I have run out o…

  • base64-like run (40+ chars): triage-c056-m162-sentiment (1): Subject: URGENT: Your store domain will expire!!! ⏎ ⏎ Dear Beloved Manager, ⏎ ⏎ I am writing to you today with a heavy heart because your…

  • URL: triage-c133-m096-sentiment (2): Subject: Strategic Partnership Proposal - Global Digital Infrastructure ⏎ ⏎ Dear Sir/Madam, ⏎ ⏎ I hope this email finds you well. My name… | triage-c146-m064-sentiment (1): Oggetto: R: Richiesta informazioni piano anticipato ⏎ ⏎ ---------- Messaggio inoltrato ---------- ⏎ Da: Assistenza Clienti <assistenza@ser… | triage-c091-m045-sentiment (4): Hi. I emailed last week about your amazing stadium. I just want to say the Allsvenskan final was incredible. Best day ever. Also, can you f…

  • Possibly cut off: 5,505 of 11,135 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: 0 42.3%, 1 46.2%, 2 48.7%, 3 56.5%, 4 56.2%

    • triage-c057-m194-sentiment: …eciate everything you do for us sellers! ⏎ ⏎ Warmest regards, ⏎ Lim Boon Kiat ⏎ Owner, KB Electronics & Furniture ⏎ +65 9182 3347
    • triage-c120-m050-sentiment: …fach im System verbleibt. ⏎ ⏎ Mit freundlichen Grüßen ⏎ Thomas Weber ⏎ Einkaufsleiter, Stahlwerk Dortmund GmbH ⏎ +49 231 555 0192
    • triage-c150-m146-sentiment: …-ce que vous allez valider ma demande d'annulation de ce double paiemnt avant 17h??? ⏎ ⏎ Sophie Martin ⏎ Comptable ⏎ 0612345678

Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):

  • ×175: "---------- Forwarded message ---------" (0 64, 4 46, 3 31)
  • ×166: "Any unauthorized review, use, disclosure, or distribution is strictly prohibited." (0 59, 4 36, 3 30)
  • ×159: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and intended solely for the use of the individu…" (4 41, 0 35, 3 34)
  • ×148: "----- Forwarded message -----" (0 76, 3 28, 4 28)
  • ×132: "I hope this message finds you well." (3 100, 4 28, 2 3)
  • ×119: "-----Original Message-----" (0 60, 2 19, 4 18)
  • ×118: "If you have received this email in error, please notify the sender immediately and delete this email from your system." (0 41, 4 27, 3 22)
  • ×116: "If you are not the intended recipient, please delete all copies and notify the sender immediately." (0 55, 3 26, 2 13)
  • ×107: "If you have received this email in error, please notify the sender immediately and delete this message from your system." (4 30, 0 25, 3 24)
  • ×107: "Topic: Billing and Invoices" (0 40, 4 27, 3 19)
  • ×91: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and may also be privileged." (0 31, 3 26, 2 13)
  • ×81: "CONFIDENTIALITY NOTICE: This email and any attachments are strictly confidential and intended solely for the addressee." (0 26, 4 17, 3 16)
  • ×66: "I hope this email finds you well." (3 41, 4 16, 2 6)
  • ×63: "----- End forwarded message -----" (0 36, 4 13, 3 12)
  • ×60: "--- Forwarded message ---" (0 34, 4 9, 3 8)

6. Samples

20 random train rows per kind: triage-sentiment-samples.txt. Reading notes are in the findings above.

QA: triage-needs_human

Checked 2026-09-30 13:39 by adapters/qa/qa.py (READY file READY-triage, ).

Verdict: PASS WITH NOTES

Triage (READY-triage 2026-09-30 13:33, train sha256 3aba76c5…3b16c0; train 66,532 rows = 16,633 messages × 4 questions). This is the first full check; the earlier triage reports checked withdrawn files. Each question type has its own report: triage-route.md, triage-urgency.md, triage-sentiment.md and triage-needs_human.md. Cue rates come from triage_extra.py.

Earlier cues, now fixed (checked)

Each rate below is the share of that label's rows containing the cue, compared across labels:

cue before now
"Subject: URGENT" 22.6% of urgency-4 against 3.0% of the rest 3.0–4.4% at every urgency level
"just wanted to" 20.4% of urgency-0 6.0–6.3% at every level, and 5.9% / 6.2% for needs_human no / yes
"by tomorrow" 14.9% of urgency-2 1.8–4.1%
"formal complaint" 7.9% of person-needed against 0.1% 2.8% / 3.8%
"Hi there" 17% of sentiment-3 4.1–5.6%
ends with "!" 39% of very positive 22–29% at every sentiment level
  • "manual", "real person" and "no reply needed" appear in 0 rows.

  • Surface-only models:

    • sentiment: 31.4% balanced accuracy (was 47.8%; chance 20%);
    • urgency: 33.4% (chance 20%);
    • needs_human: 60.5% (chance 50%);
    • route: 9.1% (chance 4.2%).

    No single surface feature beats chance by more than 1.3 points.

  • Route options: the correct team is the longest option 10.6% of the time against 9.8% chance, positions are uniform, and the option picker is at chance.

  • Duplicates: there are no near duplicates in train or across splits, and no identical prompt carries two different answers.

  • The edited messages read naturally. In my sample, urgency-4 messages now contain "just wanted to" ("I just wanted to say I need my boxes from unit 12 today!"), low-urgency messages carry "Subject: Urgent" (a spam investment offer; a confirmation email), and very negative messages end with "!".

Notes (meaning; for the model card)

  1. The remaining surface signal is meaning:
    • sentiment: unmatched ")" versus "(" counts, which are emoticons like :) and :( — and message length;
    • urgency: fewer "?" in urgent messages (questions are rarely emergencies);
    • needs_human: weak length and punctuation mixtures, with no single feature above chance.
  2. The standard phrase flags are meaning words:
    • urgency 4: "right now", "fix this", "deadline", "tonight", "court", "lawyer", "breach";
    • urgency 1: "no rush at all", the customer's own statement (kept by agreement);
    • urgency 0: "confirm that", "say thank you";
    • sentiment: "thrilled", "amazing", "warmest regards", "unacceptable", "nobody";
    • needs_human no: "a quick note", "to let you know that", "confirm that the", "perfectly";
    • route "other": "my thesis", "journalism student", "sponsoring" (student and sponsorship requests belong to no team).
  3. The label mix changed with the re-rating: urgency level 4 is now 34.7% of rows, and sentiment level 0 (very negative) is 35.6%. Dev and calibration are small: 1,368 and 1,208 rows, which is 342 and 302 messages (5 organisations each), so calibrate with care.
  4. About 51% of messages had small cue edits by qwen3.8-flash, recorded in source.decorrelated_by. Score targets are the mean of three raters (the plan, qwen3.8-max and deepseek-v4-flash).

Automatic flags (for the reviewer to judge; not all are problems)

  • 25 strong phrase flags (see list): review whether they are meaning or leakage
  • 35 standard phrase flags (≥2% of a class, mostly that class): review

Data checked

split rows families file
train 16,633 260 train.jsonl
dev 342 5 dev.jsonl
calibration 302 5 calibration.jsonl
test 1,814 30 test.jsonl
  • Train sha256: 3aba76c53cd16a1c2822d1b5e4090d183cc77fd94a233898c53501f25c3b16c0 (READY file gives no checksum)
  • Main text field (the text the phrase and length checks use): state.message.
  • Label classes: False, True (<listed option> = one of the per-row listed options such as t3 or o12). Row kinds (source.kind): needs_human.

3. Balance

Label class share per split

lclass train dev calibration test train rows
False 40.4% 41.2% 40.1% 42.4% 6,719
True 59.6% 58.8% 59.9% 57.6% 9,914

Row kind share per split

kind train dev calibration test train rows
needs_human 100.0% 100.0% 100.0% 100.0% 16,633

4. Format

split row-level format problems
train none
dev none
calibration none
test none

Prompt length in tokens

split measure median p99 max > 8192
train estimate: characters / 3 (upper bound for English) 255 925 1385 0
dev estimate: characters / 3 (upper bound for English) 273 927 1136 0
calibration estimate: characters / 3 (upper bound for English) 262 1026 1203 0
test estimate: characters / 3 (upper bound for English) 253 893 1215 0

State key sets (train)

keys rows
channel, company, message 16,633 (100.0%)

Instructions (train)

  • Canonical (the most common text) 69.7%, reworded 27.1% (70 distinct rewordings), none 3.2%. Target about 70 / 27 / 3.
  • Canonical text: "Does this need a person to act on it now, rather than an automatic reply? Answer yes for complaints that could escalate, legal or safety issues, or requests an automatic system cannot resolve."
  • source instruction tag: canonical 69.7%, none 3.2%, variant-52 0.5%, variant-23 0.5%, variant-29 0.5%, variant-11 0.5%
class canonical none
False 70.0% 3.4%
True 69.5% 3.1%

1. Shortcuts

Phrase statistics and models use a label-stratified sample of 16,633 train rows; models are scored on the full test file (1,814 rows).

Text length by label class (main text, characters)

split class rows p10 median p90 mean
train False 6719 119 404 1352 596
train True 9914 145 452 1418 649
test False 770 108 403 1436 605
test True 1044 143 439 1392 638

By row kind (train): main-text length, length of the rest of the state, options

kind rows median chars mean chars median other-state chars median options
needs_human 16633 433 628 145 0

Source fields by label class (train)

Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.

Overall majority: 59.6%.

source field values purity top values → classes
human_reason 7 99.8% manual: True 100.0%; legal_safety: True 100.0%; routine: False 100.0%; acknowledge: False 100.0%; escalation: True 100.0%; no_reply: False 100.0%
situation 17 87.4% null: True 58.2%, False 41.8%; the writer reports or confirms something that needs no decision, such as that a delivery arrived, a payment was made, a problem sorted itself out, or a form was sent: False 100.0%; the writer is about to take it further and says what they will do next, for example cancel, go to a regulator or ombudsman, post a public review, contact the press, or make a formal compl…
tone 25 75.0% politely formal: False 61.8%, True 38.2%; cheerful: True 51.4%, False 48.6%; polite and friendly: False 57.6%, True 42.4%; warm: False 53.5%, True 46.5%; disappointed: True 82.1%, False 17.9%; appreciative: False 54.5%, True 45.5%
other_kind 9 65.8% null: True 63.6%, False 36.4%; a job application or question about working there: False 97.5%, True 2.5%; a message clearly meant for a different organisation: False 98.3%, True 1.7%; a sales pitch from another business offering its services: False 97.4%, True 2.6%; a request to sponsor or donate to a local event: False 92.6%, True 7.4%; a journalist or student asking for an interview or informat…
channel 6 59.6% email: True 60.8%, False 39.2%; web form: True 61.6%, False 38.4%; chat: True 58.0%, False 42.0%; phone transcript: True 59.3%, False 40.7%; app review: True 54.9%, False 45.1%; social media reply: True 57.5%, False 42.5%
language 11 59.6% English: True 60.0%, False 40.0%; Danish: True 53.0%, False 47.0%; Czech: True 55.4%, False 44.6%; French: True 56.5%, False 43.5%; Swedish: True 50.3%, False 49.7%; Dutch: True 57.0%, False 43.0%
length_target 17 59.6% 60: True 61.0%, False 39.0%; 120: True 61.7%, False 38.3%; 25: True 58.5%, False 41.5%; 220: True 62.8%, False 37.2%; 50: True 58.3%, False 41.7%; 10: True 54.3%, False 45.7%
prompt_version 3 59.6% 3: True 60.2%, False 39.8%; 2: True 57.8%, False 42.2%; 1: True 74.1%, False 25.9%
decorrelated_by 2 59.6% qwen3.8-flash: True 57.6%, False 42.4%; null: True 61.7%, False 38.3%
regenerated 2 59.6% false: True 58.0%, False 42.0%; true: True 63.9%, False 36.1%

Formatting by label class (main text, share of rows)

feature False True
ends with ? 3.8% 4.8%
ends with . 36.2% 39.0%
ends with ! 13.7% 13.1%
no end punctuation 45.9% 43.0%
starts lowercase 6.7% 6.3%
all lowercase 5.3% 3.6%
has a digit 83.9% 90.8%
has newline 76.2% 77.5%
has quotes 3.2% 8.1%
has markup (HTML/markdown) 3.6% 2.9%
has URL 0.1% 0.0%
non-ASCII 46.7% 54.3%
non-Latin script 0.0% 0.0%
emoji 5.9% 5.1%
ALL-CAPS word (4+) 16.1% 16.2%
contains ' - ' or — 14.2% 19.8%

Same, by row kind

feature needs_human
ends with ? 4%
ends with . 38%
ends with ! 13%
no end punctuation 44%
starts lowercase 6%
all lowercase 4%
has a digit 88%
has newline 77%
has quotes 6%
has markup (HTML/markdown) 3%
has URL 0%
non-ASCII 51%
non-Latin script 0%
emoji 5%
ALL-CAPS word (4+) 16%
contains ' - ' or — 18%

Over-represented words and phrases per label class (main text)

Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.

Words, False: thank z=30 36.6% vs 15.2%; thanks z=24 20.4% vs 7.9%; much z=23 27.5% vs 13.3%; perfectly z=22 9.3% vs 0.9%; so z=20 46.0% vs 30.0%; everything z=19 19.0% vs 9.0%; happy z=19 10.9% vs 3.3%; quick z=19 8.9% vs 2.3%; wanted z=18 16.3% vs 7.5%; all z=18 25.8% vs 15.0%; best z=18 10.3% vs 3.4%; anyway z=18 10.9% vs 3.9%; just z=18 36.7% vs 24.2%; rush z=18 7.5% vs 1.7%; such z=17 9.6% vs 3.2%; say z=17 13.9% vs 6.3%; went z=17 7.5% vs 2.1%; question z=16 8.6% vs 2.9%; fine z=16 7.0% vs 1.8%; regards z=16 16.6% vs 9.1%

Words, True: today z=30 39.3% vs 10.2%; before z=26 29.0% vs 7.8%; tomorrow z=22 18.7% vs 3.9%; fix z=22 15.3% vs 1.6%; if z=22 41.4% vs 18.8%; deadline z=20 13.9% vs 1.0%; cannot z=20 14.2% vs 2.6%; someone z=20 15.0% vs 3.1%; right z=19 19.4% vs 6.6%; will z=19 22.9% vs 8.8%; immediately z=18 16.5% vs 5.0%; legal z=17 9.9% vs 1.1%; look z=16 12.8% vs 4.0%; by z=16 27.5% vs 13.6%; must z=15 10.4% vs 2.9%; contract z=15 10.3% vs 3.0%; tonight z=14 7.0% vs 0.5%; need z=14 21.5% vs 10.4%; filing z=14 7.2% vs 1.3%; please z=14 32.5% vs 18.7%

2–4-word phrases, False: thank you z=28 35.5% vs 14.9%; to say z=23 11.4% vs 2.1%; i wanted z=19 7.8% vs 1.0%; i wanted to z=19 7.7% vs 1.0%; so much z=18 20.8% vs 10.7%; to confirm z=17 6.3% vs 0.8%; wanted to z=17 16.0% vs 7.5%; thank you for z=17 12.4% vs 5.0%; you for z=17 12.8% vs 5.4%; no rush z=17 5.9% vs 0.9%; writing to z=16 9.9% vs 3.7%; you know z=16 5.4% vs 0.7%; am writing to z=16 9.3% vs 3.4%; i am writing to z=16 9.3% vs 3.4%; everything is z=15 5.2% vs 0.9%; such a z=15 6.6% vs 1.9%; again for z=15 5.6% vs 1.3%; that the z=15 8.8% vs 3.4%; a quick z=15 4.4% vs 0.5%; at all z=15 7.1% vs 2.4%

2–4-word phrases, True: right now z=22 15.8% vs 2.4%; i need z=16 11.3% vs 3.3%; this is z=16 15.0% vs 5.6%; if this z=16 8.0% vs 1.1%; before the z=15 7.6% vs 1.4%; fix this z=15 9.0% vs 0.3%; i cannot z=14 7.3% vs 1.7%; will be z=14 7.7% vs 2.0%; today i z=14 6.3% vs 1.0%; is not z=13 7.2% vs 2.1%; look into z=12 4.9% vs 0.7%; tomorrow morning z=12 5.0% vs 0.2%; i will z=12 8.6% vs 3.3%; if i z=11 6.0% vs 1.9%; this to z=11 3.8% vs 0.4%; look at z=11 3.9% vs 0.7%; filing a z=11 3.8% vs 0.7%; if we z=10 3.7% vs 0.7%; within the z=10 3.4% vs 0.6%; 17 00 z=10 4.0% vs 0.1%

Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):

  • True: tomorrow 18.7% vs 3.9%
  • True: right now 15.8% vs 2.4%
  • True: fix 15.3% vs 1.6%
  • True: someone 15.0% vs 3.1%
  • True: cannot 14.2% vs 2.6%
  • True: deadline 13.9% vs 1.0%
  • False: to say 11.4% vs 2.1%
  • True: legal 9.9% vs 1.1%
  • False: perfectly 9.3% vs 0.9%
  • True: fix this 9.0% vs 0.3%
  • True: if this 8.0% vs 1.1%
  • False: i wanted 7.8% vs 1.0%
  • False: i wanted to 7.7% vs 1.0%
  • True: before the 7.6% vs 1.4%
  • False: rush 7.5% vs 1.7%
  • True: i cannot 7.3% vs 1.7%
  • True: filing 7.2% vs 1.3%
  • True: tonight 7.0% vs 0.5%
  • False: to confirm 6.3% vs 0.8%
  • True: today i 6.3% vs 1.0%
  • False: no rush 5.9% vs 0.9%
  • False: again for 5.6% vs 1.3%
  • False: you know 5.4% vs 0.7%
  • False: everything is 5.2% vs 0.9%
  • True: tomorrow morning 5.0% vs 0.2%

Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):

  • all rows (16,633 rows): False: perfectly 9.3% of class, 87% of its 717 rows; False: i wanted to 7.7% of class, 85% of its 614 rows; False: to confirm 6.3% of class, 83% of its 509 rows; False: no rush 5.9% of class, 81% of its 492 rows; False: you know 5.4% of class, 83% of its 434 rows; False: confirm that 5.1% of class, 95% of its 363 rows; False: to confirm that 4.6% of class, 97% of its 321 rows; False: a quick 4.4% of class, 85% of its 346 rows; False: bye 4.3% of class, 86% of its 338 rows; False: to let 4.0% of class, 86% of its 313 rows; False: perfect 3.6% of class, 85% of its 287 rows; False: rush at all 3.5% of class, 90% of its 264 rows; False: to let you know 3.5% of class, 93% of its 254 rows; False: successfully 3.2% of class, 88% of its 246 rows; False: no rush at all 3.1% of class, 90% of its 234 rows; False: know that 3.0% of class, 89% of its 228 rows; False: gratitude 3.0% of class, 83% of its 242 rows; False: went through 3.0% of class, 84% of its 239 rows; False: for making 2.8% of class, 87% of its 215 rows; False: say thank you 2.7% of class, 95% of its 191 rows; False: thank you for the 2.7% of class, 88% of its 205 rows; False: to share 2.7% of class, 85% of its 214 rows; False: confirm that the 2.7% of class, 97% of its 185 rows; False: let you know that 2.7% of class, 95% of its 188 rows; False: a quick note 2.5% of class, 96% of its 174 rows; False: to confirm that the 2.4% of class, 98% of its 163 rows; False: to say thank you 2.4% of class, 95% of its 167 rows; False: am writing to confirm 2.3% of class, 99% of its 156 rows; False: note to 2.3% of class, 98% of its 158 rows; False: keep up the 2.2% of class, 88% of its 170 rows; False: writing to confirm that 2.2% of class, 99% of its 148 rows; False: and everything 2.1% of class, 83% of its 169 rows; False: 5 5 stars 2.1% of class, 82% of its 168 rows; False: subject thank you 2.0% of class, 96% of its 142 rows; False: to drop 2.0% of class, 94% of its 144 rows
  • state.channel = email (6,611 rows): False: perfectly 12.1% of class, 84% of its 374 rows; False: to confirm 11.0% of class, 86% of its 329 rows; False: to say 10.1% of class, 79% of its 334 rows; False: confirm that 9.2% of class, 94% of its 252 rows; False: i wanted to 9.1% of class, 79% of its 297 rows; False: to confirm that 8.5% of class, 97% of its 226 rows; False: you know 7.4% of class, 91% of its 211 rows; False: gratitude 6.6% of class, 85% of its 202 rows; False: to let 6.5% of class, 86% of its 195 rows; False: to let you know 6.1% of class, 93% of its 168 rows; False: successfully 5.8% of class, 85% of its 176 rows; False: a quick 5.7% of class, 86% of its 171 rows; False: subject thank you 5.2% of class, 96% of its 142 rows; False: know that 5.2% of class, 87% of its 154 rows; False: to share 5.0% of class, 82% of its 158 rows; False: thank you for the 5.0% of class, 89% of its 145 rows; False: let you know that 4.7% of class, 95% of its 129 rows; False: quick note 4.6% of class, 94% of its 125 rows; False: am writing to confirm 4.5% of class, 99% of its 118 rows; False: for making 4.4% of class, 86% of its 133 rows; False: to confirm that the 4.4% of class, 98% of its 116 rows; False: writing to confirm that 4.3% of class, 99% of its 112 rows; False: note to 4.1% of class, 98% of its 108 rows; False: a quick note 4.0% of class, 96% of its 109 rows; False: beautifully 3.5% of class, 80% of its 115 rows; False: rush at all 3.5% of class, 88% of its 104 rows; False: subject thank you for 3.4% of class, 97% of its 92 rows; False: a quick note to 3.4% of class, 98% of its 90 rows; False: no rush at all 3.2% of class, 89% of its 92 rows; False: to drop 3.1% of class, 94% of its 85 rows; False: my end 3.0% of class, 84% of its 94 rows; False: and everything 2.9% of class, 88% of its 84 rows; False: gratitude for 2.9% of class, 90% of its 82 rows; False: was so 2.9% of class, 79% of its 94 rows; False: on my end 2.8% of class, 88% of its 83 rows; False: our end 2.8% of class, 84% of its 86 rows; False: deepest 2.7% of class, 83% of its 86 rows; False: sincere 2.7% of class, 82% of its 85 rows; False: keep up the 2.7% of class, 88% of its 78 rows; False: everything went 2.6% of class, 92% of its 74 rows
  • state.channel = web form (3,009 rows): False: to say 11.0% of class, 82% of its 155 rows; False: perfectly 10.1% of class, 91% of its 129 rows; False: i am writing to 9.9% of class, 79% of its 145 rows; False: i wanted to 9.4% of class, 85% of its 128 rows; False: such a 8.1% of class, 78% of its 121 rows; False: quick 7.5% of class, 78% of its 112 rows; False: rush 7.3% of class, 82% of its 102 rows; False: again for 6.6% of class, 86% of its 88 rows; False: whenever 6.6% of class, 79% of its 96 rows; False: to confirm 6.2% of class, 78% of its 91 rows; False: no rush 5.9% of class, 85% of its 80 rows; False: everything is 5.7% of class, 79% of its 84 rows; False: you know 5.2% of class, 91% of its 66 rows; False: to let 5.0% of class, 85% of its 68 rows; False: a quick 4.9% of class, 88% of its 65 rows; False: finally 4.9% of class, 77% of its 74 rows; False: smoothly 4.6% of class, 80% of its 66 rows; False: to confirm that 4.5% of class, 98% of its 53 rows; False: to let you know 4.4% of class, 91% of its 56 rows; False: share 4.2% of class, 79% of its 61 rows; False: successfully 4.1% of class, 98% of its 48 rows; False: drop 4.0% of class, 88% of its 52 rows; False: a quick note 3.7% of class, 98% of its 44 rows; False: whenever you 3.7% of class, 84% of its 51 rows; False: no rush at all 3.6% of class, 91% of its 46 rows; False: you again 3.6% of class, 84% of its 50 rows; False: let you know that 3.5% of class, 95% of its 42 rows; False: team i am 3.5% of class, 77% of its 52 rows; False: note to 3.4% of class, 100% of its 39 rows; False: to drop 3.4% of class, 95% of its 41 rows; False: went through 3.3% of class, 83% of its 46 rows; False: completed 3.2% of class, 82% of its 45 rows; False: keep up the 3.2% of class, 90% of its 41 rows; False: thank you again 3.2% of class, 84% of its 44 rows; False: the standard 3.2% of class, 77% of its 48 rows; False: warmest regards 3.2% of class, 82% of its 45 rows; False: thanks again 3.1% of class, 88% of its 41 rows; False: worked 3.1% of class, 88% of its 41 rows; False: for making 2.9% of class, 89% of its 38 rows; False: beautifully 2.9% of class, 82% of its 40 rows
  • state.channel = chat (2,529 rows): False: link 6.8% of class, 85% of its 85 rows; False: no rush 6.4% of class, 89% of its 76 rows; False: perfectly 4.2% of class, 94% of its 48 rows; False: thank you for 3.9% of class, 95% of its 43 rows; False: i wanted to 3.4% of class, 95% of its 38 rows; False: whenever 3.4% of class, 90% of its 40 rows; False: hi just 3.1% of class, 94% of its 35 rows; False: no rush at all 2.9% of class, 94% of its 33 rows; False: is there a 2.8% of class, 91% of its 33 rows; False: send me 2.6% of class, 93% of its 30 rows; False: wondering 2.4% of class, 87% of its 30 rows; False: copy 2.3% of class, 92% of its 26 rows; False: just confirming 2.2% of class, 100% of its 23 rows
  • state.channel = phone transcript (1,846 rows): False: bye 37.6% of class, 86% of its 329 rows; False: message 19.9% of class, 85% of its 177 rows; False: i wanted to 16.5% of class, 91% of its 136 rows; False: anyway i 16.2% of class, 82% of its 149 rows; False: wanted to say 15.0% of class, 82% of its 138 rows; False: just calling 12.9% of class, 86% of its 113 rows; False: to leave 10.8% of class, 93% of its 87 rows; False: perfectly 10.5% of class, 88% of its 90 rows; False: say thank you 10.4% of class, 95% of its 82 rows; False: leave a 10.0% of class, 90% of its 83 rows; False: no rush 9.8% of class, 86% of its 86 rows; False: calling to 9.4% of class, 92% of its 77 rows; False: goodbye 9.0% of class, 91% of its 75 rows; False: wanted to leave 9.0% of class, 100% of its 68 rows; False: to say thank you 8.8% of class, 94% of its 70 rows; False: a quick 7.7% of class, 87% of its 67 rows; False: to leave a 7.2% of class, 96% of its 56 rows; False: message to 6.9% of class, 100% of its 52 rows; False: perfect 6.9% of class, 98% of its 53 rows; False: s all 6.9% of class, 96% of its 54 rows; False: this message 6.9% of class, 88% of its 59 rows; False: thanks bye 6.8% of class, 91% of its 56 rows; False: a message 6.6% of class, 83% of its 60 rows; False: i wanted to say 6.6% of class, 94% of its 53 rows; False: just um 6.5% of class, 84% of its 58 rows; False: rush at all 6.2% of class, 89% of its 53 rows; False: i m just calling 6.1% of class, 88% of its 52 rows; False: wanted to leave a 6.1% of class, 100% of its 46 rows; False: i am just 6.0% of class, 82% of its 55 rows; False: let you 5.9% of class, 98% of its 45 rows; False: yeah that 5.7% of class, 86% of its 50 rows; False: is all 5.6% of class, 88% of its 48 rows; False: say thank you so 5.6% of class, 93% of its 45 rows; False: just calling to 5.3% of class, 98% of its 41 rows; False: leave this 5.3% of class, 93% of its 43 rows; False: no rush at all 5.3% of class, 89% of its 45 rows; False: message to say 5.2% of class, 100% of its 39 rows; False: just thought 5.1% of class, 86% of its 44 rows; False: just uh 5.1% of class, 83% of its 46 rows; False: thank you very much 5.1% of class, 88% of its 43 rows
  • state.channel = app review (1,430 rows): False: went through 3.4% of class, 96% of its 23 rows; False: everything is 3.3% of class, 95% of its 22 rows; False: perfect 3.1% of class, 95% of its 21 rows; False: whenever 3.1% of class, 100% of its 20 rows
  • state.channel = social media reply (1,208 rows): False: do i 7.8% of class, 85% of its 47 rows; False: happy 5.1% of class, 90% of its 29 rows; False: best 4.1% of class, 91% of its 23 rows; False: perfectly 4.1% of class, 100% of its 21 rows; False: how do i 3.9% of class, 87% of its 23 rows; False: much for 3.9% of class, 95% of its 21 rows

Shortcut models

Predicting the label class on test (1,814 rows). Chance 50.0%, majority class ('True') 57.6%; balanced chance 50.0%.

model (logistic regression, trained on the train sample) test accuracy balanced accuracy (mean recall)
bag of words, whole state (words and word pairs) 91.5% 90.6%
bag of words, main text only (message) 93.6% 93.1%
surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) 64.6% 60.5%
surface features of the main text only 64.1% 60.0%

Strongest single surface features (logistic regression on one feature, balanced accuracy on test):

feature accuracy balanced accuracy
chars(log) 57.6% 50.1%
words(log) 57.6% 50.1%
url 57.6% 50.1%
markdown 57.6% 50.1%
count_* 57.6% 50.1%
nonlatin 57.6% 50.0%
upper_ratio 57.6% 50.0%
digit_ratio 57.6% 50.0%
nonascii_ratio 57.6% 50.0%
emoji 57.6% 50.0%

Other state fields alone (predicting the label class on test from one field, without the main text):

field treated as accuracy balanced accuracy
company text: bag of words / length+empty 57.6% / 57.6% 50.0% / 50.0%
channel categorical, 6 values 57.6% 50.0%

2. Duplicates and split separation

Families shared between splits

splits shared families examples
train ∩ dev 0
train ∩ calibration 0
train ∩ test 0
dev ∩ calibration 0
dev ∩ test 0
calibration ∩ test 0
  • Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
  • Train rows identical in the whole prompt (state, options, instructions): 0.
  • Identical whole prompt, different answer: 0 groups (0 rows).
  • Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)

Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):

split rows examples
dev 0 (0.0%)
calibration 0 (0.0%)
test 0 (0.0%)

Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)

  • Train: 0 near-duplicate pairs; 0 rows (0.0%) sit in 0 clusters; largest cluster 0; excess rows (cluster size − 1) 0 (0.0%).
  • Clusters with more than one label class: 0 (0 rows).
  • Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 0 (0.0%), test 0 (0.0%)

5. Junk

split empty main text main text under 10 characters
train 0 0
dev 0 0
calibration 0 0
test 0 0

Pattern scan of train main texts (count, then the share of each class's rows):

pattern rows by class
placeholder [NAME]-style 0 (0.0%)
lorem ipsum 0 (0.0%)
TODO/TBD/FIXME 0 (0.0%)
'As an AI' / refusal 1 (0.0%) True 0.0%
chat preamble ('Here is/are...', 'Sure!') 0 (0.0%)
meta words (example/variation/message:) 0 (0.0%)
model thinking tags 0 (0.0%)
JSON/code-fence leftovers 0 (0.0%)
encoding garbage (mojibake/replacement char) 0 (0.0%)
HTML tag 511 (3.1%) False 3.4%, True 2.8%
HTML entity 0 (0.0%)
base64-like run (40+ chars) 1 (0.0%) False 0.0%
URL 8 (0.0%) False 0.1%, True 0.0%
  • 'As an AI' / refusal: triage-c290-m140-needs_human (True): Subject: Concern regarding stonework at Hrubý Rohozec ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I hope this message finds you well, though I must confes…

  • HTML tag: triage-c089-m228-needs_human (False): Subject: Fwd: Your consultation is booked ⏎ ⏎ ---------- Forwarded message --------- ⏎ From: SecureHome India sales@securehome.in ⏎ Date… | triage-c221-m174-needs_human (True): Subject: Fwd: Police notification - Accident on A1 ⏎ ⏎ From: PSP Transit Authority transito@psp.pt ⏎ Date: 14 May 2024 08:15 ⏎ To: Maria… | triage-c012-m239-needs_human (True): Subject: Fwd: RE: Booking REF #PK-88291 ⏎ ⏎ ---------- Forwarded message --------- ⏎ From: ParkFr noreply@parkfr.com ⏎ Date: Mon, 14 Oct…

  • base64-like run (40+ chars): triage-c056-m162-needs_human (False): Subject: URGENT: Your store domain will expire!!! ⏎ ⏎ Dear Beloved Manager, ⏎ ⏎ I am writing to you today with a heavy heart because your…

  • URL: triage-c238-m151-needs_human (False): Subject: Meine Bougainvillea blüht wunderschön. ⏎ ⏎ Liebes Team, ⏎ ⏎ ich wollte Ihnen nur mitteilen, dass Ihre Tipps zum Zurückschneiden … | triage-c133-m096-needs_human (False): Subject: Strategic Partnership Proposal - Global Digital Infrastructure ⏎ ⏎ Dear Sir/Madam, ⏎ ⏎ I hope this email finds you well. My name… | triage-c074-m101-needs_human (False): Subject: We offer printing services for your charity ⏎ ⏎ Hello friends, ⏎ ⏎ My name is Laszlo and I am working for GreenPrint Solutions K…

  • Possibly cut off: 5,505 of 11,135 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: False 53.3%, True 47.0%

    • triage-c065-m038-needs_human: …to where everything is right now? I am very worried about this situation. ⏎ ⏎ Thank you, ⏎ Lukas Brunner ⏎ Account: 0881-442-11
    • triage-c032-m188-needs_human: …before Friday, as I need to finalize our winter menu suppliers by then. ⏎ ⏎ Thomas Varga ⏎ Owner & Head Buyer ⏎ +36 30 555 0192
    • triage-c201-m195-needs_human: … discharge tomorrow. ⏎ ⏎ Thank you so much for your wonderful collaboration. ⏎ ⏎ Warm regards, ⏎ Marco Ferretti ⏎ Practice Mana…

Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):

  • ×175: "---------- Forwarded message ---------" (True 108, False 67)
  • ×166: "Any unauthorized review, use, disclosure, or distribution is strictly prohibited." (True 100, False 66)
  • ×159: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and intended solely for the use of the individu…" (True 82, False 77)
  • ×148: "----- Forwarded message -----" (True 106, False 42)
  • ×132: "I hope this message finds you well." (True 84, False 48)
  • ×119: "-----Original Message-----" (True 81, False 38)
  • ×118: "If you have received this email in error, please notify the sender immediately and delete this email from your system." (True 68, False 50)
  • ×116: "If you are not the intended recipient, please delete all copies and notify the sender immediately." (True 87, False 29)
  • ×107: "Topic: Billing and Invoices" (True 62, False 45)
  • ×107: "If you have received this email in error, please notify the sender immediately and delete this message from your system." (True 59, False 48)
  • ×91: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and may also be privileged." (True 62, False 29)
  • ×81: "CONFIDENTIALITY NOTICE: This email and any attachments are strictly confidential and intended solely for the addressee." (True 47, False 34)
  • ×66: "I hope this email finds you well." (False 35, True 31)
  • ×63: "----- End forwarded message -----" (True 46, False 17)
  • ×60: "--- Forwarded message ---" (True 45, False 15)

6. Samples

20 random train rows per kind: triage-needs_human-samples.txt. Reading notes are in the findings above.