QA report: triage
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.
QA: triage
Verdict: PASS WITH NOTES
Triage (READY-triage 2026-09-30 13:33, train sha256 3aba76c5…3b16c0; train 66,532 rows = 16,633 messages × 4 questions). This is the first full check; the earlier triage reports checked withdrawn files. Each question type has its own report: triage-route.md, triage-urgency.md, triage-sentiment.md and triage-needs_human.md. Cue rates come from triage_extra.py.
Earlier cues, now fixed (checked)
Each rate below is the share of that label's rows containing the cue, compared across labels:
| cue | before | now |
|---|---|---|
| "Subject: URGENT" | 22.6% of urgency-4 against 3.0% of the rest | 3.0–4.4% at every urgency level |
| "just wanted to" | 20.4% of urgency-0 | 6.0–6.3% at every level, and 5.9% / 6.2% for needs_human no / yes |
| "by tomorrow" | 14.9% of urgency-2 | 1.8–4.1% |
| "formal complaint" | 7.9% of person-needed against 0.1% | 2.8% / 3.8% |
| "Hi there" | 17% of sentiment-3 | 4.1–5.6% |
| ends with "!" | 39% of very positive | 22–29% at every sentiment level |
"manual", "real person" and "no reply needed" appear in 0 rows.
Surface-only models:
- sentiment: 31.4% balanced accuracy (was 47.8%; chance 20%);
- urgency: 33.4% (chance 20%);
- needs_human: 60.5% (chance 50%);
- route: 9.1% (chance 4.2%).
No single surface feature beats chance by more than 1.3 points.
Route options: the correct team is the longest option 10.6% of the time against 9.8% chance, positions are uniform, and the option picker is at chance.
Duplicates: there are no near duplicates in train or across splits, and no identical prompt carries two different answers.
The edited messages read naturally. In my sample, urgency-4 messages now contain "just wanted to" ("I just wanted to say I need my boxes from unit 12 today!"), low-urgency messages carry "Subject: Urgent" (a spam investment offer; a confirmation email), and very negative messages end with "!".
Notes (meaning; for the model card)
- The remaining surface signal is meaning:
- sentiment: unmatched ")" versus "(" counts, which are emoticons like :) and :( — and message length;
- urgency: fewer "?" in urgent messages (questions are rarely emergencies);
- needs_human: weak length and punctuation mixtures, with no single feature above chance.
- The standard phrase flags are meaning words:
- urgency 4: "right now", "fix this", "deadline", "tonight", "court", "lawyer", "breach";
- urgency 1: "no rush at all", the customer's own statement (kept by agreement);
- urgency 0: "confirm that", "say thank you";
- sentiment: "thrilled", "amazing", "warmest regards", "unacceptable", "nobody";
- needs_human no: "a quick note", "to let you know that", "confirm that the", "perfectly";
- route "other": "my thesis", "journalism student", "sponsoring" (student and sponsorship requests belong to no team).
- The label mix changed with the re-rating: urgency level 4 is now 34.7% of rows, and sentiment level 0 (very negative) is 35.6%. Dev and calibration are small: 1,368 and 1,208 rows, which is 342 and 302 messages (5 organisations each), so calibrate with care.
- About 51% of messages had small cue edits by qwen3.8-flash, recorded in
source.decorrelated_by. Score targets are the mean of three raters (the plan, qwen3.8-max and deepseek-v4-flash).
QA: triage-route
Checked 2026-09-30 13:37 by adapters/qa/qa.py (READY file READY-triage, ).
Verdict: PASS WITH NOTES (full notes in triage-needs_human.md and triage.md)
Automatic flags (for the reviewer to judge; not all are problems)
- 16 strong phrase flags (see list): review whether they are meaning or leakage
- 7 standard phrase flags (≥2% of a class, mostly that class): review
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 16,633 | 260 | train.jsonl |
| dev | 342 | 5 | dev.jsonl |
| calibration | 302 | 5 | calibration.jsonl |
| test | 1,814 | 30 | test.jsonl |
- Train sha256:
3aba76c53cd16a1c2822d1b5e4090d183cc77fd94a233898c53501f25c3b16c0(READY file gives no checksum) - Main text field (the text the phrase and length checks use):
state.message. - Label classes: <listed option>, other (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): route.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| <listed option> | 93.2% | 94.4% | 95.4% | 94.0% | 15,509 |
| other | 6.8% | 5.6% | 4.6% | 6.0% | 1,124 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| route | 100.0% | 100.0% | 100.0% | 100.0% | 16,633 |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Options per choice row
| split | min | median | p99 | max |
|---|---|---|---|---|
| train | 5 | 13 | 38 | 39 |
| dev | 12 | 16 | 27 | 27 |
| calibration | 5 | 11 | 36 | 36 |
| test | 6 | 15 | 36 | 36 |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | estimate: characters / 3 (upper bound for English) | 841 | 2039 | 2524 | 0 |
| dev | estimate: characters / 3 (upper bound for English) | 893 | 1614 | 1967 | 0 |
| calibration | estimate: characters / 3 (upper bound for English) | 729 | 2327 | 2433 | 0 |
| test | estimate: characters / 3 (upper bound for English) | 820 | 1958 | 2394 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| channel, company, message | 16,633 (100.0%) |
Instructions (train)
- Canonical (the most common text) 70.1%, reworded 26.7% (70 distinct rewordings), none 3.2%. Target about 70 / 27 / 3.
- Canonical text: "Which team should handle this message? If it covers several issues, choose the team for the most important one."
sourceinstruction tag: canonical 70.1%, none 3.2%, variant-42 0.5%, variant-21 0.5%, variant-68 0.5%, variant-30 0.5%
| class | canonical | none |
|---|---|---|
| <listed option> | 70.0% | 3.2% |
| other | 70.6% | 2.9% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 16,633 train rows; models are scored on the full test file (1,814 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | <listed option> | 15509 | 136 | 434 | 1388 | 627 |
| train | other | 1124 | 127 | 425 | 1431 | 637 |
| test | <listed option> | 1705 | 124 | 423 | 1415 | 626 |
| test | other | 109 | 110 | 410 | 1297 | 592 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| route | 16633 | 433 | 628 | 145 | 13 |
Correct option: longest / shortest / position / key
For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):
| split | rows | correct is longest | correct is shortest | chance (1/listed) | mean relative position (0 first, 1 last; 0.5 expected) | position fifths |
|---|---|---|---|---|---|---|
| train | 15509 | 10.6% | 10.9% | 9.8% | 0.497 | 22% / 18% / 18% / 18% / 24% |
| dev | 323 | 9.0% | 10.5% | 6.9% | 0.484 | 18% / 22% / 21% / 17% / 22% |
| calibration | 288 | 13.9% | 9.0% | 12.0% | 0.520 | 19% / 19% / 18% / 22% / 22% |
| test | 1705 | 11.1% | 11.3% | 10.1% | 0.500 | 21% / 18% / 20% / 18% / 24% |
Correct key and position by option count
| split | options | rows | mean options | top correct keys | most common position (0-based) |
|---|---|---|---|---|---|
| train | 2-5 | 775 | 5.0 | k4 23.9%, k3 23.1%, k2 22.8%, k1 21.0%, other 9.2% | 4 (23.9%) |
| train | 6-10 | 5116 | 7.8 | k1 14.7%, k2 14.3%, k5 13.9%, k3 13.8%, k4 13.7% | 1 (14.7%) |
| train | 11-30 | 9463 | 17.7 | k1 6.6%, k8 6.5%, k3 6.4%, k7 6.4%, k5 6.4% | 1 (6.6%) |
| train | 31-80 | 1279 | 35.4 | other 4.7%, k27 3.7%, k29 3.6%, k13 3.6%, k1 3.4% | 0 (4.7%) |
| test | 6-10 | 709 | 7.3 | k4 18.1%, k2 16.9%, k3 15.4%, k5 14.5%, k1 13.7% | 4 (18.1%) |
| test | 11-30 | 982 | 18.4 | k3 7.3%, k8 6.7%, k2 6.6%, k1 6.3%, k9 6.2% | 3 (7.3%) |
| test | 31-80 | 123 | 33.6 | k22 6.5%, k26 4.9%, k9 4.9%, k15 4.9%, k12 4.1% | 22 (6.5%) |
- Train rows whose correct listed key is the first listed key (t1/o1): 10.2%, chance 9.8%.
Option count by label class (train)
| class | rows | min | median | mean | max |
|---|---|---|---|---|---|
| <listed option> | 15509 | 5 | 13 | 15.6 | 39 |
| other | 1124 | 5 | 12 | 13.5 | 39 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 93.2%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| other_kind | 9 | 100.0% | null: <listed option> 100.0%; a job application or question about working there: other 100.0%; a message clearly meant for a different organisation: other 100.0%; a sales pitch from another business offering its services: other 100.0%; a request to sponsor or donate to a local event: other 100.0%; a journalist or student asking for an interview or information for a project: other 100.0% |
| channel | 6 | 93.2% | email: <listed option> 93.3%, other 6.7%; web form: <listed option> 93.2%, other 6.8%; chat: <listed option> 93.0%, other 7.0%; phone transcript: <listed option> 92.7%, other 7.3%; app review: <listed option> 94.3%, other 5.7%; social media reply: <listed option> 92.9%, other 7.1% |
| language | 11 | 93.2% | English: <listed option> 93.2%, other 6.8%; Danish: <listed option> 92.9%, other 7.1%; Czech: <listed option> 93.8%, other 6.2%; French: <listed option> 95.3%, other 4.7%; Swedish: <listed option> 92.0%, other 8.0%; Dutch: <listed option> 91.1%, other 8.9% |
| length_target | 17 | 93.2% | 60: <listed option> 93.6%, other 6.4%; 120: <listed option> 93.2%, other 6.8%; 25: <listed option> 93.5%, other 6.5%; 220: <listed option> 93.0%, other 7.0%; 50: <listed option> 92.2%, other 7.8%; 10: <listed option> 92.5%, other 7.5% |
| tone | 25 | 93.2% | politely formal: <listed option> 90.2%, other 9.8%; cheerful: <listed option> 92.5%, other 7.5%; polite and friendly: <listed option> 89.0%, other 11.0%; warm: <listed option> 90.5%, other 9.5%; disappointed: <listed option> 95.8%, other 4.2%; appreciative: <listed option> 90.6%, other 9.4% |
| prompt_version | 3 | 93.2% | 3: <listed option> 93.4%, other 6.6%; 2: <listed option> 92.9%, other 7.1%; 1: <listed option> 93.1%, other 6.9% |
| human_reason | 7 | 93.2% | manual: <listed option> 98.9%, other 1.1%; legal_safety: <listed option> 100.0%; routine: <listed option> 88.3%, other 11.7%; acknowledge: <listed option> 81.5%, other 18.5%; escalation: <listed option> 100.0%; no_reply: <listed option> 80.6%, other 19.4% |
| decorrelated_by | 2 | 93.2% | qwen3.8-flash: <listed option> 92.3%, other 7.7%; null: <listed option> 94.3%, other 5.7% |
| situation | 17 | 93.2% | null: <listed option> 92.9%, other 7.1%; the writer reports or confirms something that needs no decision, such as that a delivery arrived, a payment was made, a problem sorted itself out, or a form was sent: <listed option> 80.9%, other 19.1%; the writer is about to take it further and says what they will do next, for example cancel, go to a regulator or ombudsman, post a public review, contact t… |
| regenerated | 2 | 93.2% | false: <listed option> 93.0%, other 7.0%; true: <listed option> 94.0%, other 6.0% |
Formatting by label class (main text, share of rows)
| feature | <listed option> | other | |
|---|---|---|---|
| ends with ? | 4.6% | 2.0% | |
| ends with . | 38.0% | 35.2% | |
| ends with ! | 13.3% | 14.5% | |
| no end punctuation | 43.9% | 48.0% | |
| starts lowercase | 6.4% | 7.0% | |
| all lowercase | 4.2% | 6.4% | |
| has a digit | 88.8% | 76.7% | |
| has newline | 77.3% | 73.4% | |
| has quotes | 6.5% | 1.5% | |
| has markup (HTML/markdown) | 3.1% | 3.3% | |
| has URL | 0.0% | 0.4% | |
| non-ASCII | 51.7% | 44.0% | |
| non-Latin script | 0.0% | 0.1% | |
| emoji | 5.4% | 5.0% | |
| ALL-CAPS word (4+) | 16.1% | 17.2% | |
| contains ' - ' or — | 17.8% | 14.0% |
Same, by row kind
| feature | route |
|---|---|
| ends with ? | 4% |
| ends with . | 38% |
| ends with ! | 13% |
| no end punctuation | 44% |
| starts lowercase | 6% |
| all lowercase | 4% |
| has a digit | 88% |
| has newline | 77% |
| has quotes | 6% |
| has markup (HTML/markdown) | 3% |
| has URL | 0% |
| non-ASCII | 51% |
| non-Latin script | 0% |
| emoji | 5% |
| ALL-CAPS word (4+) | 16% |
| contains ' - ' or — | 18% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, <listed option>: today z=11 28.9% vs 9.4%; before z=11 21.5% vs 5.3%; now z=10 32.3% vs 14.1%; account z=8 25.7% vs 12.3%; also z=8 20.5% vs 9.2%; cannot z=7 10.0% vs 2.3%; fix z=7 10.5% vs 0.4%; my z=7 63.3% vs 44.8%; tomorrow z=7 13.3% vs 5.1%; deadline z=7 9.2% vs 2.6%; right z=6 14.8% vs 6.9%; legal z=6 6.7% vs 1.1%; but z=6 29.1% vs 18.3%; immediately z=6 12.4% vs 5.4%; need z=6 17.6% vs 9.3%; card z=6 6.5% vs 1.7%; into z=6 9.8% vs 4.1%; t z=5 10.1% vs 4.4%; charge z=5 5.6% vs 1.3%; says z=5 5.1% vs 1.0%
Words, other: interview z=16 7.5% vs 0.8%; role z=16 6.3% vs 0.5%; university z=15 6.4% vs 0.7%; opportunity z=14 4.5% vs 0.2%; thesis z=14 4.7% vs 0.1%; delivery z=13 10.7% vs 2.7%; cv z=13 4.3% vs 0.2%; research z=13 4.8% vs 0.4%; hiring z=13 4.3% vs 0.3%; solutions z=13 6.1% vs 1.0%; proposal z=13 4.8% vs 0.5%; sponsorship z=12 4.4% vs 0.5%; sponsor z=12 3.6% vs 0.1%; company z=12 14.7% vs 5.4%; offer z=12 5.2% vs 0.8%; student z=12 6.2% vs 1.2%; position z=11 5.0% vs 0.8%; testing z=11 3.2% vs 0.2%; journalism z=11 3.0% vs 0.1%; best z=11 14.4% vs 5.6%
2–4-word phrases, <listed option>: right now z=8 11.0% vs 2.2%; i need z=6 8.5% vs 2.3%; my account z=6 7.7% vs 1.9%; i cannot z=5 5.3% vs 1.2%; before the z=5 5.3% vs 1.5%; today i z=5 4.4% vs 0.8%; need to z=5 7.5% vs 3.1%; is not z=5 5.4% vs 1.8%; if i z=5 4.6% vs 1.2%; to the z=5 12.4% vs 7.0%; look into z=4 3.4% vs 0.4%; instead of z=4 3.2% vs 0.5%; to my z=4 6.2% vs 2.8%; but the z=4 3.9% vs 1.2%; when i z=4 6.1% vs 2.8%; tomorrow morning z=4 3.3% vs 0.2%; my bank z=4 3.5% vs 1.0%; i can z=4 5.3% vs 2.3%; could you z=4 9.1% vs 5.2%; you please z=4 4.1% vs 1.6%
2–4-word phrases, other: your company z=12 6.9% vs 1.4%; to your z=11 9.5% vs 3.1%; to say z=11 13.3% vs 5.3%; my cv z=11 2.7% vs 0.1%; my thesis z=10 2.8% vs 0.1%; my application z=10 4.5% vs 0.9%; working with z=10 3.6% vs 0.5%; hr 2024 z=10 2.4% vs 0.1%; dear hiring z=10 2.2% vs 0.1%; reaching out z=10 4.8% vs 1.1%; we are z=10 17.4% vs 8.6%; the delivery z=9 3.3% vs 0.5%; interview request z=9 2.0% vs 0.1%; application for z=9 3.6% vs 0.7%; i wanted to z=9 8.8% vs 3.3%; i wanted z=9 8.8% vs 3.4%; hiring team z=9 2.0% vs 0.1%; my application for the z=9 2.1% vs 0.2%; my application for z=9 2.3% vs 0.3%; dear hiring team z=9 1.8% vs 0.1%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- <listed option>:
before21.5% vs 5.3% - <listed option>:
right now11.0% vs 2.2% - <listed option>:
fix10.5% vs 0.4% - <listed option>:
cannot10.0% vs 2.3% - <listed option>:
my account7.7% vs 1.9% - other:
interview7.5% vs 0.8% - other:
your company6.9% vs 1.4% - <listed option>:
legal6.7% vs 1.1% - other:
university6.4% vs 0.7% - other:
role6.3% vs 0.5% - other:
student6.2% vs 1.2% - other:
solutions6.1% vs 1.0% - <listed option>:
charge5.6% vs 1.3% - <listed option>:
i cannot5.3% vs 1.2% - other:
offer5.2% vs 0.8% - <listed option>:
says5.1% vs 1.0%
Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):
- all rows (16,633 rows): other:
thesis4.7% of class, 75% of its 71 rows; other:sponsor3.6% of class, 71% of its 56 rows; other:student at3.4% of class, 95% of its 40 rows; other:my thesis2.8% of class, 80% of its 40 rows; other:journalism student2.7% of class, 100% of its 30 rows; other:careers2.5% of class, 80% of its 35 rows; other:sponsoring2.4% of class, 77% of its 35 rows - state.channel = email (6,611 rows): other:
thesis6.4% of class, 72% of its 39 rows; other:student at4.8% of class, 100% of its 21 rows - state.channel = web form (3,009 rows): none
- state.channel = chat (2,529 rows): none
- state.channel = phone transcript (1,846 rows): none
- state.channel = app review (1,430 rows): none
- state.channel = social media reply (1,208 rows): none
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
- Train rows identical in the whole prompt (state, options, instructions): 0.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 0 (0.0%) | |
| calibration | 0 (0.0%) | |
| test | 0 (0.0%) |
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 0 |
| dev | 0 | 0 |
| calibration | 0 | 0 |
| test | 0 | 0 |
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 0 (0.0%) | |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 1 (0.0%) | <listed option> 0.0% |
| chat preamble ('Here is/are...', 'Sure!') | 0 (0.0%) | |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 0 (0.0%) | |
| encoding garbage (mojibake/replacement char) | 0 (0.0%) | |
| HTML tag | 511 (3.1%) | <listed option> 3.1%, other 3.2% |
| HTML entity | 0 (0.0%) | |
| base64-like run (40+ chars) | 1 (0.0%) | other 0.1% |
| URL | 8 (0.0%) | <listed option> 0.0%, other 0.4% |
'As an AI' / refusal:
triage-c290-m140-route(<listed option>): Subject: Concern regarding stonework at Hrubý Rohozec ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I hope this message finds you well, though I must confes…HTML tag:
triage-c227-m067-route(<listed option>): ---------- Forwarded message --------- ⏎ From: Berlin Dental Group noreply@berlindental.eu ⏎ Date: 24 May 2024 at 08:15 ⏎ Subject: RE: Ne… |triage-c196-m133-route(<listed option>): Subject: Fwd: RE: Redemption statement completion - 12 Maple Road ⏎ ⏎ From: Helen Croft helen.croft@smithsolicitors.co.uk ⏎ Sent: 21 Nov… |triage-c151-m155-route(<listed option>): ---------- Forwarded message --------- ⏎ From: Billing Dept billing@secureguard.cz ⏎ Date: Mon, Oct 14, 2024 at 8:00 AM ⏎ Subject: Your i…base64-like run (40+ chars):
triage-c056-m162-route(other): Subject: URGENT: Your store domain will expire!!! ⏎ ⏎ Dear Beloved Manager, ⏎ ⏎ I am writing to you today with a heavy heart because your…URL:
triage-c133-m096-route(other): Subject: Strategic Partnership Proposal - Global Digital Infrastructure ⏎ ⏎ Dear Sir/Madam, ⏎ ⏎ I hope this email finds you well. My name… |triage-c073-m068-route(other): Name: Rajesh Kumer ⏎ Email: rajesh.kumer82@yahho.co.in ⏎ Topic: Claims Processing ⏎ Message: ⏎ Dear Sir or Madam, ⏎ ⏎ I am writting to you… |triage-c074-m101-route(other): Subject: We offer printing services for your charity ⏎ ⏎ Hello friends, ⏎ ⏎ My name is Laszlo and I am working for GreenPrint Solutions K…Possibly cut off: 5,505 of 11,135 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: <listed option> 49.1%, other 54.6%
triage-c195-m064-route: …t pop by next month for a routine checkup. ⏎ ⏎ Thank you again for everyting. ⏎ ⏎ With much appriciation, ⏎ Katarzyna Wisniewskatriage-c142-m013-route: …ay. It truly makes a difference to families like ours. ⏎ ⏎ Warmest regards, ⏎ ⏎ Kristin Hagen ⏎ Parent / Guardian ⏎ +47 912 34 …triage-c206-m086-route: …l double charge on my last invoice, probably just a glitch. No rush at all. ⏎ Warmly, ⏎ Sarah Jenkins ⏎ Homeowner ⏎ 082 555 1234
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- ×175: "---------- Forwarded message ---------" (<listed option> 163, other 12)
- ×166: "Any unauthorized review, use, disclosure, or distribution is strictly prohibited." (<listed option> 155, other 11)
- ×159: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and intended solely for the use of the individu…" (<listed option> 146, other 13)
- ×148: "----- Forwarded message -----" (<listed option> 142, other 6)
- ×132: "I hope this message finds you well." (<listed option> 121, other 11)
- ×119: "-----Original Message-----" (<listed option> 111, other 8)
- ×118: "If you have received this email in error, please notify the sender immediately and delete this email from your system." (<listed option> 112, other 6)
- ×116: "If you are not the intended recipient, please delete all copies and notify the sender immediately." (<listed option> 112, other 4)
- ×107: "If you have received this email in error, please notify the sender immediately and delete this message from your system." (<listed option> 97, other 10)
- ×107: "Topic: Billing and Invoices" (<listed option> 99, other 8)
- ×91: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and may also be privileged." (<listed option> 87, other 4)
- ×81: "CONFIDENTIALITY NOTICE: This email and any attachments are strictly confidential and intended solely for the addressee." (<listed option> 71, other 10)
- ×66: "I hope this email finds you well." (<listed option> 59, other 7)
- ×63: "----- End forwarded message -----" (<listed option> 60, other 3)
- ×60: "--- Forwarded message ---" (<listed option> 56, other 4)
6. Samples
20 random train rows per kind: triage-route-samples.txt. Reading notes are in the findings above.
QA: triage-urgency
Checked 2026-09-30 13:36 by adapters/qa/qa.py (READY file READY-triage, ).
Verdict: PASS WITH NOTES (full notes in triage-needs_human.md and triage.md)
Automatic flags (for the reviewer to judge; not all are problems)
- 56 strong phrase flags (see list): review whether they are meaning or leakage
- 60 standard phrase flags (≥2% of a class, mostly that class): review
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 16,633 | 260 | train.jsonl |
| dev | 342 | 5 | dev.jsonl |
| calibration | 302 | 5 | calibration.jsonl |
| test | 1,814 | 30 | test.jsonl |
- Train sha256:
3aba76c53cd16a1c2822d1b5e4090d183cc77fd94a233898c53501f25c3b16c0(READY file gives no checksum) - Main text field (the text the phrase and length checks use):
state.message. - Label classes: 0, 1, 2, 3, 4 (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): urgency.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| 0 | 17.9% | 18.4% | 18.2% | 18.2% | 2,976 |
| 1 | 11.5% | 10.2% | 10.6% | 12.1% | 1,914 |
| 2 | 22.4% | 25.4% | 24.2% | 21.3% | 3,724 |
| 3 | 13.5% | 13.7% | 13.2% | 13.2% | 2,253 |
| 4 | 34.7% | 32.2% | 33.8% | 35.1% | 5,766 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| urgency | 100.0% | 100.0% | 100.0% | 100.0% | 16,633 |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | estimate: characters / 3 (upper bound for English) | 205 | 881 | 1334 | 0 |
| dev | estimate: characters / 3 (upper bound for English) | 221 | 880 | 1088 | 0 |
| calibration | estimate: characters / 3 (upper bound for English) | 215 | 975 | 1152 | 0 |
| test | estimate: characters / 3 (upper bound for English) | 203 | 842 | 1165 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| channel, company, message | 16,633 (100.0%) |
Instructions (train)
- Canonical (the most common text) 69.7%, reworded 27.1% (70 distinct rewordings), none 3.2%. Target about 70 / 27 / 3.
- Canonical text: "How urgently does this need a response?"
sourceinstruction tag: canonical 69.7%, none 3.2%, variant-62 0.5%, variant-25 0.5%, variant-16 0.5%, variant-6 0.5%
| class | canonical | none |
|---|---|---|
| 0 | 69.4% | 2.9% |
| 1 | 69.7% | 3.2% |
| 2 | 70.5% | 3.2% |
| 3 | 70.6% | 2.8% |
| 4 | 69.0% | 3.6% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 16,633 train rows; models are scored on the full test file (1,814 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | 0 | 2976 | 121 | 439 | 1395 | 628 |
| train | 1 | 1914 | 170 | 622 | 1532 | 743 |
| train | 2 | 3724 | 126 | 384 | 1306 | 566 |
| train | 3 | 2253 | 120 | 405 | 1338 | 582 |
| train | 4 | 5766 | 143 | 450 | 1413 | 647 |
| test | 0 | 331 | 117 | 403 | 1576 | 629 |
| test | 1 | 220 | 150 | 667 | 1542 | 745 |
| test | 2 | 386 | 114 | 396 | 1277 | 546 |
| test | 3 | 240 | 100 | 423 | 1401 | 608 |
| test | 4 | 637 | 143 | 435 | 1370 | 633 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| urgency | 16633 | 433 | 628 | 145 | 0 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 34.7%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| human_reason | 7 | 60.5% | manual: 4 46.0%, 2 29.0%; legal_safety: 4 66.0%, 2 17.2%; routine: 2 53.0%, 1 35.5%; acknowledge: 0 63.3%, 1 23.0%; escalation: 4 71.2%, 3 21.8%; no_reply: 0 86.4%, 1 13.6% |
| situation | 17 | 53.7% | null: 4 33.0%, 2 26.6%; the writer reports or confirms something that needs no decision, such as that a delivery arrived, a payment was made, a problem sorted itself out, or a form was sent: 0 73.9%, 1 19.1%; the writer is about to take it further and says what they will do next, for example cancel, go to a regulator or ombudsman, post a public review, contact the press, or make a formal complain… |
| tone | 25 | 49.5% | politely formal: 2 35.2%, 0 28.8%; cheerful: 2 31.9%, 0 21.3%; polite and friendly: 2 28.3%, 0 23.8%; warm: 2 30.5%, 0 23.5%; disappointed: 4 53.8%, 2 18.1%; appreciative: 2 29.4%, 0 26.8% |
| other_kind | 9 | 38.8% | null: 4 37.2%, 2 23.3%; a job application or question about working there: 1 45.0%, 0 41.1%; a message clearly meant for a different organisation: 0 77.9%, 1 13.8%; a sales pitch from another business offering its services: 0 67.8%, 1 28.3%; a request to sponsor or donate to a local event: 1 45.9%, 0 38.5%; a journalist or student asking for an interview or information for a project: 0 42.4%, 1 4… |
| channel | 6 | 34.7% | email: 4 33.6%, 2 21.4%; web form: 4 34.6%, 2 22.4%; chat: 4 38.6%, 2 26.3%; phone transcript: 4 36.5%, 0 20.9%; app review: 4 32.9%, 0 22.2%; social media reply: 4 32.0%, 2 27.5% |
| language | 11 | 34.7% | English: 4 34.9%, 2 22.5%; Danish: 4 31.7%, 0 25.1%; Czech: 4 28.8%, 2 23.2%; French: 4 34.1%, 2 24.1%; Swedish: 4 27.6%, 2 24.5%; Dutch: 4 34.2%, 0 21.5% |
| length_target | 17 | 34.7% | 60: 4 34.2%, 2 22.8%; 120: 4 34.6%, 2 20.5%; 25: 4 33.3%, 2 28.1%; 220: 4 33.6%, 2 19.4%; 50: 4 36.6%, 2 26.5%; 10: 4 32.8%, 2 23.6% |
| prompt_version | 3 | 34.7% | 3: 4 35.4%, 2 20.6%; 2: 4 32.7%, 2 26.7%; 1: 4 45.7%, 2 21.6% |
| decorrelated_by | 2 | 34.7% | qwen3.8-flash: 4 32.1%, 2 23.1%; null: 4 37.4%, 2 21.6% |
| regenerated | 2 | 34.7% | false: 4 33.9%, 2 23.5%; true: 4 36.7%, 0 20.6% |
| target_kind | 2 | 34.7% | soft: 4 31.1%, 2 26.2%; hard: 4 41.5%, 0 28.5% |
Formatting by label class (main text, share of rows)
| feature | 0 | 1 | 2 | 3 | 4 | |
|---|---|---|---|---|---|---|
| ends with ? | 0.1% | 2.9% | 8.3% | 6.5% | 3.8% | |
| ends with . | 38.9% | 33.9% | 33.4% | 36.5% | 41.9% | |
| ends with ! | 14.1% | 12.9% | 13.3% | 12.9% | 13.3% | |
| no end punctuation | 46.4% | 49.9% | 44.6% | 44.0% | 40.9% | |
| starts lowercase | 5.4% | 4.0% | 7.1% | 6.5% | 7.4% | |
| all lowercase | 5.2% | 3.3% | 4.7% | 4.9% | 3.7% | |
| has a digit | 84.0% | 84.7% | 87.1% | 89.7% | 91.0% | |
| has newline | 76.9% | 82.1% | 76.5% | 76.7% | 75.8% | |
| has quotes | 2.1% | 5.1% | 5.2% | 8.4% | 8.4% | |
| has markup (HTML/markdown) | 4.8% | 4.0% | 2.0% | 2.4% | 3.1% | |
| has URL | 0.2% | 0.1% | 0.0% | 0.0% | 0.0% | |
| non-ASCII | 46.1% | 51.2% | 50.6% | 53.8% | 53.3% | |
| non-Latin script | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | |
| emoji | 4.9% | 4.1% | 6.6% | 5.9% | 5.2% | |
| ALL-CAPS word (4+) | 16.4% | 16.1% | 15.8% | 15.8% | 16.5% | |
| contains ' - ' or — | 14.9% | 18.0% | 15.0% | 18.2% | 20.1% |
Same, by row kind
| feature | urgency |
|---|---|
| ends with ? | 4% |
| ends with . | 38% |
| ends with ! | 13% |
| no end punctuation | 44% |
| starts lowercase | 6% |
| all lowercase | 4% |
| has a digit | 88% |
| has newline | 77% |
| has quotes | 6% |
| has markup (HTML/markdown) | 3% |
| has URL | 0% |
| non-ASCII | 51% |
| non-Latin script | 0% |
| emoji | 5% |
| ALL-CAPS word (4+) | 16% |
| contains ' - ' or — | 18% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, 0: thank z=25 43.8% vs 19.5%; perfectly z=25 13.9% vs 2.2%; everything z=22 26.1% vs 10.2%; arrived z=21 12.1% vs 2.7%; say z=20 19.8% vs 7.1%; went z=20 11.5% vs 2.7%; bye z=17 6.7% vs 1.0%; perfect z=17 6.0% vs 0.8%; wanted z=17 20.3% vs 9.0%; successfully z=17 5.6% vs 0.6%; completed z=16 5.7% vs 1.0%; smoothly z=15 6.1% vs 1.2%; was z=15 43.0% vs 27.1%; note z=15 11.0% vs 3.9%; all z=15 29.9% vs 17.0%; made z=15 9.4% vs 3.0%; finally z=15 6.4% vs 1.5%; sorted z=14 8.9% vs 3.1%; such z=13 11.0% vs 4.6%; settled z=13 4.3% vs 0.8%
Words, 1: rush z=27 19.5% vs 2.0%; whenever z=24 15.4% vs 1.6%; question z=17 15.4% vs 3.8%; wondering z=16 8.5% vs 1.2%; moment z=14 6.9% vs 1.2%; small z=13 13.9% vs 4.6%; thought z=12 7.8% vs 1.9%; would z=12 20.3% vs 8.5%; such z=11 13.2% vs 4.8%; curious z=11 3.0% vs 0.2%; maybe z=11 10.5% vs 3.5%; much z=11 33.2% vs 17.2%; fine z=11 9.7% vs 3.2%; all z=10 32.9% vs 17.6%; warmest z=10 10.7% vs 3.9%; anyway z=10 14.1% vs 5.8%; quick z=10 11.2% vs 4.2%; recently z=10 5.1% vs 1.2%; how z=10 27.5% vs 14.3%; thank z=10 38.9% vs 21.9%
Words, 2: could z=27 31.4% vs 12.6%; question z=16 10.1% vs 3.8%; kindly z=14 6.8% vs 2.2%; would z=13 14.7% vs 8.4%; can z=13 23.2% vs 15.4%; thanks z=12 18.1% vs 11.4%; love z=12 10.0% vs 5.3%; hope z=11 9.6% vs 5.1%; standard z=11 7.1% vs 3.3%; hello z=11 22.6% vs 16.0%; possible z=11 4.3% vs 1.5%; ask z=11 6.9% vs 3.3%; plan z=11 7.2% vs 3.6%; much z=10 24.0% vs 17.6%; advise z=10 3.7% vs 1.2%; next z=10 15.8% vs 10.6%; hi z=10 23.8% vs 17.7%; want z=10 8.6% vs 4.9%; arrange z=10 3.1% vs 1.0%; please z=10 31.9% vs 25.5%
Words, 3: before z=17 34.2% vs 18.3%; today z=16 42.1% vs 25.3%; need z=12 26.1% vs 15.6%; tomorrow z=11 19.7% vs 11.6%; closes z=8 3.4% vs 1.2%; afternoon z=8 9.7% vs 5.7%; friday z=8 7.0% vs 3.7%; thursday z=8 5.5% vs 2.7%; could z=7 21.8% vs 16.0%; tomorow z=7 2.6% vs 0.8%; can z=7 22.0% vs 16.4%; 17 z=7 5.5% vs 2.8%; still z=7 12.5% vs 8.6%; evening z=6 4.8% vs 2.6%; trustpilot z=6 1.4% vs 0.4%; day z=6 10.4% vs 7.3%; be z=6 26.5% vs 21.8%; also z=6 23.5% vs 19.1%; approved z=6 3.9% vs 2.1%; please z=6 31.0% vs 26.3%
Words, 4: today z=36 49.9% vs 15.7%; right z=30 27.7% vs 7.1%; now z=29 50.1% vs 20.9%; fix z=29 20.8% vs 3.9%; immediately z=26 22.7% vs 6.1%; deadline z=26 18.0% vs 3.8%; tomorrow z=23 22.2% vs 7.7%; cannot z=22 17.4% vs 5.3%; tonight z=21 10.1% vs 1.3%; legal z=20 12.4% vs 3.1%; nobody z=18 8.2% vs 1.4%; filing z=18 9.6% vs 2.3%; court z=18 6.9% vs 0.8%; before z=17 29.6% vs 15.6%; lose z=17 6.4% vs 0.7%; if z=17 43.8% vs 26.1%; hour z=17 6.1% vs 0.6%; completely z=17 12.3% vs 4.4%; breach z=16 5.8% vs 0.5%; will z=16 25.1% vs 13.0%
2–4-word phrases, 0: to say z=26 17.6% vs 3.3%; thank you z=23 42.2% vs 19.1%; confirm that z=21 8.9% vs 0.7%; to confirm z=21 10.2% vs 1.5%; that the z=21 14.4% vs 3.6%; to confirm that z=21 8.2% vs 0.6%; i wanted z=20 10.8% vs 2.2%; i wanted to z=20 10.7% vs 2.2%; everything is z=19 8.7% vs 1.4%; let you z=18 6.7% vs 0.7%; you know z=18 8.0% vs 1.4%; to let z=18 6.6% vs 0.9%; let you know z=17 6.1% vs 0.6%; you for z=17 17.0% vs 6.5%; thank you for z=17 16.4% vs 6.2%; to let you z=17 6.0% vs 0.6%; to let you know z=17 5.7% vs 0.6%; this morning z=17 16.1% vs 6.2%; wanted to z=16 19.9% vs 8.9%; confirm that the z=16 5.1% vs 0.2%
2–4-word phrases, 1: no rush z=24 15.4% vs 1.3%; at all z=21 16.2% vs 2.7%; rush at z=21 10.3% vs 0.5%; rush at all z=21 10.3% vs 0.4%; no rush at z=19 9.1% vs 0.4%; no rush at all z=19 9.1% vs 0.4%; whenever you z=18 9.0% vs 0.8%; you get z=13 4.8% vs 0.5%; whenever you get z=13 4.2% vs 0.3%; question about z=13 7.9% vs 1.7%; you get a z=12 4.1% vs 0.4%; whenever you have z=12 4.0% vs 0.4%; you have a z=12 5.0% vs 0.7%; whenever you get a z=12 3.6% vs 0.2%; one small z=12 4.4% vs 0.6%; quick question z=12 6.1% vs 1.2%; perfectly fine z=11 3.3% vs 0.2%; just wondering z=11 3.7% vs 0.4%; wondering if z=11 3.9% vs 0.6%; whenever you have a z=11 3.1% vs 0.3%
2–4-word phrases, 2: could you z=23 18.6% vs 6.0%; could you please z=14 7.0% vs 2.2%; you please z=14 7.7% vs 2.8%; let me know z=12 5.4% vs 1.8%; me know z=12 5.4% vs 1.8%; do i z=12 4.9% vs 1.5%; question about z=12 5.0% vs 1.7%; i would z=11 6.4% vs 2.9%; you kindly z=11 3.7% vs 1.2%; look into z=11 5.7% vs 2.4%; let me know what z=11 2.4% vs 0.4%; me know what z=11 2.4% vs 0.4%; within the next day z=10 2.4% vs 0.5%; so much z=10 19.5% vs 13.4%; like to z=10 2.8% vs 0.7%; the next day z=10 2.5% vs 0.5%; next day z=10 2.5% vs 0.5%; how do i z=10 2.1% vs 0.4%; how do z=10 2.4% vs 0.5%; quick question z=10 3.7% vs 1.3%
2–4-word phrases, 3: end of z=14 5.6% vs 1.2%; before the end of z=14 4.0% vs 0.5%; before the end z=14 4.0% vs 0.5%; the end of z=13 4.8% vs 1.0%; the end z=12 4.8% vs 1.1%; of thursday z=12 2.9% vs 0.3%; end of thursday z=12 2.9% vs 0.3%; the end of thursday z=12 2.9% vs 0.3%; before the z=11 9.5% vs 4.4%; need to z=10 12.2% vs 6.4%; i need z=10 12.9% vs 7.3%; today before z=10 2.8% vs 0.6%; next day z=9 2.8% vs 0.7%; the next day z=9 2.8% vs 0.7%; within the next day z=9 2.7% vs 0.6%; before then z=8 2.3% vs 0.5%; we need z=8 4.3% vs 1.8%; i need to z=8 6.2% vs 3.1%; need to know z=8 3.4% vs 1.3%; could someone z=7 3.5% vs 1.4%
2–4-word phrases, 4: right now z=34 24.3% vs 3.1%; fix this z=25 13.4% vs 1.3%; this is z=22 19.3% vs 6.9%; tomorrow morning z=18 7.2% vs 0.9%; today i z=18 8.5% vs 1.9%; if this z=17 9.7% vs 2.9%; i cannot z=17 9.3% vs 2.8%; or i z=16 5.4% vs 0.8%; 1 5 z=15 4.9% vs 0.5%; is not z=15 8.9% vs 3.1%; 1 5 stars z=15 4.6% vs 0.5%; 17 00 z=14 5.1% vs 1.0%; my lawyer z=14 4.3% vs 0.5%; 00 today z=14 4.6% vs 0.8%; or we z=14 4.4% vs 0.2%; right now i z=14 4.1% vs 0.5%; filing a z=14 5.1% vs 1.2%; complaint with z=14 4.1% vs 0.6%; this to z=14 4.9% vs 1.1%; is completely z=14 4.0% vs 0.6%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- 4:
right now24.3% vs 3.1% - 4:
fix20.8% vs 3.9% - 1:
rush19.5% vs 2.0% - 4:
deadline18.0% vs 3.8% - 0:
to say17.6% vs 3.3% - 1:
at all16.2% vs 2.7% - 1:
whenever15.4% vs 1.6% - 1:
question15.4% vs 3.8% - 1:
no rush15.4% vs 1.3% - 0:
perfectly13.9% vs 2.2% - 4:
fix this13.4% vs 1.3% - 0:
arrived12.1% vs 2.7% - 0:
went11.5% vs 2.7% - 0:
i wanted10.8% vs 2.2% - 0:
i wanted to10.7% vs 2.2% - 1:
rush at10.3% vs 0.5% - 1:
rush at all10.3% vs 0.4% - 0:
to confirm10.2% vs 1.5% - 4:
tonight10.1% vs 1.3% - 4:
filing9.6% vs 2.3% - 1:
no rush at9.1% vs 0.4% - 1:
no rush at all9.1% vs 0.4% - 1:
whenever you9.0% vs 0.8% - 0:
confirm that8.9% vs 0.7% - 0:
everything is8.7% vs 1.4% - 1:
wondering8.5% vs 1.2% - 4:
today i8.5% vs 1.9% - 0:
to confirm that8.2% vs 0.6% - 4:
nobody8.2% vs 1.4% - 0:
you know8.0% vs 1.4% - 1:
question about7.9% vs 1.7% - 1:
thought7.8% vs 1.9% - 4:
tomorrow morning7.2% vs 0.9% - 4:
court6.9% vs 0.8% - 1:
moment6.9% vs 1.2% - 0:
bye6.7% vs 1.0% - 0:
let you6.7% vs 0.7% - 0:
to let6.6% vs 0.9% - 4:
lose6.4% vs 0.7% - 0:
finally6.4% vs 1.5%
Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):
- all rows (16,633 rows): 4:
right now24.3% of class, 81% of its 1732 rows; 4:fix20.8% of class, 74% of its 1626 rows; 4:deadline18.0% of class, 71% of its 1451 rows; 4:fix this13.4% of class, 85% of its 911 rows; 1:rush at all10.3% of class, 75% of its 264 rows; 4:tonight10.1% of class, 80% of its 727 rows; 1:no rush at all9.1% of class, 74% of its 234 rows; 0:confirm that8.9% of class, 73% of its 363 rows; 4:today i8.5% of class, 71% of its 690 rows; 0:to confirm that8.2% of class, 76% of its 321 rows; 4:nobody8.2% of class, 76% of its 619 rows; 4:tomorrow morning7.2% of class, 81% of its 513 rows; 4:court6.7% of class, 81% of its 476 rows; 4:emergency6.7% of class, 74% of its 525 rows; 4:lose6.4% of class, 83% of its 445 rows; 4:lawyer5.9% of class, 78% of its 441 rows; 4:hour5.8% of class, 87% of its 384 rows; 4:breach5.6% of class, 86% of its 375 rows; 4:or i5.4% of class, 77% of its 400 rows; 4:noon5.2% of class, 84% of its 360 rows; 4:filing a5.1% of class, 70% of its 421 rows; 4:17 005.1% of class, 73% of its 403 rows; 0:confirm that the5.1% of class, 82% of its 185 rows; 4:unacceptable5.1% of class, 76% of its 383 rows; 4:midnight5.0% of class, 83% of its 344 rows; 4:consumer4.9% of class, 74% of its 381 rows; 4:immediate4.9% of class, 78% of its 358 rows; 4:shaking4.8% of class, 95% of its 293 rows; 0:say thank you4.6% of class, 72% of its 191 rows; 4:1 5 stars4.6% of class, 84% of its 315 rows; 0:to confirm that the4.6% of class, 83% of its 163 rows; 0:let you know that4.5% of class, 72% of its 188 rows; 4:00 today4.5% of class, 76% of its 345 rows; 4:5pm4.5% of class, 81% of its 318 rows; 4:posting4.5% of class, 74% of its 349 rows; 4:or we4.4% of class, 91% of its 280 rows; 4:my lawyer4.3% of class, 81% of its 308 rows; 4:this now4.3% of class, 94% of its 263 rows; 0:am writing to confirm4.3% of class, 81% of its 156 rows; 4:begging4.3% of class, 95% of its 260 rows - state.channel = email (6,611 rows): 4:
right now26.7% of class, 81% of its 736 rows; 4:fix22.6% of class, 76% of its 664 rows; 4:fix this15.3% of class, 85% of its 397 rows; 0:confirm that15.2% of class, 73% of its 252 rows; 0:to confirm that14.2% of class, 76% of its 226 rows; 4:tonight12.2% of class, 80% of its 339 rows; 4:emergency10.0% of class, 75% of its 299 rows; 1:rush at all9.2% of class, 75% of its 104 rows; 4:tomorrow morning9.2% of class, 78% of its 260 rows; 0:subject thank you9.0% of class, 76% of its 142 rows; 4:nobody9.0% of class, 75% of its 264 rows; 4:immediate8.9% of class, 77% of its 256 rows; 4:breach8.8% of class, 86% of its 226 rows; 4:17 008.2% of class, 70% of its 260 rows; 4:hour8.2% of class, 88% of its 208 rows; 0:to confirm that the8.2% of class, 85% of its 116 rows; 1:no rush at all8.1% of class, 75% of its 92 rows; 0:am writing to confirm8.0% of class, 81% of its 118 rows; 4:lose7.9% of class, 78% of its 224 rows; 0:writing to confirm that7.7% of class, 82% of its 112 rows; 0:you know that7.7% of class, 70% of its 131 rows; 4:00 today7.2% of class, 73% of its 220 rows; 4:court7.1% of class, 76% of its 208 rows; 4:unacceptable7.0% of class, 73% of its 213 rows; 4:noon6.6% of class, 83% of its 178 rows; 4:consumer6.5% of class, 70% of its 206 rows; 4:lawyer6.2% of class, 74% of its 187 rows; 4:shaking6.0% of class, 96% of its 140 rows; 4:filing a6.0% of class, 71% of its 187 rows; 0:subject thank you for5.9% of class, 77% of its 92 rows; 4:midnight5.7% of class, 79% of its 161 rows; 4:this immediately5.7% of class, 92% of its 138 rows; 4:is completely5.5% of class, 77% of its 160 rows; 4:complaint with5.5% of class, 82% of its 148 rows; 4:begging5.4% of class, 93% of its 129 rows; 4:this is a5.4% of class, 86% of its 138 rows; 4:solicitor5.3% of class, 70% of its 167 rows; 4:risk5.2% of class, 77% of its 151 rows; 4:pm today5.2% of class, 80% of its 143 rows; 4:calling5.1% of class, 80% of its 141 rows - state.channel = web form (3,009 rows): 4:
immediately23.8% of class, 89% of its 280 rows; 4:right now20.3% of class, 78% of its 271 rows; 4:this is19.6% of class, 71% of its 289 rows; 4:fix19.5% of class, 70% of its 288 rows; 4:deadline19.0% of class, 76% of its 261 rows; 4:fix this13.3% of class, 78% of its 177 rows; 4:legal10.8% of class, 73% of its 153 rows; 0:you know9.6% of class, 73% of its 66 rows; 1:rush at all9.2% of class, 70% of its 50 rows; 4:nobody9.1% of class, 79% of its 120 rows; 4:tonight9.0% of class, 78% of its 121 rows; 0:to confirm that8.2% of class, 77% of its 53 rows; 0:to let you know8.0% of class, 71% of its 56 rows; 0:successfully7.8% of class, 81% of its 48 rows; 4:tomorrow morning7.5% of class, 85% of its 92 rows; 4:pm7.4% of class, 75% of its 103 rows; 4:177.3% of class, 78% of its 98 rows; 4:court7.3% of class, 85% of its 89 rows; 4:00 today6.9% of class, 83% of its 87 rows; 4:noon6.9% of class, 85% of its 85 rows; 4:hospital6.8% of class, 71% of its 100 rows; 4:17 006.6% of class, 81% of its 85 rows; 0:a quick note6.6% of class, 75% of its 44 rows; 0:note to6.6% of class, 85% of its 39 rows; 0:let you know that6.4% of class, 76% of its 42 rows; 4:or i6.1% of class, 72% of its 87 rows; 0:to drop6.0% of class, 73% of its 41 rows; 4:lawyer5.7% of class, 79% of its 75 rows; 4:this immediately5.3% of class, 96% of its 57 rows; 0:to drop a5.2% of class, 79% of its 33 rows; 4:5pm5.2% of class, 86% of its 63 rows; 4:lose5.1% of class, 90% of its 59 rows; 4:unacceptable5.1% of class, 79% of its 67 rows; 0:a quick note to5.0% of class, 83% of its 30 rows; 0:wanted to drop5.0% of class, 86% of its 29 rows; 4:calling4.9% of class, 89% of its 57 rows; 4:tomorrow at4.9% of class, 85% of its 60 rows; 4:consumer4.8% of class, 77% of its 65 rows; 0:am writing to confirm4.8% of class, 77% of its 31 rows; 0:note to say4.8% of class, 80% of its 30 rows - state.channel = chat (2,529 rows): 4:
fix19.1% of class, 78% of its 237 rows; 4:right16.7% of class, 84% of its 194 rows; 4:right now15.8% of class, 90% of its 172 rows; 4:fix this9.6% of class, 90% of its 104 rows; 0:perfectly9.1% of class, 73% of its 48 rows; 4:filing8.4% of class, 81% of its 101 rows; 4:immediately8.0% of class, 91% of its 86 rows; 4:or i7.2% of class, 91% of its 77 rows; 4:twice7.2% of class, 78% of its 90 rows; 4:now or6.5% of class, 95% of its 66 rows; 4:tonight6.5% of class, 85% of its 74 rows; 4:this now6.2% of class, 100% of its 60 rows; 4:court6.1% of class, 88% of its 67 rows; 4:oh6.1% of class, 83% of its 71 rows; 4:posting5.7% of class, 85% of its 66 rows; 2:question5.3% of class, 71% of its 49 rows; 4:now i5.2% of class, 77% of its 66 rows; 4:cannot5.0% of class, 79% of its 62 rows; 4:fix this now5.0% of class, 100% of its 49 rows; 0:say the4.9% of class, 79% of its 24 rows; 4:calling4.7% of class, 87% of its 53 rows; 4:filing a4.6% of class, 85% of its 53 rows; 4:lose4.4% of class, 96% of its 45 rows; 4:if this4.2% of class, 77% of its 53 rows; 4:going4.1% of class, 82% of its 49 rows; 4:or we3.9% of class, 95% of its 40 rows; 4:deadline is3.8% of class, 92% of its 40 rows; 4:now or i3.8% of class, 97% of its 38 rows; 0:i wanted to say3.6% of class, 70% of its 20 rows; 0:perfect3.6% of class, 70% of its 20 rows; 4:midnight3.6% of class, 92% of its 38 rows; 4:reporting3.6% of class, 90% of its 39 rows; 4:down3.5% of class, 83% of its 41 rows; 4:lawyer3.5% of class, 77% of its 44 rows; 2:send me3.5% of class, 77% of its 30 rows; 2:copy3.3% of class, 85% of its 26 rows; 4:everywhere3.3% of class, 89% of its 36 rows; 4:hour3.3% of class, 97% of its 33 rows; 4:it now3.3% of class, 97% of its 33 rows; 4:nobody3.3% of class, 78% of its 41 rows - state.channel = phone transcript (1,846 rows): 4:
right now52.1% of class, 78% of its 449 rows; 1:whenever30.2% of class, 70% of its 105 rows; 4:fix19.3% of class, 74% of its 176 rows; 0:to leave18.1% of class, 80% of its 87 rows; 1:rush at all18.0% of class, 83% of its 53 rows; 1:whenever you18.0% of class, 72% of its 61 rows; 0:say thank you17.1% of class, 80% of its 82 rows; 4:deadline16.0% of class, 81% of its 133 rows; 4:hello is15.9% of class, 78% of its 138 rows; 0:wanted to leave15.8% of class, 90% of its 68 rows; 4:look i15.7% of class, 76% of its 140 rows; 1:no rush at all15.5% of class, 84% of its 45 rows; 4:anyone15.4% of class, 76% of its 137 rows; 0:leave a15.3% of class, 71% of its 83 rows; 4:nobody15.3% of class, 76% of its 136 rows; 0:calling to14.8% of class, 74% of its 77 rows; 0:to say thank you14.8% of class, 81% of its 70 rows; 4:listen13.6% of class, 87% of its 106 rows; 4:tonight13.6% of class, 84% of its 110 rows; 4:is anyone13.1% of class, 94% of its 94 rows; 4:o'clock13.1% of class, 79% of its 112 rows; 0:message to13.0% of class, 96% of its 52 rows; 4:fix this12.9% of class, 84% of its 103 rows; 4:please i12.8% of class, 92% of its 93 rows; 0:to leave a12.7% of class, 88% of its 56 rows; 4:by five11.6% of class, 91% of its 86 rows; 4:tomorrow morning11.4% of class, 81% of its 95 rows; 4:right now i11.3% of class, 75% of its 101 rows; 0:perfect11.1% of class, 81% of its 53 rows; 0:wanted to leave a11.1% of class, 93% of its 46 rows; 4:literally11.1% of class, 75% of its 100 rows; 4:clock11.0% of class, 84% of its 88 rows; 0:s all10.9% of class, 78% of its 54 rows; 4:today i10.7% of class, 82% of its 88 rows; 0:i wanted to say10.6% of class, 77% of its 53 rows; 4:shaking10.5% of class, 92% of its 77 rows; 4:immediately10.2% of class, 93% of its 74 rows; 4:to me10.2% of class, 77% of its 90 rows; 0:message to say10.1% of class, 100% of its 39 rows; 4:you have to10.1% of class, 97% of its 70 rows - state.channel = app review (1,430 rows): 4:
1 5 stars56.6% of class, 84% of its 315 rows; 4:right17.7% of class, 78% of its 107 rows; 4:right now15.1% of class, 88% of its 81 rows; 4:fix this14.3% of class, 88% of its 76 rows; 4:immediately13.8% of class, 88% of its 74 rows; 0:perfectly12.6% of class, 73% of its 55 rows; 4:this is12.3% of class, 83% of its 70 rows; 1:rush12.2% of class, 73% of its 33 rows; 1:no rush9.6% of class, 76% of its 25 rows; 4:tonight8.7% of class, 79% of its 52 rows; 4:lawyer7.7% of class, 88% of its 41 rows; 4:legal7.7% of class, 73% of its 49 rows; 4:nobody7.7% of class, 75% of its 48 rows; 1:whenever7.6% of class, 75% of its 20 rows; 4:times7.2% of class, 76% of its 45 rows; 4:this to6.4% of class, 79% of its 38 rows; 4:unacceptable6.4% of class, 88% of its 34 rows; 4:today i6.2% of class, 81% of its 36 rows; 4:calling6.0% of class, 85% of its 33 rows; 4:never6.0% of class, 70% of its 40 rows; 4:my lawyer5.7% of class, 90% of its 30 rows; 0:everything is5.7% of class, 82% of its 22 rows; 4:ago5.5% of class, 76% of its 34 rows; 4:broken5.5% of class, 70% of its 37 rows; 4:emergency5.5% of class, 81% of its 32 rows; 0:perfect5.4% of class, 81% of its 21 rows; 4:tomorrow morning5.3% of class, 83% of its 30 rows; 4:midnight5.1% of class, 83% of its 29 rows; 4:he4.9% of class, 72% of its 32 rows; 4:lose4.9% of class, 82% of its 28 rows; 4:absolute4.7% of class, 76% of its 29 rows; 4:consumer4.7% of class, 79% of its 28 rows; 4:dont4.7% of class, 76% of its 29 rows; 4:or we4.7% of class, 92% of its 24 rows; 4:reporting4.7% of class, 76% of its 29 rows; 4:this now4.7% of class, 96% of its 23 rows; 4:court4.5% of class, 75% of its 28 rows; 4:this to the4.5% of class, 91% of its 23 rows; 4:expires4.3% of class, 87% of its 23 rows; 4:5 stars unacceptable4.0% of class, 83% of its 23 rows - state.channel = social media reply (1,208 rows): 4:
fix this11.9% of class, 85% of its 54 rows; 4:deadline11.7% of class, 73% of its 62 rows; 4:twice11.1% of class, 75% of its 57 rows; 0:perfectly10.7% of class, 95% of its 21 rows; 2:do i10.5% of class, 74% of its 47 rows; 4:this now8.0% of class, 97% of its 32 rows; 2:loving7.5% of class, 74% of its 34 rows; 4:immediately7.3% of class, 90% of its 31 rows; 4:fix this now6.5% of class, 96% of its 26 rows; 2:loving the6.3% of class, 81% of its 26 rows; 4:5pm6.2% of class, 86% of its 28 rows; 4:lawyer6.0% of class, 79% of its 29 rows; 4:or i6.0% of class, 72% of its 32 rows; 4:emailed twice5.7% of class, 79% of its 28 rows; 4:now or5.7% of class, 88% of its 25 rows; 4:right5.7% of class, 73% of its 30 rows; 4:tonight5.7% of class, 71% of its 31 rows; 2:hi i5.4% of class, 82% of its 22 rows; 2:how do i5.1% of class, 74% of its 23 rows; 4:court4.9% of class, 90% of its 21 rows; 4:locked4.9% of class, 73% of its 26 rows; 4:right now4.9% of class, 83% of its 23 rows; 4:or we4.7% of class, 90% of its 20 rows; 4:this is4.7% of class, 75% of its 24 rows; 4:003.9% of class, 75% of its 20 rows
Shortcut models
Predicting the label class on test (1,814 rows). Chance 20.0%, majority class ('4') 35.1%; balanced chance 20.0%.
| model (logistic regression, trained on the train sample) | test accuracy | balanced accuracy (mean recall) |
|---|---|---|
| bag of words, whole state (words and word pairs) | 70.2% | 60.0% |
bag of words, main text only (message) |
72.3% | 62.6% |
| surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) | 43.7% | 33.4% |
| surface features of the main text only | 42.8% | 32.7% |
Strongest single surface features (logistic regression on one feature, balanced accuracy on test):
| feature | accuracy | balanced accuracy |
|---|---|---|
| ends_? | 35.9% | 21.3% |
| count_? | 34.5% | 20.7% |
| chars(log) | 35.2% | 20.1% |
| words(log) | 35.1% | 20.0% |
| upper_ratio | 35.1% | 20.0% |
| digit_ratio | 35.1% | 20.0% |
| nonascii_ratio | 35.1% | 20.0% |
| emoji | 35.1% | 20.0% |
| newlines | 35.1% | 20.0% |
| html_tag | 35.1% | 20.0% |
Other state fields alone (predicting the label class on test from one field, without the main text):
| field | treated as | accuracy | balanced accuracy |
|---|---|---|---|
| company | text: bag of words / length+empty | 35.1% / 35.1% | 20.0% / 20.0% |
| channel | categorical, 6 values | 35.1% | 20.0% |
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
- Train rows identical in the whole prompt (state, options, instructions): 0.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 0 (0.0%) | |
| calibration | 0 (0.0%) | |
| test | 0 (0.0%) |
Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)
- Train: 0 near-duplicate pairs; 0 rows (0.0%) sit in 0 clusters; largest cluster 0; excess rows (cluster size − 1) 0 (0.0%).
- Clusters with more than one label class: 0 (0 rows).
- Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 0 (0.0%), test 0 (0.0%)
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 0 |
| dev | 0 | 0 |
| calibration | 0 | 0 |
| test | 0 | 0 |
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 0 (0.0%) | |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 1 (0.0%) | 2 0.0% |
| chat preamble ('Here is/are...', 'Sure!') | 0 (0.0%) | |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 0 (0.0%) | |
| encoding garbage (mojibake/replacement char) | 0 (0.0%) | |
| HTML tag | 511 (3.1%) | 0 4.6%, 1 3.9%, 2 1.9%, 3 2.4%, 4 3.1% |
| HTML entity | 0 (0.0%) | |
| base64-like run (40+ chars) | 1 (0.0%) | 0 0.0% |
| URL | 8 (0.0%) | 0 0.2%, 1 0.1%, 2 0.0%, 4 0.0% |
'As an AI' / refusal:
triage-c290-m140-urgency(2): Subject: Concern regarding stonework at Hrubý Rohozec ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I hope this message finds you well, though I must confes…HTML tag:
triage-c105-m100-urgency(0): Subject: URGENT - Confirmation of received replacement card - Account 4401 22XX XXXX 8812 ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I am writing to conf… |triage-c218-m004-urgency(0): Subject: Re: Corporate Fleet Inquiry ⏎ ⏎ Noted! As mentioned when I emailed last week, we've decided against the bulk employee plan for no… |triage-c185-m219-urgency(1): Subject: URGENT - Fwd: Missing artist catalog request ⏎ ⏎ ----- Forwarded message ----- ⏎ From: Janusz Lewandowski <j.lewandowski@interia.…base64-like run (40+ chars):
triage-c056-m162-urgency(0): Subject: URGENT: Your store domain will expire!!! ⏎ ⏎ Dear Beloved Manager, ⏎ ⏎ I am writing to you today with a heavy heart because your…URL:
triage-c133-m096-urgency(0): Subject: Strategic Partnership Proposal - Global Digital Infrastructure ⏎ ⏎ Dear Sir/Madam, ⏎ ⏎ I hope this email finds you well. My name… |triage-c146-m064-urgency(2): Oggetto: R: Richiesta informazioni piano anticipato ⏎ ⏎ ---------- Messaggio inoltrato ---------- ⏎ Da: Assistenza Clienti <assistenza@ser… |triage-c238-m151-urgency(0): Subject: Meine Bougainvillea blüht wunderschön. ⏎ ⏎ Liebes Team, ⏎ ⏎ ich wollte Ihnen nur mitteilen, dass Ihre Tipps zum Zurückschneiden …Possibly cut off: 5,505 of 11,135 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: 0 52.1%, 1 54.9%, 2 53.8%, 3 50.5%, 4 43.3%
triage-c156-m218-urgency: …ine is maintained. ⏎ ⏎ Awaiting your direction. ⏎ ⏎ Marcus Teo ⏎ Director of Facilities, Meridian Holdings Pte Ltd ⏎ +65 6823 9…triage-c089-m048-urgency: …nita, your request to revise margins for Q4 (ref PO-2291) has been approved. ⏎ ⏎ Anita Desai ⏎ Store Manager ⏎ +91 20 2555 0198triage-c088-m025-urgency: …plan befor next month starts so we can add more users. Plese tell me how. ⏎ Marco Bianchi, Sales Director, +39 06 492 8173
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- ×175: "---------- Forwarded message ---------" (4 65, 0 34, 1 27)
- ×166: "Any unauthorized review, use, disclosure, or distribution is strictly prohibited." (4 65, 0 35, 1 25)
- ×159: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and intended solely for the use of the individu…" (2 39, 4 34, 1 34)
- ×148: "----- Forwarded message -----" (4 85, 0 22, 1 16)
- ×132: "I hope this message finds you well." (2 55, 1 30, 3 22)
- ×119: "-----Original Message-----" (4 55, 0 21, 2 16)
- ×118: "If you have received this email in error, please notify the sender immediately and delete this email from your system." (4 37, 2 26, 1 21)
- ×116: "If you are not the intended recipient, please delete all copies and notify the sender immediately." (4 58, 0 19, 2 17)
- ×107: "Topic: Billing and Invoices" (4 32, 0 21, 3 21)
- ×107: "If you have received this email in error, please notify the sender immediately and delete this message from your system." (0 25, 4 25, 2 22)
- ×91: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and may also be privileged." (4 33, 0 20, 2 18)
- ×81: "CONFIDENTIALITY NOTICE: This email and any attachments are strictly confidential and intended solely for the addressee." (4 24, 2 19, 0 18)
- ×66: "I hope this email finds you well." (2 27, 1 24, 3 8)
- ×63: "----- End forwarded message -----" (4 38, 0 10, 3 5)
- ×60: "--- Forwarded message ---" (4 31, 2 10, 0 10)
6. Samples
20 random train rows per kind: triage-urgency-samples.txt. Reading notes are in the findings above.
QA: triage-sentiment
Checked 2026-09-30 13:36 by adapters/qa/qa.py (READY file READY-triage, ).
Verdict: PASS WITH NOTES (full notes in triage-needs_human.md and triage.md)
Automatic flags (for the reviewer to judge; not all are problems)
- lclass '2' share varies across splits by more than 5 points: train 13.7%, dev 10.2%, calibration 17.2%, test 13.0%
- 74 strong phrase flags (see list): review whether they are meaning or leakage
- 60 standard phrase flags (≥2% of a class, mostly that class): review
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 16,633 | 260 | train.jsonl |
| dev | 342 | 5 | dev.jsonl |
| calibration | 302 | 5 | calibration.jsonl |
| test | 1,814 | 30 | test.jsonl |
- Train sha256:
3aba76c53cd16a1c2822d1b5e4090d183cc77fd94a233898c53501f25c3b16c0(READY file gives no checksum) - Main text field (the text the phrase and length checks use):
state.message. - Label classes: 0, 1, 2, 3, 4 (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): sentiment.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| 0 | 35.6% | 33.0% | 32.8% | 36.5% | 5,924 |
| 1 | 9.8% | 10.5% | 7.6% | 10.3% | 1,636 |
| 2 | 13.7% | 10.2% | 17.2% | 13.0% | 2,285 |
| 3 | 20.8% | 24.0% | 21.9% | 20.0% | 3,452 |
| 4 | 20.1% | 22.2% | 20.5% | 20.3% | 3,336 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| sentiment | 100.0% | 100.0% | 100.0% | 100.0% | 16,633 |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | estimate: characters / 3 (upper bound for English) | 203 | 877 | 1330 | 0 |
| dev | estimate: characters / 3 (upper bound for English) | 223 | 878 | 1086 | 0 |
| calibration | estimate: characters / 3 (upper bound for English) | 213 | 971 | 1148 | 0 |
| test | estimate: characters / 3 (upper bound for English) | 200 | 842 | 1161 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| channel, company, message | 16,633 (100.0%) |
Instructions (train)
- Canonical (the most common text) 70.3%, reworded 26.8% (62 distinct rewordings), none 3.0%. Target about 70 / 27 / 3.
- Canonical text: "How does the customer feel?"
sourceinstruction tag: canonical 70.3%, none 3.0%, variant-17 0.6%, variant-23 0.6%, variant-21 0.5%, variant-40 0.5%
| class | canonical | none |
|---|---|---|
| 0 | 70.4% | 2.8% |
| 1 | 70.7% | 3.2% |
| 2 | 70.3% | 3.1% |
| 3 | 70.5% | 2.9% |
| 4 | 69.5% | 3.2% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 16,633 train rows; models are scored on the full test file (1,814 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | 0 | 5924 | 140 | 421 | 1357 | 612 |
| train | 1 | 1636 | 131 | 415 | 1349 | 608 |
| train | 2 | 2285 | 113 | 377 | 1271 | 553 |
| train | 3 | 3452 | 132 | 433 | 1369 | 616 |
| train | 4 | 3336 | 146 | 561 | 1518 | 728 |
| test | 0 | 662 | 141 | 417 | 1364 | 614 |
| test | 1 | 186 | 104 | 394 | 1409 | 570 |
| test | 2 | 235 | 103 | 396 | 1315 | 577 |
| test | 3 | 363 | 111 | 414 | 1389 | 590 |
| test | 4 | 368 | 153 | 521 | 1523 | 733 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| sentiment | 16633 | 433 | 628 | 145 | 0 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 35.6%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| tone | 25 | 77.6% | politely formal: 3 67.4%, 4 20.2%; cheerful: 3 65.8%, 4 30.8%; polite and friendly: 3 70.2%, 4 23.6%; warm: 3 67.0%, 4 30.0%; disappointed: 0 63.4%, 1 33.9%; appreciative: 3 66.4%, 4 29.7% |
| human_reason | 7 | 51.7% | manual: 0 45.4%, 3 22.3%; legal_safety: 0 56.4%, 3 23.5%; routine: 4 34.2%, 2 28.1%; acknowledge: 4 40.6%, 3 26.7%; escalation: 0 88.0%, 1 10.8%; no_reply: 4 52.3%, 3 24.1% |
| situation | 17 | 47.7% | null: 0 34.0%, 3 21.4%; the writer reports or confirms something that needs no decision, such as that a delivery arrived, a payment was made, a problem sorted itself out, or a form was sent: 4 43.2%, 3 26.2%; the writer is about to take it further and says what they will do next, for example cancel, go to a regulator or ombudsman, post a public review, contact the press, or make a formal complain… |
| other_kind | 9 | 38.2% | null: 0 38.1%, 3 20.4%; a job application or question about working there: 4 35.6%, 3 29.2%; a message clearly meant for a different organisation: 4 51.9%, 2 16.6%; a sales pitch from another business offering its services: 2 36.2%, 3 23.7%; a request to sponsor or donate to a local event: 4 40.7%, 3 25.9%; a journalist or student asking for an interview or information for a project: 4 37.9%, 3 3… |
| channel | 6 | 35.6% | email: 0 34.6%, 4 22.0%; web form: 0 37.6%, 3 20.5%; chat: 0 36.1%, 3 22.1%; phone transcript: 0 33.0%, 3 19.7%; app review: 0 37.1%, 3 21.7%; social media reply: 0 37.4%, 3 21.4% |
| language | 11 | 35.6% | English: 0 35.7%, 3 20.9%; Danish: 0 33.3%, 4 22.4%; Czech: 0 35.0%, 3 21.5%; French: 0 37.1%, 4 22.9%; Swedish: 0 31.3%, 3 23.9%; Dutch: 0 36.7%, 3 19.6% |
| length_target | 17 | 35.6% | 60: 0 37.7%, 3 20.9%; 120: 0 34.8%, 4 23.6%; 25: 0 37.2%, 3 21.1%; 220: 0 32.0%, 4 27.5%; 50: 0 34.2%, 3 21.0%; 10: 0 36.2%, 2 19.4% |
| prompt_version | 3 | 35.6% | 3: 0 36.3%, 4 21.5%; 2: 0 33.8%, 3 21.5%; 1: 0 42.2%, 3 17.2% |
| decorrelated_by | 2 | 35.6% | qwen3.8-flash: 0 31.9%, 4 25.3%; null: 0 39.6%, 3 19.5% |
| regenerated | 2 | 35.6% | false: 0 35.1%, 3 20.7%; true: 0 37.2%, 4 21.0% |
| target_kind | 2 | 35.6% | soft: 0 33.4%, 3 24.1%; hard: 0 38.4%, 4 21.5% |
Formatting by label class (main text, share of rows)
| feature | 0 | 1 | 2 | 3 | 4 | |
|---|---|---|---|---|---|---|
| ends with ? | 3.6% | 7.3% | 7.7% | 4.0% | 2.6% | |
| ends with . | 43.3% | 37.5% | 36.1% | 33.4% | 34.0% | |
| ends with ! | 13.2% | 14.1% | 14.8% | 13.0% | 12.7% | |
| no end punctuation | 39.8% | 41.1% | 41.3% | 49.1% | 50.2% | |
| starts lowercase | 7.0% | 7.9% | 7.6% | 5.6% | 4.9% | |
| all lowercase | 4.0% | 5.0% | 5.2% | 4.3% | 4.0% | |
| has a digit | 91.0% | 86.6% | 87.6% | 87.1% | 84.7% | |
| has newline | 75.9% | 74.6% | 75.4% | 77.7% | 80.4% | |
| has quotes | 9.9% | 5.3% | 3.9% | 4.3% | 3.4% | |
| has markup (HTML/markdown) | 2.7% | 2.1% | 2.7% | 3.1% | 4.8% | |
| has URL | 0.0% | 0.1% | 0.1% | 0.0% | 0.1% | |
| non-ASCII | 53.0% | 47.6% | 42.5% | 53.0% | 53.9% | |
| non-Latin script | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | |
| emoji | 4.9% | 4.9% | 3.2% | 7.0% | 6.5% | |
| ALL-CAPS word (4+) | 16.3% | 16.6% | 15.6% | 15.8% | 16.5% | |
| contains ' - ' or — | 18.8% | 13.4% | 12.4% | 18.0% | 20.3% |
Same, by row kind
| feature | sentiment |
|---|---|
| ends with ? | 4% |
| ends with . | 38% |
| ends with ! | 13% |
| no end punctuation | 44% |
| starts lowercase | 6% |
| all lowercase | 4% |
| has a digit | 88% |
| has newline | 77% |
| has quotes | 6% |
| has markup (HTML/markdown) | 3% |
| has URL | 0% |
| non-ASCII | 51% |
| non-Latin script | 0% |
| emoji | 5% |
| ALL-CAPS word (4+) | 16% |
| contains ' - ' or — | 18% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, 0: fix z=33 22.4% vs 2.8%; now z=32 49.4% vs 20.9%; today z=30 44.2% vs 18.4%; immediately z=27 22.1% vs 6.2%; right z=25 24.5% vs 8.6%; cannot z=25 17.8% vs 4.9%; tomorrow z=22 21.0% vs 8.1%; nobody z=21 8.8% vs 0.9%; if z=20 43.5% vs 26.0%; will z=20 25.6% vs 12.5%; filing z=19 9.6% vs 2.2%; this z=19 65.1% vs 44.4%; money z=19 9.6% vs 2.4%; twice z=19 8.6% vs 1.8%; tonight z=18 8.7% vs 2.0%; completely z=17 12.0% vs 4.4%; oh z=17 10.8% vs 3.7%; lawyer z=17 5.9% vs 0.9%; lose z=17 5.9% vs 0.9%; consumer z=16 5.5% vs 0.6%
Words, 1: why z=15 14.1% vs 4.5%; disappointed z=13 4.8% vs 0.7%; frustrating z=13 3.7% vs 0.3%; still z=13 18.1% vs 8.2%; disappointing z=12 3.1% vs 0.2%; quite z=12 7.3% vs 2.1%; anyway z=12 13.9% vs 5.9%; annoying z=11 2.5% vs 0.1%; but z=11 41.0% vs 27.0%; confusing z=10 2.3% vs 0.2%; took z=10 5.4% vs 1.6%; yeah z=9 9.8% vs 4.4%; confused z=9 3.5% vs 0.8%; long z=9 4.5% vs 1.4%; not z=8 30.5% vs 20.7%; honestly z=8 7.0% vs 3.0%; need z=8 24.8% vs 16.2%; annoyed z=8 1.5% vs 0.0%; weeks z=8 6.8% vs 3.0%; finally z=8 5.3% vs 2.1%
Words, 2: advise z=14 5.3% vs 1.2%; confirm z=13 12.3% vs 5.8%; hey z=12 4.5% vs 1.2%; current z=12 6.7% vs 2.5%; standard z=11 8.0% vs 3.6%; regarding z=11 15.8% vs 9.7%; noting z=11 2.1% vs 0.2%; require z=10 4.3% vs 1.4%; clarify z=10 2.2% vs 0.4%; noted z=10 2.9% vs 0.7%; form z=9 7.4% vs 3.7%; required z=9 4.6% vs 1.8%; documentation z=9 3.4% vs 1.1%; sure z=9 7.7% vs 4.0%; 3 z=9 12.5% vs 8.0%; madam z=9 6.2% vs 3.0%; fine z=9 6.8% vs 3.5%; additionally z=9 2.7% vs 0.8%; sir z=9 6.3% vs 3.2%; further z=8 4.2% vs 1.8%
Words, 3: could z=30 36.5% vs 11.6%; hope z=30 19.0% vs 2.7%; much z=26 37.1% vs 14.3%; thanks z=24 26.9% vs 9.3%; appreciate z=22 13.6% vs 3.0%; kindly z=22 10.3% vs 1.4%; love z=20 14.5% vs 4.2%; lovely z=20 10.5% vs 2.3%; warm z=19 9.5% vs 1.9%; team z=18 28.9% vs 14.3%; great z=18 10.2% vs 2.6%; having z=17 11.0% vs 3.2%; thank z=17 36.6% vs 20.5%; good z=17 11.1% vs 3.5%; so z=17 51.7% vs 32.5%; always z=16 10.2% vs 3.1%; would z=16 17.7% vs 7.8%; well z=16 13.0% vs 4.9%; hello z=15 27.1% vs 15.0%; finds z=15 5.2% vs 0.8%
Words, 4: thank z=37 58.2% vs 15.2%; much z=36 48.6% vs 11.6%; absolutely z=34 29.5% vs 4.0%; wonderful z=30 27.7% vs 5.1%; thrilled z=30 19.9% vs 1.0%; such z=30 20.3% vs 2.2%; happy z=29 20.6% vs 2.8%; so z=29 69.8% vs 28.1%; perfectly z=28 16.6% vs 1.2%; warmest z=27 16.9% vs 1.6%; amazing z=26 15.4% vs 0.8%; team z=25 37.2% vs 12.4%; everything z=24 29.5% vs 8.9%; best z=24 17.5% vs 3.4%; say z=22 21.8% vs 6.2%; incredibly z=21 10.0% vs 0.8%; made z=20 12.1% vs 2.1%; quick z=20 13.4% vs 2.9%; guys z=19 14.4% vs 3.8%; grateful z=18 11.1% vs 2.3%
2–4-word phrases, 0: right now z=33 22.4% vs 3.8%; this is z=26 20.1% vs 6.3%; fix this z=25 14.3% vs 0.6%; if this z=23 11.0% vs 2.1%; i cannot z=20 9.8% vs 2.4%; today i z=20 8.5% vs 1.8%; i will z=19 11.4% vs 3.7%; is not z=19 9.7% vs 2.6%; tomorrow morning z=18 6.6% vs 1.2%; this to z=17 5.8% vs 0.6%; filing a z=17 5.8% vs 0.7%; or i z=16 6.4% vs 0.2%; my lawyer z=15 4.4% vs 0.5%; complaint with z=15 4.7% vs 0.2%; is completely z=15 4.2% vs 0.4%; a formal z=14 8.0% vs 3.2%; or we z=14 3.9% vs 0.5%; if i z=14 7.2% vs 2.7%; i will be z=14 3.9% vs 0.6%; filing a formal z=14 4.0% vs 0.7%
2–4-word phrases, 1: 2 5 z=11 6.5% vs 1.8%; it took z=11 2.9% vs 0.1%; 2 5 stars z=11 5.8% vs 1.6%; understand why z=10 2.8% vs 0.5%; hi yeah z=9 3.2% vs 0.7%; need to know z=9 4.4% vs 1.3%; to know if z=9 3.4% vs 0.8%; need to know if z=9 2.6% vs 0.5%; am quite z=9 1.8% vs 0.0%; i am quite z=9 1.8% vs 0.0%; caller um hi yeah z=9 2.6% vs 0.5%; um hi yeah z=9 2.6% vs 0.5%; anyway the z=9 3.1% vs 0.7%; but the z=8 7.4% vs 3.3%; need to z=8 12.1% vs 6.6%; caller um z=8 7.1% vs 3.2%; caller um hi z=8 4.0% vs 1.3%; can you z=8 5.0% vs 1.9%; um hi z=8 4.2% vs 1.4%; i need to z=8 6.9% vs 3.2%
2–4-word phrases, 2: not sure z=16 5.0% vs 0.7%; 3 5 stars z=16 4.9% vs 0.7%; 3 5 z=15 5.3% vs 1.0%; sure if z=13 3.0% vs 0.4%; please advise z=12 3.2% vs 0.6%; no further z=12 2.6% vs 0.3%; need to z=11 11.3% vs 6.5%; not sure if z=11 2.1% vs 0.3%; please confirm z=11 3.2% vs 0.8%; you for your time z=11 2.0% vs 0.2%; dear sir z=10 6.3% vs 2.9%; am not sure z=10 1.9% vs 0.2%; i am not sure z=10 1.9% vs 0.2%; do i z=10 4.7% vs 1.9%; to confirm z=10 5.6% vs 2.6%; i need to z=10 6.1% vs 3.1%; or if z=9 3.9% vs 1.6%; advise on z=9 1.6% vs 0.2%; madam i z=9 5.4% vs 2.7%; i am not z=9 3.9% vs 1.6%
2–4-word phrases, 3: so much z=25 30.5% vs 10.7%; thanks so much z=24 11.8% vs 1.3%; thanks so z=24 11.8% vs 1.3%; i hope z=23 13.1% vs 2.1%; could you z=22 19.5% vs 6.0%; hope you z=22 10.1% vs 1.3%; much for your z=21 9.3% vs 1.0%; having a z=21 9.3% vs 1.1%; for your z=20 18.2% vs 6.0%; much for z=19 16.7% vs 5.5%; warm regards z=18 8.1% vs 1.3%; so much for your z=18 6.9% vs 0.8%; i hope you z=17 6.8% vs 1.1%; so much for z=17 13.7% vs 4.8%; you kindly z=16 5.7% vs 0.7%; love the z=16 5.4% vs 0.7%; thank you z=16 35.8% vs 19.9%; hope you are z=16 5.8% vs 0.9%; 4 5 z=16 5.6% vs 0.8%; team i hope z=16 5.4% vs 0.7%
2–4-word phrases, 4: thank you z=33 56.3% vs 15.0%; so much z=32 40.3% vs 8.4%; you so z=31 28.4% vs 4.2%; thank you so z=31 27.9% vs 3.9%; you so much z=29 25.6% vs 3.9%; thank you so much z=29 25.3% vs 3.8%; to say z=25 18.6% vs 2.6%; such a z=24 13.8% vs 1.3%; much for z=22 20.8% vs 4.6%; so much for z=22 18.4% vs 3.7%; again for z=21 11.2% vs 1.0%; you so much for z=21 14.6% vs 2.5%; team i z=21 16.4% vs 3.3%; warmest regards z=20 10.3% vs 1.0%; so happy z=19 9.1% vs 0.6%; absolutely thrilled z=19 9.5% vs 0.3%; i wanted to z=19 11.2% vs 1.8%; i wanted z=19 11.3% vs 1.8%; team i am z=19 9.5% vs 1.1%; guys are z=18 8.7% vs 0.9%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- 4:
much48.6% vs 11.6% - 4:
so much40.3% vs 8.4% - 4:
absolutely29.5% vs 4.0% - 4:
you so28.4% vs 4.2% - 4:
thank you so27.9% vs 3.9% - 4:
wonderful27.7% vs 5.1% - 4:
you so much25.6% vs 3.9% - 4:
thank you so much25.3% vs 3.8% - 0:
fix22.4% vs 2.8% - 0:
right now22.4% vs 3.8% - 4:
much for20.8% vs 4.6% - 4:
happy20.6% vs 2.8% - 4:
such20.3% vs 2.2% - 4:
thrilled19.9% vs 1.0% - 3:
hope19.0% vs 2.7% - 4:
to say18.6% vs 2.6% - 4:
so much for18.4% vs 3.7% - 4:
best17.5% vs 3.4% - 4:
warmest16.9% vs 1.6% - 4:
perfectly16.6% vs 1.2% - 4:
team i16.4% vs 3.3% - 4:
amazing15.4% vs 0.8% - 4:
you so much for14.6% vs 2.5% - 0:
fix this14.3% vs 0.6% - 4:
such a13.8% vs 1.3% - 3:
appreciate13.6% vs 3.0% - 4:
quick13.4% vs 2.9% - 3:
i hope13.1% vs 2.1% - 4:
made12.1% vs 2.1% - 3:
thanks so11.8% vs 1.3% - 3:
thanks so much11.8% vs 1.3% - 4:
i wanted11.3% vs 1.8% - 4:
i wanted to11.2% vs 1.8% - 4:
again for11.2% vs 1.0% - 4:
grateful11.1% vs 2.3% - 0:
if this11.0% vs 2.1% - 3:
lovely10.5% vs 2.3% - 4:
warmest regards10.3% vs 1.0% - 3:
kindly10.3% vs 1.4% - 3:
hope you10.1% vs 1.3%
Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):
- all rows (16,633 rows): 0:
fix22.4% of class, 82% of its 1626 rows; 0:right now22.4% of class, 77% of its 1732 rows; 4:such20.3% of class, 70% of its 963 rows; 4:thrilled19.9% of class, 83% of its 799 rows; 4:warmest16.9% of class, 73% of its 772 rows; 4:perfectly16.6% of class, 77% of its 717 rows; 4:amazing15.4% of class, 84% of its 615 rows; 0:fix this14.3% of class, 93% of its 911 rows; 4:such a13.8% of class, 73% of its 628 rows; 3:thanks so much11.8% of class, 70% of its 579 rows; 4:again for11.2% of class, 75% of its 500 rows; 0:if this11.0% of class, 74% of its 872 rows; 4:warmest regards10.3% of class, 71% of its 482 rows; 4:incredibly10.0% of class, 75% of its 447 rows; 4:absolutely thrilled9.5% of class, 87% of its 362 rows; 3:much for your9.3% of class, 71% of its 454 rows; 4:so happy9.1% of class, 80% of its 379 rows; 0:nobody8.8% of class, 85% of its 619 rows; 0:twice8.6% of class, 73% of its 698 rows; 4:you guys are8.5% of class, 71% of its 400 rows; 0:today i8.5% of class, 73% of its 690 rows; 4:you again8.2% of class, 74% of its 367 rows; 4:thank you again7.8% of class, 82% of its 316 rows; 4:incredible6.9% of class, 94% of its 245 rows; 4:perfect6.9% of class, 80% of its 287 rows; 4:gratitude6.7% of class, 92% of its 242 rows; 4:the best6.7% of class, 82% of its 273 rows; 0:tomorrow morning6.6% of class, 76% of its 513 rows; 0:unacceptable6.4% of class, 99% of its 383 rows; 0:or i6.4% of class, 94% of its 400 rows; 4:thank you again for6.2% of class, 81% of its 257 rows; 4:thrilled with5.9% of class, 86% of its 230 rows; 0:lawyer5.9% of class, 79% of its 441 rows; 0:lose5.9% of class, 78% of its 445 rows; 0:filing a5.8% of class, 81% of its 421 rows; 0:this to5.8% of class, 84% of its 407 rows; 0:consumer5.5% of class, 85% of its 381 rows; 4:for making5.5% of class, 85% of its 215 rows; 0:hour5.4% of class, 83% of its 384 rows; 0:1 5 stars5.3% of class, 99% of its 315 rows - state.channel = email (6,611 rows): 4:
warmest34.2% of class, 73% of its 680 rows; 4:such30.4% of class, 70% of its 628 rows; 4:thrilled26.7% of class, 82% of its 474 rows; 0:right now25.0% of class, 78% of its 736 rows; 0:this is24.5% of class, 72% of its 780 rows; 0:fix24.4% of class, 84% of its 664 rows; 4:warmest regards21.2% of class, 71% of its 434 rows; 4:such a20.0% of class, 72% of its 402 rows; 4:perfectly19.7% of class, 76% of its 374 rows; 4:again for19.2% of class, 75% of its 374 rows; 3:much for your16.7% of class, 72% of its 310 rows; 0:fix this16.3% of class, 94% of its 397 rows; 4:incredibly15.7% of class, 73% of its 314 rows; 4:absolutely thrilled14.9% of class, 86% of its 251 rows; 0:if this14.6% of class, 72% of its 464 rows; 4:you again13.3% of class, 71% of its 270 rows; 4:gratitude13.1% of class, 94% of its 202 rows; 4:thank you again12.6% of class, 79% of its 232 rows; 4:amazing12.4% of class, 85% of its 213 rows; 0:today i12.2% of class, 71% of its 393 rows; 0:oh11.9% of class, 83% of its 328 rows; 4:thank you again for11.1% of class, 77% of its 208 rows; 0:tonight10.5% of class, 71% of its 339 rows; 3:thanks so much10.3% of class, 73% of its 190 rows; 4:i am so10.1% of class, 72% of its 204 rows; 0:nobody9.8% of class, 85% of its 264 rows; 4:dear team i am9.3% of class, 73% of its 186 rows; 4:subject thank you9.2% of class, 94% of its 142 rows; 0:unacceptable9.2% of class, 99% of its 213 rows; 4:incredible8.8% of class, 93% of its 138 rows; 4:a quick8.5% of class, 72% of its 171 rows; 0:breach8.4% of class, 85% of its 226 rows; 4:so happy8.3% of class, 80% of its 151 rows; 4:to share8.3% of class, 77% of its 158 rows; 0:tomorrow morning8.3% of class, 73% of its 260 rows; 4:delighted8.2% of class, 75% of its 161 rows; 0:expect8.2% of class, 81% of its 231 rows; 0:this to8.1% of class, 81% of its 230 rows; 4:i am absolutely thrilled8.1% of class, 91% of its 129 rows; 4:for making8.0% of class, 88% of its 133 rows - state.channel = web form (3,009 rows): 4:
such22.6% of class, 74% of its 184 rows; 4:thrilled22.1% of class, 85% of its 157 rows; 0:immediately22.0% of class, 89% of its 280 rows; 0:this is21.0% of class, 82% of its 289 rows; 3:hope20.8% of class, 76% of its 168 rows; 0:fix20.4% of class, 80% of its 288 rows; 4:to say19.7% of class, 77% of its 155 rows; 0:right now18.8% of class, 78% of its 271 rows; 4:perfectly16.9% of class, 79% of its 129 rows; 4:such a16.1% of class, 80% of its 121 rows; 3:i hope15.3% of class, 75% of its 126 rows; 0:fix this14.0% of class, 89% of its 177 rows; 0:if this12.6% of class, 79% of its 180 rows; 3:hope you12.2% of class, 74% of its 102 rows; 4:amazing12.1% of class, 74% of its 98 rows; 3:having a11.7% of class, 77% of its 94 rows; 4:again for10.8% of class, 74% of its 88 rows; 4:absolutely thrilled10.6% of class, 91% of its 70 rows; 4:warmest10.3% of class, 73% of its 85 rows; 0:nobody9.5% of class, 89% of its 120 rows; 0:message my9.2% of class, 83% of its 125 rows; 0:oh8.9% of class, 87% of its 116 rows; 4:i am so8.6% of class, 76% of its 68 rows; 0:tonight8.1% of class, 76% of its 121 rows; 0:message oh7.9% of class, 89% of its 100 rows; 4:thrilled with7.6% of class, 94% of its 49 rows; 4:the best7.5% of class, 80% of its 56 rows; 0:or i7.4% of class, 97% of its 87 rows; 4:you again7.3% of class, 88% of its 50 rows; 0:tomorrow morning6.8% of class, 84% of its 92 rows; 4:thank you again6.8% of class, 93% of its 44 rows; 4:to let you know6.8% of class, 73% of its 56 rows; 4:so happy6.6% of class, 73% of its 55 rows; 4:team i am6.6% of class, 77% of its 52 rows; 4:thrilled to6.3% of class, 73% of its 52 rows; 0:fixed6.2% of class, 76% of its 92 rows; 3:your help6.2% of class, 72% of its 53 rows; 4:i am absolutely thrilled6.1% of class, 90% of its 41 rows; 4:you guys are6.1% of class, 77% of its 48 rows; 0:filing a5.9% of class, 77% of its 87 rows - state.channel = chat (2,529 rows): 0:
now53.8% of class, 73% of its 676 rows; 0:fix21.6% of class, 83% of its 237 rows; 4:amazing20.3% of class, 84% of its 102 rows; 0:right16.4% of class, 77% of its 194 rows; 4:absolutely16.1% of class, 86% of its 79 rows; 0:right now15.9% of class, 84% of its 172 rows; 4:you guys are14.7% of class, 74% of its 84 rows; 4:best14.4% of class, 72% of its 85 rows; 0:this is11.6% of class, 78% of its 136 rows; 4:so happy11.6% of class, 78% of its 63 rows; 0:fix this10.6% of class, 93% of its 104 rows; 3:kindly9.3% of class, 71% of its 73 rows; 4:perfectly8.7% of class, 77% of its 48 rows; 3:love the8.6% of class, 71% of its 68 rows; 0:filing8.5% of class, 77% of its 101 rows; 0:immediately8.5% of class, 91% of its 86 rows; 0:twice8.5% of class, 87% of its 90 rows; 0:or i8.0% of class, 95% of its 77 rows; 3:hope7.3% of class, 79% of its 52 rows; 0:now or7.1% of class, 98% of its 66 rows; 3:you kindly7.0% of class, 72% of its 54 rows; 0:oh6.9% of class, 89% of its 71 rows; 0:posting6.7% of class, 92% of its 66 rows; 0:this now6.5% of class, 98% of its 60 rows; 4:i wanted to6.4% of class, 71% of its 38 rows; 4:i am so6.1% of class, 74% of its 35 rows; 4:you guys are amazing6.1% of class, 90% of its 29 rows; 4:the best5.9% of class, 81% of its 31 rows; 0:cannot5.5% of class, 81% of its 62 rows; 4:thrilled5.4% of class, 88% of its 26 rows; 3:having5.4% of class, 73% of its 41 rows; 0:fix this now5.4% of class, 100% of its 49 rows; 0:now i5.4% of class, 74% of its 66 rows; 0:calling5.0% of class, 87% of its 53 rows; 3:having a4.8% of class, 90% of its 30 rows; 0:filing a4.8% of class, 83% of its 53 rows; 0:i cant4.7% of class, 77% of its 56 rows; 0:if this4.7% of class, 81% of its 53 rows; 4:so happy with4.5% of class, 83% of its 23 rows; 0:lose4.4% of class, 89% of its 45 rows - state.channel = phone transcript (1,846 rows): 4:
amazing27.9% of class, 86% of its 115 rows; 4:oh my27.0% of class, 96% of its 100 rows; 3:thanks so much22.6% of class, 73% of its 113 rows; 0:fix22.5% of class, 78% of its 176 rows; 4:gosh21.7% of class, 92% of its 84 rows; 4:thrilled21.4% of class, 84% of its 91 rows; 4:say thank you20.6% of class, 89% of its 82 rows; 0:cannot19.2% of class, 77% of its 151 rows; 4:perfectly18.6% of class, 73% of its 90 rows; 0:nobody18.4% of class, 82% of its 136 rows; 0:hello is18.0% of class, 80% of its 138 rows; 0:look i18.0% of class, 79% of its 140 rows; 4:best17.7% of class, 80% of its 79 rows; 4:to say thank you17.2% of class, 87% of its 70 rows; 0:anyone17.0% of class, 76% of its 137 rows; 4:oh my gosh16.9% of class, 98% of its 61 rows; 4:so so16.6% of class, 97% of its 61 rows; 0:fix this15.4% of class, 91% of its 103 rows; 4:just so15.2% of class, 75% of its 72 rows; 0:if this14.6% of class, 73% of its 122 rows; 0:is anyone14.6% of class, 95% of its 94 rows; 0:please i14.3% of class, 94% of its 93 rows; 0:listen13.3% of class, 76% of its 106 rows; 4:perfect13.2% of class, 89% of its 53 rows; 4:i am just12.4% of class, 80% of its 55 rows; 0:right now i12.1% of class, 73% of its 101 rows; 0:shaking12.1% of class, 96% of its 77 rows; 4:i wanted to say12.1% of class, 81% of its 53 rows; 0:today i12.0% of class, 83% of its 88 rows; 4:say thank you so11.5% of class, 91% of its 45 rows; 4:so happy11.5% of class, 79% of its 52 rows; 4:incredibly11.3% of class, 83% of its 48 rows; 4:thank you so so11.3% of class, 98% of its 41 rows; 4:the best11.3% of class, 80% of its 50 rows; 4:you so so much11.3% of class, 98% of its 41 rows; 0:to me11.1% of class, 76% of its 90 rows; 0:you have to11.1% of class, 97% of its 70 rows; 0:by five10.8% of class, 77% of its 86 rows; 0:i cannot10.5% of class, 77% of its 83 rows; 3:ever so10.5% of class, 84% of its 45 rows - state.channel = app review (1,430 rows): 0:
1 5 stars58.8% of class, 99% of its 315 rows; 4:5 5 stars51.5% of class, 90% of its 168 rows; 4:absolutely23.7% of class, 71% of its 98 rows; 0:fix22.0% of class, 77% of its 151 rows; 3:5 stars great16.1% of class, 76% of its 66 rows; 4:best15.3% of class, 88% of its 51 rows; 4:perfectly14.9% of class, 80% of its 55 rows; 4:amazing14.2% of class, 88% of its 48 rows; 0:fix this13.9% of class, 97% of its 76 rows; 4:5 stars absolutely13.6% of class, 82% of its 49 rows; 0:right now12.8% of class, 84% of its 81 rows; 3:5 stars lovely12.5% of class, 72% of its 54 rows; 0:immediately12.4% of class, 89% of its 74 rows; 4:so happy11.9% of class, 83% of its 42 rows; 3:4 5 stars great11.6% of class, 80% of its 45 rows; 3:could you11.6% of class, 73% of its 49 rows; 3:kindly11.6% of class, 82% of its 44 rows; 4:5 5 stars absolutely11.2% of class, 100% of its 33 rows; 0:this is10.7% of class, 81% of its 70 rows; 3:overall10.0% of class, 72% of its 43 rows; 4:thrilled9.2% of class, 93% of its 29 rows; 4:ever8.5% of class, 78% of its 32 rows; 0:filing8.5% of class, 80% of its 56 rows; 4:such a8.5% of class, 76% of its 33 rows; 3:4 5 stars lovely8.4% of class, 74% of its 35 rows; 0:money8.1% of class, 75% of its 57 rows; 0:oh7.7% of class, 93% of its 44 rows; 0:nobody7.5% of class, 83% of its 48 rows; 4:5 stars absolutely brilliant7.1% of class, 95% of its 22 rows; 0:this to6.8% of class, 95% of its 38 rows; 0:times6.8% of class, 80% of its 45 rows; 4:5 stars amazing6.4% of class, 90% of its 21 rows; 0:lawyer6.4% of class, 83% of its 41 rows; 0:unacceptable6.4% of class, 100% of its 34 rows; 3:thanks so much6.1% of class, 79% of its 24 rows; 4:with how6.1% of class, 82% of its 22 rows; 0:broken5.8% of class, 84% of its 37 rows; 0:if this5.8% of class, 79% of its 39 rows; 0:5 stars brilliant5.6% of class, 86% of its 35 rows; 0:never5.6% of class, 75% of its 40 rows - state.channel = social media reply (1,208 rows): 0:
fix19.9% of class, 82% of its 110 rows; 0:or18.8% of class, 80% of its 106 rows; 4:amazing17.1% of class, 90% of its 39 rows; 4:absolutely11.7% of class, 83% of its 29 rows; 4:happy11.7% of class, 83% of its 29 rows; 4:thank you so much11.7% of class, 80% of its 30 rows; 0:fix this11.3% of class, 94% of its 54 rows; 4:you guys11.2% of class, 82% of its 28 rows; 0:twice11.1% of class, 88% of its 57 rows; 3:loving10.8% of class, 82% of its 34 rows; 4:best10.7% of class, 96% of its 23 rows; 4:much for9.8% of class, 95% of its 21 rows; 4:you guys are9.8% of class, 100% of its 20 rows; 4:perfectly9.3% of class, 90% of its 21 rows; 3:loving the8.9% of class, 88% of its 26 rows; 0:filing8.8% of class, 83% of its 48 rows; 4:thrilled7.8% of class, 73% of its 22 rows; 3:love the6.9% of class, 75% of its 24 rows; 3:loved6.9% of class, 82% of its 22 rows; 0:charged6.9% of class, 89% of its 35 rows; 0:or i6.9% of class, 97% of its 32 rows; 0:this now6.9% of class, 97% of its 32 rows; 0:immediately6.6% of class, 97% of its 31 rows; 0:oh6.2% of class, 90% of its 31 rows; 0:fix this now5.8% of class, 100% of its 26 rows; 0:emailed twice5.5% of class, 89% of its 28 rows; 0:now or5.5% of class, 100% of its 25 rows; 0:oh brilliant5.1% of class, 100% of its 23 rows; 0:this is5.1% of class, 96% of its 24 rows; 0:lawyer4.9% of class, 76% of its 29 rows; 0:locked4.9% of class, 85% of its 26 rows; 0:or we4.4% of class, 100% of its 20 rows; 0:right now4.0% of class, 78% of its 23 rows; 0:court3.8% of class, 81% of its 21 rows; 0:filing a3.5% of class, 80% of its 20 rows
Shortcut models
Predicting the label class on test (1,814 rows). Chance 20.0%, majority class ('0') 36.5%; balanced chance 20.0%.
| model (logistic regression, trained on the train sample) | test accuracy | balanced accuracy (mean recall) |
|---|---|---|
| bag of words, whole state (words and word pairs) | 76.0% | 66.1% |
bag of words, main text only (message) |
77.2% | 68.3% |
| surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) | 44.5% | 31.4% |
| surface features of the main text only | 44.3% | 31.2% |
Strongest single surface features (logistic regression on one feature, balanced accuracy on test):
| feature | accuracy | balanced accuracy |
|---|---|---|
| count_, | 36.7% | 20.7% |
| count_; | 36.5% | 20.1% |
| digit_ratio | 36.5% | 20.1% |
| url | 36.5% | 20.1% |
| count_( | 36.5% | 20.1% |
| count_) | 36.5% | 20.1% |
| newlines | 36.5% | 20.0% |
| chars(log) | 36.5% | 20.0% |
| words(log) | 36.5% | 20.0% |
| upper_ratio | 36.5% | 20.0% |
Other state fields alone (predicting the label class on test from one field, without the main text):
| field | treated as | accuracy | balanced accuracy |
|---|---|---|---|
| company | text: bag of words / length+empty | 36.5% / 36.5% | 20.0% / 20.0% |
| channel | categorical, 6 values | 36.5% | 20.0% |
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
- Train rows identical in the whole prompt (state, options, instructions): 0.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 0 (0.0%) | |
| calibration | 0 (0.0%) | |
| test | 0 (0.0%) |
Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)
- Train: 0 near-duplicate pairs; 0 rows (0.0%) sit in 0 clusters; largest cluster 0; excess rows (cluster size − 1) 0 (0.0%).
- Clusters with more than one label class: 0 (0 rows).
- Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 0 (0.0%), test 0 (0.0%)
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 0 |
| dev | 0 | 0 |
| calibration | 0 | 0 |
| test | 0 | 0 |
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 0 (0.0%) | |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 1 (0.0%) | 1 0.1% |
| chat preamble ('Here is/are...', 'Sure!') | 0 (0.0%) | |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 0 (0.0%) | |
| encoding garbage (mojibake/replacement char) | 0 (0.0%) | |
| HTML tag | 511 (3.1%) | 0 2.7%, 1 2.0%, 2 2.6%, 3 3.0%, 4 4.7% |
| HTML entity | 0 (0.0%) | |
| base64-like run (40+ chars) | 1 (0.0%) | 1 0.1% |
| URL | 8 (0.0%) | 0 0.0%, 1 0.1%, 2 0.1%, 4 0.1% |
'As an AI' / refusal:
triage-c290-m140-sentiment(1): Subject: Concern regarding stonework at Hrubý Rohozec ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I hope this message finds you well, though I must confes…HTML tag:
triage-c008-m090-sentiment(4): Subject: Fwd: Content Request - Please add Ilo Ilo to your library! ⏎ ⏎ From: Sarah Lim sarah.lim82@gmail.com ⏎ Date: 14 November 2023 a… |triage-c205-m124-sentiment(3): Subject: Fwd: Výsledek zkoušky – BUS402 ⏎ ⏎ ---------- Přeposlaná zpráva ---------- ⏎ Od: Zkouškové oddělení zkousky@edulearn.co.za ⏎ Da… |triage-c016-m054-sentiment(0): Subject: Incorrect Invoice #4409281 - Kronoberg VA - Legal action deadline tomorrow ⏎ ⏎ I am writing to you again because I have run out o…base64-like run (40+ chars):
triage-c056-m162-sentiment(1): Subject: URGENT: Your store domain will expire!!! ⏎ ⏎ Dear Beloved Manager, ⏎ ⏎ I am writing to you today with a heavy heart because your…URL:
triage-c133-m096-sentiment(2): Subject: Strategic Partnership Proposal - Global Digital Infrastructure ⏎ ⏎ Dear Sir/Madam, ⏎ ⏎ I hope this email finds you well. My name… |triage-c146-m064-sentiment(1): Oggetto: R: Richiesta informazioni piano anticipato ⏎ ⏎ ---------- Messaggio inoltrato ---------- ⏎ Da: Assistenza Clienti <assistenza@ser… |triage-c091-m045-sentiment(4): Hi. I emailed last week about your amazing stadium. I just want to say the Allsvenskan final was incredible. Best day ever. Also, can you f…Possibly cut off: 5,505 of 11,135 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: 0 42.3%, 1 46.2%, 2 48.7%, 3 56.5%, 4 56.2%
triage-c057-m194-sentiment: …eciate everything you do for us sellers! ⏎ ⏎ Warmest regards, ⏎ Lim Boon Kiat ⏎ Owner, KB Electronics & Furniture ⏎ +65 9182 3347triage-c120-m050-sentiment: …fach im System verbleibt. ⏎ ⏎ Mit freundlichen Grüßen ⏎ Thomas Weber ⏎ Einkaufsleiter, Stahlwerk Dortmund GmbH ⏎ +49 231 555 0192triage-c150-m146-sentiment: …-ce que vous allez valider ma demande d'annulation de ce double paiemnt avant 17h??? ⏎ ⏎ Sophie Martin ⏎ Comptable ⏎ 0612345678
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- ×175: "---------- Forwarded message ---------" (0 64, 4 46, 3 31)
- ×166: "Any unauthorized review, use, disclosure, or distribution is strictly prohibited." (0 59, 4 36, 3 30)
- ×159: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and intended solely for the use of the individu…" (4 41, 0 35, 3 34)
- ×148: "----- Forwarded message -----" (0 76, 3 28, 4 28)
- ×132: "I hope this message finds you well." (3 100, 4 28, 2 3)
- ×119: "-----Original Message-----" (0 60, 2 19, 4 18)
- ×118: "If you have received this email in error, please notify the sender immediately and delete this email from your system." (0 41, 4 27, 3 22)
- ×116: "If you are not the intended recipient, please delete all copies and notify the sender immediately." (0 55, 3 26, 2 13)
- ×107: "If you have received this email in error, please notify the sender immediately and delete this message from your system." (4 30, 0 25, 3 24)
- ×107: "Topic: Billing and Invoices" (0 40, 4 27, 3 19)
- ×91: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and may also be privileged." (0 31, 3 26, 2 13)
- ×81: "CONFIDENTIALITY NOTICE: This email and any attachments are strictly confidential and intended solely for the addressee." (0 26, 4 17, 3 16)
- ×66: "I hope this email finds you well." (3 41, 4 16, 2 6)
- ×63: "----- End forwarded message -----" (0 36, 4 13, 3 12)
- ×60: "--- Forwarded message ---" (0 34, 4 9, 3 8)
6. Samples
20 random train rows per kind: triage-sentiment-samples.txt. Reading notes are in the findings above.
QA: triage-needs_human
Checked 2026-09-30 13:39 by adapters/qa/qa.py (READY file READY-triage, ).
Verdict: PASS WITH NOTES
Triage (READY-triage 2026-09-30 13:33, train sha256 3aba76c5…3b16c0; train 66,532 rows = 16,633 messages × 4 questions). This is the first full check; the earlier triage reports checked withdrawn files. Each question type has its own report: triage-route.md, triage-urgency.md, triage-sentiment.md and triage-needs_human.md. Cue rates come from triage_extra.py.
Earlier cues, now fixed (checked)
Each rate below is the share of that label's rows containing the cue, compared across labels:
| cue | before | now |
|---|---|---|
| "Subject: URGENT" | 22.6% of urgency-4 against 3.0% of the rest | 3.0–4.4% at every urgency level |
| "just wanted to" | 20.4% of urgency-0 | 6.0–6.3% at every level, and 5.9% / 6.2% for needs_human no / yes |
| "by tomorrow" | 14.9% of urgency-2 | 1.8–4.1% |
| "formal complaint" | 7.9% of person-needed against 0.1% | 2.8% / 3.8% |
| "Hi there" | 17% of sentiment-3 | 4.1–5.6% |
| ends with "!" | 39% of very positive | 22–29% at every sentiment level |
"manual", "real person" and "no reply needed" appear in 0 rows.
Surface-only models:
- sentiment: 31.4% balanced accuracy (was 47.8%; chance 20%);
- urgency: 33.4% (chance 20%);
- needs_human: 60.5% (chance 50%);
- route: 9.1% (chance 4.2%).
No single surface feature beats chance by more than 1.3 points.
Route options: the correct team is the longest option 10.6% of the time against 9.8% chance, positions are uniform, and the option picker is at chance.
Duplicates: there are no near duplicates in train or across splits, and no identical prompt carries two different answers.
The edited messages read naturally. In my sample, urgency-4 messages now contain "just wanted to" ("I just wanted to say I need my boxes from unit 12 today!"), low-urgency messages carry "Subject: Urgent" (a spam investment offer; a confirmation email), and very negative messages end with "!".
Notes (meaning; for the model card)
- The remaining surface signal is meaning:
- sentiment: unmatched ")" versus "(" counts, which are emoticons like :) and :( — and message length;
- urgency: fewer "?" in urgent messages (questions are rarely emergencies);
- needs_human: weak length and punctuation mixtures, with no single feature above chance.
- The standard phrase flags are meaning words:
- urgency 4: "right now", "fix this", "deadline", "tonight", "court", "lawyer", "breach";
- urgency 1: "no rush at all", the customer's own statement (kept by agreement);
- urgency 0: "confirm that", "say thank you";
- sentiment: "thrilled", "amazing", "warmest regards", "unacceptable", "nobody";
- needs_human no: "a quick note", "to let you know that", "confirm that the", "perfectly";
- route "other": "my thesis", "journalism student", "sponsoring" (student and sponsorship requests belong to no team).
- The label mix changed with the re-rating: urgency level 4 is now 34.7% of rows, and sentiment level 0 (very negative) is 35.6%. Dev and calibration are small: 1,368 and 1,208 rows, which is 342 and 302 messages (5 organisations each), so calibrate with care.
- About 51% of messages had small cue edits by qwen3.8-flash, recorded in
source.decorrelated_by. Score targets are the mean of three raters (the plan, qwen3.8-max and deepseek-v4-flash).
Automatic flags (for the reviewer to judge; not all are problems)
- 25 strong phrase flags (see list): review whether they are meaning or leakage
- 35 standard phrase flags (≥2% of a class, mostly that class): review
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 16,633 | 260 | train.jsonl |
| dev | 342 | 5 | dev.jsonl |
| calibration | 302 | 5 | calibration.jsonl |
| test | 1,814 | 30 | test.jsonl |
- Train sha256:
3aba76c53cd16a1c2822d1b5e4090d183cc77fd94a233898c53501f25c3b16c0(READY file gives no checksum) - Main text field (the text the phrase and length checks use):
state.message. - Label classes: False, True (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): needs_human.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| False | 40.4% | 41.2% | 40.1% | 42.4% | 6,719 |
| True | 59.6% | 58.8% | 59.9% | 57.6% | 9,914 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| needs_human | 100.0% | 100.0% | 100.0% | 100.0% | 16,633 |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | estimate: characters / 3 (upper bound for English) | 255 | 925 | 1385 | 0 |
| dev | estimate: characters / 3 (upper bound for English) | 273 | 927 | 1136 | 0 |
| calibration | estimate: characters / 3 (upper bound for English) | 262 | 1026 | 1203 | 0 |
| test | estimate: characters / 3 (upper bound for English) | 253 | 893 | 1215 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| channel, company, message | 16,633 (100.0%) |
Instructions (train)
- Canonical (the most common text) 69.7%, reworded 27.1% (70 distinct rewordings), none 3.2%. Target about 70 / 27 / 3.
- Canonical text: "Does this need a person to act on it now, rather than an automatic reply? Answer yes for complaints that could escalate, legal or safety issues, or requests an automatic system cannot resolve."
sourceinstruction tag: canonical 69.7%, none 3.2%, variant-52 0.5%, variant-23 0.5%, variant-29 0.5%, variant-11 0.5%
| class | canonical | none |
|---|---|---|
| False | 70.0% | 3.4% |
| True | 69.5% | 3.1% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 16,633 train rows; models are scored on the full test file (1,814 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | False | 6719 | 119 | 404 | 1352 | 596 |
| train | True | 9914 | 145 | 452 | 1418 | 649 |
| test | False | 770 | 108 | 403 | 1436 | 605 |
| test | True | 1044 | 143 | 439 | 1392 | 638 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| needs_human | 16633 | 433 | 628 | 145 | 0 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 59.6%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| human_reason | 7 | 99.8% | manual: True 100.0%; legal_safety: True 100.0%; routine: False 100.0%; acknowledge: False 100.0%; escalation: True 100.0%; no_reply: False 100.0% |
| situation | 17 | 87.4% | null: True 58.2%, False 41.8%; the writer reports or confirms something that needs no decision, such as that a delivery arrived, a payment was made, a problem sorted itself out, or a form was sent: False 100.0%; the writer is about to take it further and says what they will do next, for example cancel, go to a regulator or ombudsman, post a public review, contact the press, or make a formal compl… |
| tone | 25 | 75.0% | politely formal: False 61.8%, True 38.2%; cheerful: True 51.4%, False 48.6%; polite and friendly: False 57.6%, True 42.4%; warm: False 53.5%, True 46.5%; disappointed: True 82.1%, False 17.9%; appreciative: False 54.5%, True 45.5% |
| other_kind | 9 | 65.8% | null: True 63.6%, False 36.4%; a job application or question about working there: False 97.5%, True 2.5%; a message clearly meant for a different organisation: False 98.3%, True 1.7%; a sales pitch from another business offering its services: False 97.4%, True 2.6%; a request to sponsor or donate to a local event: False 92.6%, True 7.4%; a journalist or student asking for an interview or informat… |
| channel | 6 | 59.6% | email: True 60.8%, False 39.2%; web form: True 61.6%, False 38.4%; chat: True 58.0%, False 42.0%; phone transcript: True 59.3%, False 40.7%; app review: True 54.9%, False 45.1%; social media reply: True 57.5%, False 42.5% |
| language | 11 | 59.6% | English: True 60.0%, False 40.0%; Danish: True 53.0%, False 47.0%; Czech: True 55.4%, False 44.6%; French: True 56.5%, False 43.5%; Swedish: True 50.3%, False 49.7%; Dutch: True 57.0%, False 43.0% |
| length_target | 17 | 59.6% | 60: True 61.0%, False 39.0%; 120: True 61.7%, False 38.3%; 25: True 58.5%, False 41.5%; 220: True 62.8%, False 37.2%; 50: True 58.3%, False 41.7%; 10: True 54.3%, False 45.7% |
| prompt_version | 3 | 59.6% | 3: True 60.2%, False 39.8%; 2: True 57.8%, False 42.2%; 1: True 74.1%, False 25.9% |
| decorrelated_by | 2 | 59.6% | qwen3.8-flash: True 57.6%, False 42.4%; null: True 61.7%, False 38.3% |
| regenerated | 2 | 59.6% | false: True 58.0%, False 42.0%; true: True 63.9%, False 36.1% |
Formatting by label class (main text, share of rows)
| feature | False | True | |
|---|---|---|---|
| ends with ? | 3.8% | 4.8% | |
| ends with . | 36.2% | 39.0% | |
| ends with ! | 13.7% | 13.1% | |
| no end punctuation | 45.9% | 43.0% | |
| starts lowercase | 6.7% | 6.3% | |
| all lowercase | 5.3% | 3.6% | |
| has a digit | 83.9% | 90.8% | |
| has newline | 76.2% | 77.5% | |
| has quotes | 3.2% | 8.1% | |
| has markup (HTML/markdown) | 3.6% | 2.9% | |
| has URL | 0.1% | 0.0% | |
| non-ASCII | 46.7% | 54.3% | |
| non-Latin script | 0.0% | 0.0% | |
| emoji | 5.9% | 5.1% | |
| ALL-CAPS word (4+) | 16.1% | 16.2% | |
| contains ' - ' or — | 14.2% | 19.8% |
Same, by row kind
| feature | needs_human |
|---|---|
| ends with ? | 4% |
| ends with . | 38% |
| ends with ! | 13% |
| no end punctuation | 44% |
| starts lowercase | 6% |
| all lowercase | 4% |
| has a digit | 88% |
| has newline | 77% |
| has quotes | 6% |
| has markup (HTML/markdown) | 3% |
| has URL | 0% |
| non-ASCII | 51% |
| non-Latin script | 0% |
| emoji | 5% |
| ALL-CAPS word (4+) | 16% |
| contains ' - ' or — | 18% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, False: thank z=30 36.6% vs 15.2%; thanks z=24 20.4% vs 7.9%; much z=23 27.5% vs 13.3%; perfectly z=22 9.3% vs 0.9%; so z=20 46.0% vs 30.0%; everything z=19 19.0% vs 9.0%; happy z=19 10.9% vs 3.3%; quick z=19 8.9% vs 2.3%; wanted z=18 16.3% vs 7.5%; all z=18 25.8% vs 15.0%; best z=18 10.3% vs 3.4%; anyway z=18 10.9% vs 3.9%; just z=18 36.7% vs 24.2%; rush z=18 7.5% vs 1.7%; such z=17 9.6% vs 3.2%; say z=17 13.9% vs 6.3%; went z=17 7.5% vs 2.1%; question z=16 8.6% vs 2.9%; fine z=16 7.0% vs 1.8%; regards z=16 16.6% vs 9.1%
Words, True: today z=30 39.3% vs 10.2%; before z=26 29.0% vs 7.8%; tomorrow z=22 18.7% vs 3.9%; fix z=22 15.3% vs 1.6%; if z=22 41.4% vs 18.8%; deadline z=20 13.9% vs 1.0%; cannot z=20 14.2% vs 2.6%; someone z=20 15.0% vs 3.1%; right z=19 19.4% vs 6.6%; will z=19 22.9% vs 8.8%; immediately z=18 16.5% vs 5.0%; legal z=17 9.9% vs 1.1%; look z=16 12.8% vs 4.0%; by z=16 27.5% vs 13.6%; must z=15 10.4% vs 2.9%; contract z=15 10.3% vs 3.0%; tonight z=14 7.0% vs 0.5%; need z=14 21.5% vs 10.4%; filing z=14 7.2% vs 1.3%; please z=14 32.5% vs 18.7%
2–4-word phrases, False: thank you z=28 35.5% vs 14.9%; to say z=23 11.4% vs 2.1%; i wanted z=19 7.8% vs 1.0%; i wanted to z=19 7.7% vs 1.0%; so much z=18 20.8% vs 10.7%; to confirm z=17 6.3% vs 0.8%; wanted to z=17 16.0% vs 7.5%; thank you for z=17 12.4% vs 5.0%; you for z=17 12.8% vs 5.4%; no rush z=17 5.9% vs 0.9%; writing to z=16 9.9% vs 3.7%; you know z=16 5.4% vs 0.7%; am writing to z=16 9.3% vs 3.4%; i am writing to z=16 9.3% vs 3.4%; everything is z=15 5.2% vs 0.9%; such a z=15 6.6% vs 1.9%; again for z=15 5.6% vs 1.3%; that the z=15 8.8% vs 3.4%; a quick z=15 4.4% vs 0.5%; at all z=15 7.1% vs 2.4%
2–4-word phrases, True: right now z=22 15.8% vs 2.4%; i need z=16 11.3% vs 3.3%; this is z=16 15.0% vs 5.6%; if this z=16 8.0% vs 1.1%; before the z=15 7.6% vs 1.4%; fix this z=15 9.0% vs 0.3%; i cannot z=14 7.3% vs 1.7%; will be z=14 7.7% vs 2.0%; today i z=14 6.3% vs 1.0%; is not z=13 7.2% vs 2.1%; look into z=12 4.9% vs 0.7%; tomorrow morning z=12 5.0% vs 0.2%; i will z=12 8.6% vs 3.3%; if i z=11 6.0% vs 1.9%; this to z=11 3.8% vs 0.4%; look at z=11 3.9% vs 0.7%; filing a z=11 3.8% vs 0.7%; if we z=10 3.7% vs 0.7%; within the z=10 3.4% vs 0.6%; 17 00 z=10 4.0% vs 0.1%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- True:
tomorrow18.7% vs 3.9% - True:
right now15.8% vs 2.4% - True:
fix15.3% vs 1.6% - True:
someone15.0% vs 3.1% - True:
cannot14.2% vs 2.6% - True:
deadline13.9% vs 1.0% - False:
to say11.4% vs 2.1% - True:
legal9.9% vs 1.1% - False:
perfectly9.3% vs 0.9% - True:
fix this9.0% vs 0.3% - True:
if this8.0% vs 1.1% - False:
i wanted7.8% vs 1.0% - False:
i wanted to7.7% vs 1.0% - True:
before the7.6% vs 1.4% - False:
rush7.5% vs 1.7% - True:
i cannot7.3% vs 1.7% - True:
filing7.2% vs 1.3% - True:
tonight7.0% vs 0.5% - False:
to confirm6.3% vs 0.8% - True:
today i6.3% vs 1.0% - False:
no rush5.9% vs 0.9% - False:
again for5.6% vs 1.3% - False:
you know5.4% vs 0.7% - False:
everything is5.2% vs 0.9% - True:
tomorrow morning5.0% vs 0.2%
Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):
- all rows (16,633 rows): False:
perfectly9.3% of class, 87% of its 717 rows; False:i wanted to7.7% of class, 85% of its 614 rows; False:to confirm6.3% of class, 83% of its 509 rows; False:no rush5.9% of class, 81% of its 492 rows; False:you know5.4% of class, 83% of its 434 rows; False:confirm that5.1% of class, 95% of its 363 rows; False:to confirm that4.6% of class, 97% of its 321 rows; False:a quick4.4% of class, 85% of its 346 rows; False:bye4.3% of class, 86% of its 338 rows; False:to let4.0% of class, 86% of its 313 rows; False:perfect3.6% of class, 85% of its 287 rows; False:rush at all3.5% of class, 90% of its 264 rows; False:to let you know3.5% of class, 93% of its 254 rows; False:successfully3.2% of class, 88% of its 246 rows; False:no rush at all3.1% of class, 90% of its 234 rows; False:know that3.0% of class, 89% of its 228 rows; False:gratitude3.0% of class, 83% of its 242 rows; False:went through3.0% of class, 84% of its 239 rows; False:for making2.8% of class, 87% of its 215 rows; False:say thank you2.7% of class, 95% of its 191 rows; False:thank you for the2.7% of class, 88% of its 205 rows; False:to share2.7% of class, 85% of its 214 rows; False:confirm that the2.7% of class, 97% of its 185 rows; False:let you know that2.7% of class, 95% of its 188 rows; False:a quick note2.5% of class, 96% of its 174 rows; False:to confirm that the2.4% of class, 98% of its 163 rows; False:to say thank you2.4% of class, 95% of its 167 rows; False:am writing to confirm2.3% of class, 99% of its 156 rows; False:note to2.3% of class, 98% of its 158 rows; False:keep up the2.2% of class, 88% of its 170 rows; False:writing to confirm that2.2% of class, 99% of its 148 rows; False:and everything2.1% of class, 83% of its 169 rows; False:5 5 stars2.1% of class, 82% of its 168 rows; False:subject thank you2.0% of class, 96% of its 142 rows; False:to drop2.0% of class, 94% of its 144 rows - state.channel = email (6,611 rows): False:
perfectly12.1% of class, 84% of its 374 rows; False:to confirm11.0% of class, 86% of its 329 rows; False:to say10.1% of class, 79% of its 334 rows; False:confirm that9.2% of class, 94% of its 252 rows; False:i wanted to9.1% of class, 79% of its 297 rows; False:to confirm that8.5% of class, 97% of its 226 rows; False:you know7.4% of class, 91% of its 211 rows; False:gratitude6.6% of class, 85% of its 202 rows; False:to let6.5% of class, 86% of its 195 rows; False:to let you know6.1% of class, 93% of its 168 rows; False:successfully5.8% of class, 85% of its 176 rows; False:a quick5.7% of class, 86% of its 171 rows; False:subject thank you5.2% of class, 96% of its 142 rows; False:know that5.2% of class, 87% of its 154 rows; False:to share5.0% of class, 82% of its 158 rows; False:thank you for the5.0% of class, 89% of its 145 rows; False:let you know that4.7% of class, 95% of its 129 rows; False:quick note4.6% of class, 94% of its 125 rows; False:am writing to confirm4.5% of class, 99% of its 118 rows; False:for making4.4% of class, 86% of its 133 rows; False:to confirm that the4.4% of class, 98% of its 116 rows; False:writing to confirm that4.3% of class, 99% of its 112 rows; False:note to4.1% of class, 98% of its 108 rows; False:a quick note4.0% of class, 96% of its 109 rows; False:beautifully3.5% of class, 80% of its 115 rows; False:rush at all3.5% of class, 88% of its 104 rows; False:subject thank you for3.4% of class, 97% of its 92 rows; False:a quick note to3.4% of class, 98% of its 90 rows; False:no rush at all3.2% of class, 89% of its 92 rows; False:to drop3.1% of class, 94% of its 85 rows; False:my end3.0% of class, 84% of its 94 rows; False:and everything2.9% of class, 88% of its 84 rows; False:gratitude for2.9% of class, 90% of its 82 rows; False:was so2.9% of class, 79% of its 94 rows; False:on my end2.8% of class, 88% of its 83 rows; False:our end2.8% of class, 84% of its 86 rows; False:deepest2.7% of class, 83% of its 86 rows; False:sincere2.7% of class, 82% of its 85 rows; False:keep up the2.7% of class, 88% of its 78 rows; False:everything went2.6% of class, 92% of its 74 rows - state.channel = web form (3,009 rows): False:
to say11.0% of class, 82% of its 155 rows; False:perfectly10.1% of class, 91% of its 129 rows; False:i am writing to9.9% of class, 79% of its 145 rows; False:i wanted to9.4% of class, 85% of its 128 rows; False:such a8.1% of class, 78% of its 121 rows; False:quick7.5% of class, 78% of its 112 rows; False:rush7.3% of class, 82% of its 102 rows; False:again for6.6% of class, 86% of its 88 rows; False:whenever6.6% of class, 79% of its 96 rows; False:to confirm6.2% of class, 78% of its 91 rows; False:no rush5.9% of class, 85% of its 80 rows; False:everything is5.7% of class, 79% of its 84 rows; False:you know5.2% of class, 91% of its 66 rows; False:to let5.0% of class, 85% of its 68 rows; False:a quick4.9% of class, 88% of its 65 rows; False:finally4.9% of class, 77% of its 74 rows; False:smoothly4.6% of class, 80% of its 66 rows; False:to confirm that4.5% of class, 98% of its 53 rows; False:to let you know4.4% of class, 91% of its 56 rows; False:share4.2% of class, 79% of its 61 rows; False:successfully4.1% of class, 98% of its 48 rows; False:drop4.0% of class, 88% of its 52 rows; False:a quick note3.7% of class, 98% of its 44 rows; False:whenever you3.7% of class, 84% of its 51 rows; False:no rush at all3.6% of class, 91% of its 46 rows; False:you again3.6% of class, 84% of its 50 rows; False:let you know that3.5% of class, 95% of its 42 rows; False:team i am3.5% of class, 77% of its 52 rows; False:note to3.4% of class, 100% of its 39 rows; False:to drop3.4% of class, 95% of its 41 rows; False:went through3.3% of class, 83% of its 46 rows; False:completed3.2% of class, 82% of its 45 rows; False:keep up the3.2% of class, 90% of its 41 rows; False:thank you again3.2% of class, 84% of its 44 rows; False:the standard3.2% of class, 77% of its 48 rows; False:warmest regards3.2% of class, 82% of its 45 rows; False:thanks again3.1% of class, 88% of its 41 rows; False:worked3.1% of class, 88% of its 41 rows; False:for making2.9% of class, 89% of its 38 rows; False:beautifully2.9% of class, 82% of its 40 rows - state.channel = chat (2,529 rows): False:
link6.8% of class, 85% of its 85 rows; False:no rush6.4% of class, 89% of its 76 rows; False:perfectly4.2% of class, 94% of its 48 rows; False:thank you for3.9% of class, 95% of its 43 rows; False:i wanted to3.4% of class, 95% of its 38 rows; False:whenever3.4% of class, 90% of its 40 rows; False:hi just3.1% of class, 94% of its 35 rows; False:no rush at all2.9% of class, 94% of its 33 rows; False:is there a2.8% of class, 91% of its 33 rows; False:send me2.6% of class, 93% of its 30 rows; False:wondering2.4% of class, 87% of its 30 rows; False:copy2.3% of class, 92% of its 26 rows; False:just confirming2.2% of class, 100% of its 23 rows - state.channel = phone transcript (1,846 rows): False:
bye37.6% of class, 86% of its 329 rows; False:message19.9% of class, 85% of its 177 rows; False:i wanted to16.5% of class, 91% of its 136 rows; False:anyway i16.2% of class, 82% of its 149 rows; False:wanted to say15.0% of class, 82% of its 138 rows; False:just calling12.9% of class, 86% of its 113 rows; False:to leave10.8% of class, 93% of its 87 rows; False:perfectly10.5% of class, 88% of its 90 rows; False:say thank you10.4% of class, 95% of its 82 rows; False:leave a10.0% of class, 90% of its 83 rows; False:no rush9.8% of class, 86% of its 86 rows; False:calling to9.4% of class, 92% of its 77 rows; False:goodbye9.0% of class, 91% of its 75 rows; False:wanted to leave9.0% of class, 100% of its 68 rows; False:to say thank you8.8% of class, 94% of its 70 rows; False:a quick7.7% of class, 87% of its 67 rows; False:to leave a7.2% of class, 96% of its 56 rows; False:message to6.9% of class, 100% of its 52 rows; False:perfect6.9% of class, 98% of its 53 rows; False:s all6.9% of class, 96% of its 54 rows; False:this message6.9% of class, 88% of its 59 rows; False:thanks bye6.8% of class, 91% of its 56 rows; False:a message6.6% of class, 83% of its 60 rows; False:i wanted to say6.6% of class, 94% of its 53 rows; False:just um6.5% of class, 84% of its 58 rows; False:rush at all6.2% of class, 89% of its 53 rows; False:i m just calling6.1% of class, 88% of its 52 rows; False:wanted to leave a6.1% of class, 100% of its 46 rows; False:i am just6.0% of class, 82% of its 55 rows; False:let you5.9% of class, 98% of its 45 rows; False:yeah that5.7% of class, 86% of its 50 rows; False:is all5.6% of class, 88% of its 48 rows; False:say thank you so5.6% of class, 93% of its 45 rows; False:just calling to5.3% of class, 98% of its 41 rows; False:leave this5.3% of class, 93% of its 43 rows; False:no rush at all5.3% of class, 89% of its 45 rows; False:message to say5.2% of class, 100% of its 39 rows; False:just thought5.1% of class, 86% of its 44 rows; False:just uh5.1% of class, 83% of its 46 rows; False:thank you very much5.1% of class, 88% of its 43 rows - state.channel = app review (1,430 rows): False:
went through3.4% of class, 96% of its 23 rows; False:everything is3.3% of class, 95% of its 22 rows; False:perfect3.1% of class, 95% of its 21 rows; False:whenever3.1% of class, 100% of its 20 rows - state.channel = social media reply (1,208 rows): False:
do i7.8% of class, 85% of its 47 rows; False:happy5.1% of class, 90% of its 29 rows; False:best4.1% of class, 91% of its 23 rows; False:perfectly4.1% of class, 100% of its 21 rows; False:how do i3.9% of class, 87% of its 23 rows; False:much for3.9% of class, 95% of its 21 rows
Shortcut models
Predicting the label class on test (1,814 rows). Chance 50.0%, majority class ('True') 57.6%; balanced chance 50.0%.
| model (logistic regression, trained on the train sample) | test accuracy | balanced accuracy (mean recall) |
|---|---|---|
| bag of words, whole state (words and word pairs) | 91.5% | 90.6% |
bag of words, main text only (message) |
93.6% | 93.1% |
| surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) | 64.6% | 60.5% |
| surface features of the main text only | 64.1% | 60.0% |
Strongest single surface features (logistic regression on one feature, balanced accuracy on test):
| feature | accuracy | balanced accuracy |
|---|---|---|
| chars(log) | 57.6% | 50.1% |
| words(log) | 57.6% | 50.1% |
| url | 57.6% | 50.1% |
| markdown | 57.6% | 50.1% |
| count_* | 57.6% | 50.1% |
| nonlatin | 57.6% | 50.0% |
| upper_ratio | 57.6% | 50.0% |
| digit_ratio | 57.6% | 50.0% |
| nonascii_ratio | 57.6% | 50.0% |
| emoji | 57.6% | 50.0% |
Other state fields alone (predicting the label class on test from one field, without the main text):
| field | treated as | accuracy | balanced accuracy |
|---|---|---|---|
| company | text: bag of words / length+empty | 57.6% / 57.6% | 50.0% / 50.0% |
| channel | categorical, 6 values | 57.6% | 50.0% |
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 0 (0.0%); groups: 0; largest group 1.
- Train rows identical in the whole prompt (state, options, instructions): 0.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 0 (0.0%) | |
| calibration | 0 (0.0%) | |
| test | 0 (0.0%) |
Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)
- Train: 0 near-duplicate pairs; 0 rows (0.0%) sit in 0 clusters; largest cluster 0; excess rows (cluster size − 1) 0 (0.0%).
- Clusters with more than one label class: 0 (0 rows).
- Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 0 (0.0%), test 0 (0.0%)
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 0 |
| dev | 0 | 0 |
| calibration | 0 | 0 |
| test | 0 | 0 |
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 0 (0.0%) | |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 1 (0.0%) | True 0.0% |
| chat preamble ('Here is/are...', 'Sure!') | 0 (0.0%) | |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 0 (0.0%) | |
| encoding garbage (mojibake/replacement char) | 0 (0.0%) | |
| HTML tag | 511 (3.1%) | False 3.4%, True 2.8% |
| HTML entity | 0 (0.0%) | |
| base64-like run (40+ chars) | 1 (0.0%) | False 0.0% |
| URL | 8 (0.0%) | False 0.1%, True 0.0% |
'As an AI' / refusal:
triage-c290-m140-needs_human(True): Subject: Concern regarding stonework at Hrubý Rohozec ⏎ ⏎ Dear Sir or Madam, ⏎ ⏎ I hope this message finds you well, though I must confes…HTML tag:
triage-c089-m228-needs_human(False): Subject: Fwd: Your consultation is booked ⏎ ⏎ ---------- Forwarded message --------- ⏎ From: SecureHome India sales@securehome.in ⏎ Date… |triage-c221-m174-needs_human(True): Subject: Fwd: Police notification - Accident on A1 ⏎ ⏎ From: PSP Transit Authority transito@psp.pt ⏎ Date: 14 May 2024 08:15 ⏎ To: Maria… |triage-c012-m239-needs_human(True): Subject: Fwd: RE: Booking REF #PK-88291 ⏎ ⏎ ---------- Forwarded message --------- ⏎ From: ParkFr noreply@parkfr.com ⏎ Date: Mon, 14 Oct…base64-like run (40+ chars):
triage-c056-m162-needs_human(False): Subject: URGENT: Your store domain will expire!!! ⏎ ⏎ Dear Beloved Manager, ⏎ ⏎ I am writing to you today with a heavy heart because your…URL:
triage-c238-m151-needs_human(False): Subject: Meine Bougainvillea blüht wunderschön. ⏎ ⏎ Liebes Team, ⏎ ⏎ ich wollte Ihnen nur mitteilen, dass Ihre Tipps zum Zurückschneiden … |triage-c133-m096-needs_human(False): Subject: Strategic Partnership Proposal - Global Digital Infrastructure ⏎ ⏎ Dear Sir/Madam, ⏎ ⏎ I hope this email finds you well. My name… |triage-c074-m101-needs_human(False): Subject: We offer printing services for your charity ⏎ ⏎ Hello friends, ⏎ ⏎ My name is Laszlo and I am working for GreenPrint Solutions K…Possibly cut off: 5,505 of 11,135 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: False 53.3%, True 47.0%
triage-c065-m038-needs_human: …to where everything is right now? I am very worried about this situation. ⏎ ⏎ Thank you, ⏎ Lukas Brunner ⏎ Account: 0881-442-11triage-c032-m188-needs_human: …before Friday, as I need to finalize our winter menu suppliers by then. ⏎ ⏎ Thomas Varga ⏎ Owner & Head Buyer ⏎ +36 30 555 0192triage-c201-m195-needs_human: … discharge tomorrow. ⏎ ⏎ Thank you so much for your wonderful collaboration. ⏎ ⏎ Warm regards, ⏎ Marco Ferretti ⏎ Practice Mana…
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- ×175: "---------- Forwarded message ---------" (True 108, False 67)
- ×166: "Any unauthorized review, use, disclosure, or distribution is strictly prohibited." (True 100, False 66)
- ×159: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and intended solely for the use of the individu…" (True 82, False 77)
- ×148: "----- Forwarded message -----" (True 106, False 42)
- ×132: "I hope this message finds you well." (True 84, False 48)
- ×119: "-----Original Message-----" (True 81, False 38)
- ×118: "If you have received this email in error, please notify the sender immediately and delete this email from your system." (True 68, False 50)
- ×116: "If you are not the intended recipient, please delete all copies and notify the sender immediately." (True 87, False 29)
- ×107: "Topic: Billing and Invoices" (True 62, False 45)
- ×107: "If you have received this email in error, please notify the sender immediately and delete this message from your system." (True 59, False 48)
- ×91: "CONFIDENTIALITY NOTICE: This email and any attachments are confidential and may also be privileged." (True 62, False 29)
- ×81: "CONFIDENTIALITY NOTICE: This email and any attachments are strictly confidential and intended solely for the addressee." (True 47, False 34)
- ×66: "I hope this email finds you well." (False 35, True 31)
- ×63: "----- End forwarded message -----" (True 46, False 17)
- ×60: "--- Forwarded message ---" (True 45, False 15)
6. Samples
20 random train rows per kind: triage-needs_human-samples.txt. Reading notes are in the findings above.
