QA report: spam
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.
QA: spam-isspam
Checked 2026-09-30 12:43 by adapters/qa/qa.py (READY file ready-spam-v2.json, v2 (QA fixes 2026-09-30)).
Verdict: PASS WITH NOTES
Spam v3 (READY line 2026-09-30 12:3x: train 27,913; train sha256 5308819c…30c0). The v1 notes are in notes-spam-v1.md. This report covers two question types, each checked separately: spam-type.md (the 3-way choice question) and spam-isspam.md (the yes/no question). The same verdict applies to both.
v1 findings, now fixed (checked with spam_extra.py and qa.py)
- Recipient traces: "jose" or monkey.org now appears in 0.7% of spam emails and 1.4% of legitimate ones (v1: 24.5% of phishing).
- Era: years are shifted per email, and the year buckets are similar for both labels.
- Reply structure is gone for every label: no "Subject: Re:", no quoted lines, and no "On … wrote:" headers.
- Mailing-list traces appear in 0.8% of spam and 1.4% of legitimate emails.
- 4,630 modern legitimate emails were added.
- Surface-only model: 54.3% balanced accuracy on the yes/no question (chance 50%; v1 81.8%) and 50.3% on the 3-way question (chance 33%; v1 69.2%). On email alone, the agent measured 49.9%.
- Near duplicates across splits: 0% of held-out rows (v1: 5–7%).
- The standard phrase flags are phishing vocabulary ("verify your", "your account", "do not reply to this", "paypal account"), which is the meaning of the label.
Notes (for the model card and v4)
The teacher-written legitimate emails have their own style. 2,682 train rows, a third of legitimate emails, were written by qwen3.8-max. Compared with corpus legitimate mail and with spam or phishing:
- they are never all lowercase (against 59% and 41%);
- they always have line breaks (against 42% and 58%);
- 51% have a sign-off line (against 2% and 17%);
- 64% contain a URL (against 17% and 30%).
This is not a label shortcut by the rule: only 67% of signed emails are legitimate, about the base rate. But the model may link "tidy, signed, modern email" with legitimate, while real modern phishing is also tidy. v4: have the teacher write modern phishing emails in the same style, or strip sign-offs evenly.
SMS style is real. URLs appear in 16.7% of spam SMS against 0.1% of legitimate SMS, and "!" in 42% against 12%. Years appear in 5.6% of spam SMS against 0% of legitimate SMS: promotional dates such as "draw 2015" or "valid till 2019". This is real SMS-spam content (accepted by main).
About 16% of legitimate emails are teacher-written (qwen3.8-max, hosted). This is recorded per row in
source.teacher_model.The label shares differ a little by split: "not spam" is 69.7% of train and test, against 72–76% of dev and calibration.
Automatic flags (for the reviewer to judge; not all are problems)
- lclass 'False' share varies across splits by more than 5 points: train 69.7%, dev 76.1%, calibration 72.7%, test 69.7%
- lclass 'True' share varies across splits by more than 5 points: train 30.3%, dev 23.9%, calibration 27.3%, test 30.3%
- formatting 'ALL-CAPS word (4+)' differs by class: True 31%, False 14%
- 40 strong phrase flags (see list): review whether they are meaning or leakage
- 60 standard phrase flags (≥2% of a class, mostly that class): review
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 16,893 | 16887 | train.jsonl |
| dev | 1,178 | 1178 | dev.jsonl |
| calibration | 1,145 | 1143 | calibration.jsonl |
| test | 2,184 | 2183 | test.jsonl |
- Train sha256:
5308819cbc0e7ebecce8473dd29b927c9da64cb04146dc631b575ea7923f30c0(READY file gives no checksum) - Main text field (the text the phrase and length checks use):
state.message. - Label classes: False, True (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): is_spam.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| False | 69.7% | 76.1% | 72.7% | 69.7% | 11,766 |
| True | 30.3% | 23.9% | 27.3% | 30.3% | 5,127 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| is_spam | 100.0% | 100.0% | 100.0% | 100.0% | 16,893 |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | estimate: characters / 3 (upper bound for English) | 204 | 2056 | 2059 | 0 |
| dev | estimate: characters / 3 (upper bound for English) | 141 | 2056 | 2058 | 0 |
| calibration | estimate: characters / 3 (upper bound for English) | 132 | 1724 | 2058 | 0 |
| test | estimate: characters / 3 (upper bound for English) | 206 | 2057 | 2059 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| channel, message | 16,893 (100.0%) |
Instructions (train)
- Canonical (the most common text) 69.8%, reworded 27.5% (20 distinct rewordings), none 2.7%. Target about 70 / 27 / 3.
- Canonical text: "Is this message spam or phishing? Answer yes if it is unwanted bulk or advertising, or a scam trying to get personal details or money; answer no if it is a normal message."
sourceinstruction tag: canonical 69.8%, none 2.7%, variant-8 1.6%, variant-5 1.5%, variant-10 1.5%, variant-19 1.5%
| class | canonical | none |
|---|---|---|
| False | 69.9% | 2.6% |
| True | 69.4% | 2.9% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 16,893 train rows; models are scored on the full test file (2,184 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | False | 11766 | 39 | 394 | 1913 | 795 |
| train | True | 5127 | 139 | 590 | 2360 | 1020 |
| test | False | 1522 | 36 | 381 | 1942 | 826 |
| test | True | 662 | 140 | 523 | 2206 | 967 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| is_spam | 16893 | 448 | 863 | 6 | 0 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 69.7%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| dataset | 7 | 81.9% | zefang_phishing_email: False 65.5%, True 34.5%; sms_phishing: False 82.6%, True 17.4%; generated_legit_email: False 100.0%; nazario: True 100.0%; spamassassin: False 85.1%, True 14.9%; sms_spam: False 69.9%, True 30.1% |
| url | 7 | 81.9% | https://huggingface.co/datasets/zefang-liu/phishing-email-dataset: False 65.5%, True 34.5%; https://data.mendeley.com/datasets/f45bkkt8pr/1: False 82.6%, True 17.4%; https://dashscope-intl.aliyuncs.com (Alibaba Cloud a hosted service API): False 100.0%; https://monkey.org/~jose/phishing/: True 100.0%; https://spamassassin.apache.org/old/publiccorpus/: False 85.1%, True 14.9%; https://archive.ics.uci.edu… |
| license | 7 | 81.9% | LICENCE UNCLEAR — owner to review before release. Tagged LGPL-3.0 by the uploader; the Kaggle original does not document where its emails come from (they look like Enron and other public email corpora).: False 65.5%, True 34.5%; CC BY 4.0 (Mendeley Data): False 82.6%, True 17.4%; Generated for this project by an approved teacher (COMMON.md 'The teacher', owner decision of 2026-09-29): hosted, clo… |
| revision | 5 | 73.2% | 34085a032c123ca237f314a01a67909cdea35e34: False 65.5%, True 34.5%; version 1: False 82.6%, True 17.4%; null: True 56.6%, False 43.4%; writer qwen3.8-max, blind check qwen3.8-flash (model recorded per row): False 100.0%; 55caf9af01ae498d68ff44886ff69a4b165eb54b: True 100.0% |
| copies | 19 | 70.0% | 2: False 73.8%, True 26.2%; 1: False 65.6%, True 34.4%; 3: False 63.8%, True 36.2%; 4: True 50.0%, False 50.0%; 5: True 73.3%, False 26.7%; 7: True 81.8%, False 18.2% |
| target_kind | 2 | 69.7% | hard: False 69.6%, True 30.4%; soft: False 100.0% |
| instructions | 22 | 69.7% | canonical: False 69.8%, True 30.2%; none: False 67.2%, True 32.8%; variant-8: False 67.3%, True 32.7%; variant-5: False 69.7%, True 30.3%; variant-10: False 74.5%, True 25.5%; variant-19: False 72.6%, True 27.4% |
| channel | 2 | 69.7% | email: False 65.2%, True 34.8%; null: False 82.2%, True 17.8% |
| licence_status | 3 | 69.7% | LICENCE UNCLEAR - owner to review before release: False 55.7%, True 44.3%; null: False 82.2%, True 17.8%; generated by an approved hosted-Qwen teacher: False 100.0% |
| teacher_model | 2 | 69.7% | null: False 63.9%, True 36.1%; qwen3.8-max: False 100.0% |
| check_model | 2 | 69.7% | null: False 63.9%, True 36.1%; qwen3.8-flash: False 100.0% |
| check_verdict | 2 | 69.7% | null: False 63.9%, True 36.1%; legitimate: False 100.0% |
| generated_category | 33 | 69.7% | null: False 63.9%, True 36.1%; newsletter from a company the reader subscribed to: False 100.0%; work email thread between colleagues about a project: False 100.0%; event ticket confirmation: False 100.0%; community or neighbourhood group announcement: False 100.0%; IT department notice about planned maintenance: False 100.0% |
Formatting by label class (main text, share of rows)
| feature | False | True | |
|---|---|---|---|
| ends with ? | 4.8% | 0.5% | |
| ends with . | 26.0% | 21.4% | |
| ends with ! | 3.4% | 3.1% | |
| no end punctuation | 62.9% | 72.8% | |
| starts lowercase | 30.4% | 36.5% | |
| all lowercase | 27.6% | 34.4% | |
| has a digit | 64.2% | 89.4% | |
| has newline | 42.3% | 49.1% | |
| has quotes | 13.6% | 17.1% | |
| has markup (HTML/markdown) | 6.2% | 8.5% | |
| has URL | 23.1% | 29.6% | |
| non-ASCII | 10.0% | 21.2% | |
| non-Latin script | 0.0% | 0.0% | |
| emoji | 0.0% | 0.0% | |
| ALL-CAPS word (4+) | 13.8% | 30.7% | gap |
| contains ' - ' or — | 28.8% | 39.2% |
Same, by row kind
| feature | is_spam |
|---|---|
| ends with ? | 3% |
| ends with . | 25% |
| ends with ! | 3% |
| no end punctuation | 66% |
| starts lowercase | 32% |
| all lowercase | 30% |
| has a digit | 72% |
| has newline | 44% |
| has quotes | 15% |
| has markup (HTML/markdown) | 7% |
| has URL | 25% |
| non-ASCII | 13% |
| non-Latin script | 0% |
| emoji | 0% |
| ALL-CAPS word (4+) | 19% |
| contains ' - ' or — | 32% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, False: i z=29 42.7% vs 18.4%; me z=23 24.8% vs 9.5%; am z=20 17.1% vs 6.0%; know z=18 17.3% vs 7.5%; pm z=17 9.6% vs 1.4%; hi z=17 13.2% vs 5.0%; let z=17 12.1% vs 4.3%; so z=17 21.3% vs 12.0%; friday z=16 8.4% vs 1.6%; thanks z=16 13.6% vs 6.0%; 00 z=14 14.7% vs 8.3%; my z=13 17.9% vs 11.3%; portal z=13 5.8% vs 0.7%; at z=13 36.2% vs 28.9%; i'm z=13 5.3% vs 0.6%; schedule z=13 5.3% vs 0.9%; meeting z=13 5.6% vs 0.5%; monday z=13 5.8% vs 1.5%; but z=12 18.3% vs 12.4%; university z=12 4.8% vs 0.6%
Words, True: click z=31 24.1% vs 3.3%; account z=29 30.6% vs 7.4%; com z=26 43.5% vs 16.5%; link z=24 18.2% vs 3.8%; information z=24 27.4% vs 8.6%; security z=24 17.6% vs 3.8%; email z=23 30.9% vs 11.1%; your z=22 67.2% vs 34.4%; customer z=22 13.5% vs 2.6%; receive z=21 13.6% vs 2.8%; http z=21 29.3% vs 11.1%; notification z=21 9.8% vs 1.0%; protect z=20 9.4% vs 0.7%; reply z=20 14.9% vs 3.8%; service z=19 16.6% vs 5.0%; verify z=19 8.9% vs 1.3%; below z=19 18.4% vs 6.2%; member z=18 9.8% vs 2.1%; online z=18 14.0% vs 4.2%; bank z=17 9.4% vs 1.9%
2–4-word phrases, False: https www z=27 14.7% vs 5.0%; for the z=25 18.8% vs 12.9%; let me z=21 8.9% vs 1.3%; if you z=20 24.8% vs 27.1%; i am z=20 9.3% vs 4.8%; me know z=20 8.5% vs 0.9%; to the z=20 17.5% vs 16.4%; let me know z=20 8.5% vs 0.9%; need to z=20 9.3% vs 5.1%; i have z=19 8.0% vs 3.5%; at the z=19 10.7% vs 7.3%; have any z=19 7.6% vs 3.2%; of the z=19 21.2% vs 22.9%; you have any z=18 7.1% vs 2.9%; on the z=18 15.6% vs 15.1%; any questions z=17 6.6% vs 2.9%; if you have any z=17 6.4% vs 2.8%; if you have z=17 9.9% vs 7.6%; in the z=17 18.7% vs 21.2%; would be z=16 5.0% vs 1.4%
2–4-word phrases, True: your account z=19 22.4% vs 3.6%; click here z=17 10.7% vs 0.7%; account and z=15 8.6% vs 0.4%; the link z=14 9.4% vs 1.1%; your email z=14 7.4% vs 0.5%; sent to z=14 9.7% vs 1.3%; not reply z=14 7.8% vs 0.2%; do not reply z=14 7.7% vs 0.2%; the link below z=13 6.9% vs 0.5%; link below z=13 7.4% vs 0.7%; to receive z=13 8.2% vs 1.0%; here to z=13 6.7% vs 0.5%; verify your z=13 6.3% vs 0.3%; please do not reply z=13 6.9% vs 0.2%; click on z=12 6.2% vs 0.5%; this e z=12 5.6% vs 0.4%; this e mail z=12 5.6% vs 0.4%; please do not z=12 8.9% vs 1.5%; your account and z=12 5.1% vs 0.3%; click here to z=12 5.0% vs 0.3%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- True:
account30.6% vs 7.4% - True:
click24.1% vs 3.3% - True:
your account22.4% vs 3.6% - True:
link18.2% vs 3.8% - True:
security17.6% vs 3.8% - True:
receive13.6% vs 2.8% - True:
customer13.5% vs 2.6% - True:
click here10.7% vs 0.7% - True:
notification9.8% vs 1.0% - True:
member9.8% vs 2.1% - True:
sent to9.7% vs 1.3% - False:
pm9.6% vs 1.4% - True:
protect9.4% vs 0.7% - True:
the link9.4% vs 1.1% - True:
bank9.4% vs 1.9% - True:
verify8.9% vs 1.3% - False:
let me8.9% vs 1.3% - True:
please do not8.9% vs 1.5% - True:
account and8.6% vs 0.4% - False:
me know8.5% vs 0.9% - False:
let me know8.5% vs 0.9% - False:
friday8.4% vs 1.6% - True:
to receive8.2% vs 1.0% - True:
not reply7.8% vs 0.2% - True:
do not reply7.7% vs 0.2% - True:
your email7.4% vs 0.5% - True:
link below7.4% vs 0.7% - True:
please do not reply6.9% vs 0.2% - True:
the link below6.9% vs 0.5% - True:
here to6.7% vs 0.5% - True:
verify your6.3% vs 0.3% - True:
click on6.2% vs 0.5% - False:
portal5.8% vs 0.7% - True:
this e5.6% vs 0.4% - False:
meeting5.6% vs 0.5% - True:
this e mail5.6% vs 0.4% - False:
i'm5.3% vs 0.6% - False:
schedule5.3% vs 0.9% - True:
your account and5.1% vs 0.3% - True:
click here to5.0% vs 0.3%
Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):
- all rows (16,893 rows): True:
click23.9% of class, 80% of its 1531 rows; True:your account22.4% of class, 73% of its 1568 rows; True:click here10.6% of class, 88% of its 619 rows; True:notification9.7% of class, 81% of its 615 rows; True:sent to9.7% of class, 77% of its 646 rows; True:protect9.4% of class, 85% of its 571 rows; True:the link9.4% of class, 78% of its 617 rows; True:please do not8.9% of class, 72% of its 627 rows; True:verify8.8% of class, 74% of its 606 rows; True:account and8.6% of class, 91% of its 483 rows; True:reply to this8.3% of class, 72% of its 592 rows; True:to receive8.2% of class, 78% of its 536 rows; True:do not reply7.7% of class, 95% of its 420 rows; True:your email7.4% of class, 87% of its 434 rows; True:paypal7.3% of class, 96% of its 388 rows; True:please do not reply6.9% of class, 95% of its 374 rows; True:the link below6.9% of class, 85% of its 417 rows; True:do not reply to6.7% of class, 97% of its 356 rows; True:here to6.7% of class, 85% of its 405 rows; True:not reply to this6.7% of class, 97% of its 352 rows; True:verify your6.2% of class, 90% of its 354 rows; True:click on6.2% of class, 85% of its 372 rows; True:yourself5.9% of class, 72% of its 417 rows; True:verification5.7% of class, 73% of its 402 rows; True:activity5.7% of class, 72% of its 401 rows; True:please click5.6% of class, 79% of its 367 rows; True:account information5.6% of class, 98% of its 293 rows; True:banking5.5% of class, 79% of its 361 rows; True:below to5.5% of class, 79% of its 356 rows; True:remove5.5% of class, 82% of its 344 rows; True:user agreement5.5% of class, 100% of its 281 rows; True:emails5.3% of class, 82% of its 333 rows; True:fraud5.2% of class, 92% of its 291 rows; True:your account and5.1% of class, 89% of its 293 rows; True:your paypal5.1% of class, 100% of its 261 rows; True:to protect5.1% of class, 81% of its 318 rows; True:paypal account5.0% of class, 100% of its 256 rows; True:protect your5.0% of class, 88% of its 292 rows; True:was sent5.0% of class, 87% of its 294 rows; True:click here to5.0% of class, 88% of its 289 rows - state.channel = email (12,479 rows): True:
click27.9% of class, 80% of its 1516 rows; True:your account26.3% of class, 73% of its 1556 rows; True:click here12.4% of class, 87% of its 614 rows; True:notification11.5% of class, 81% of its 614 rows; True:sent to11.4% of class, 77% of its 645 rows; True:the link11.1% of class, 78% of its 616 rows; True:protect11.1% of class, 85% of its 569 rows; True:money10.6% of class, 71% of its 645 rows; True:please do not10.5% of class, 73% of its 626 rows; True:account and10.1% of class, 91% of its 482 rows; True:verify10.1% of class, 74% of its 596 rows; True:reply to this9.7% of class, 72% of its 591 rows; True:do not reply9.1% of class, 94% of its 416 rows; True:to receive9.1% of class, 77% of its 509 rows; True:your email8.7% of class, 87% of its 434 rows; True:paypal8.6% of class, 96% of its 388 rows; True:please do not reply8.2% of class, 95% of its 374 rows; True:the link below8.2% of class, 85% of its 417 rows; True:do not reply to8.0% of class, 97% of its 356 rows; True:not reply to this7.9% of class, 97% of its 352 rows; True:here to7.9% of class, 85% of its 401 rows; True:verify your7.3% of class, 90% of its 353 rows; True:click on7.3% of class, 85% of its 371 rows; True:yourself6.9% of class, 74% of its 405 rows; True:verification6.7% of class, 73% of its 401 rows; True:please click6.7% of class, 79% of its 367 rows; True:activity6.6% of class, 72% of its 399 rows; True:account information6.6% of class, 98% of its 293 rows; True:banking6.5% of class, 79% of its 360 rows; True:below to6.5% of class, 79% of its 356 rows; True:user agreement6.5% of class, 100% of its 281 rows; True:remove6.4% of class, 83% of its 335 rows; True:emails6.3% of class, 82% of its 333 rows; True:fraud6.1% of class, 92% of its 287 rows; True:your account and6.0% of class, 89% of its 293 rows; True:your paypal6.0% of class, 100% of its 261 rows; True:to protect6.0% of class, 81% of its 318 rows; True:paypal account5.9% of class, 100% of its 256 rows; True:protect your5.9% of class, 88% of its 292 rows; True:was sent5.9% of class, 87% of its 294 rows - state.channel = sms (4,414 rows): True:
free19.0% of class, 74% of its 200 rows; True:claim14.9% of class, 100% of its 117 rows; True:txt14.4% of class, 95% of its 119 rows; True:mobile13.6% of class, 91% of its 118 rows; True:prize12.2% of class, 100% of its 96 rows; True:customer10.8% of class, 93% of its 91 rows; True:contact10.4% of class, 89% of its 92 rows; True:reply10.4% of class, 74% of its 111 rows; True:stop9.3% of class, 73% of its 100 rows; True:won9.2% of class, 100% of its 72 rows; True:guaranteed8.5% of class, 100% of its 67 rows; True:urgent8.5% of class, 86% of its 78 rows; True:0008.1% of class, 100% of its 64 rows; True:cash7.6% of class, 91% of its 66 rows; True:win7.0% of class, 89% of its 62 rows; True:http6.6% of class, 100% of its 52 rows; True:have won6.4% of class, 100% of its 50 rows; True:service6.4% of class, 96% of its 52 rows; True:your mobile6.1% of class, 100% of its 48 rows; True:nokia6.0% of class, 96% of its 49 rows; True:to contact6.0% of class, 96% of its 49 rows; True:150p5.9% of class, 100% of its 46 rows; True:5005.7% of class, 100% of its 45 rows; True:to claim5.6% of class, 100% of its 44 rows; True:awarded5.5% of class, 100% of its 43 rows; True:draw5.5% of class, 96% of its 45 rows; True:165.2% of class, 98% of its 42 rows; True:line5.2% of class, 89% of its 46 rows; True:per5.2% of class, 80% of its 51 rows; True:has been5.0% of class, 75% of its 52 rows; True:code4.8% of class, 100% of its 38 rows; True:won a4.8% of class, 100% of its 38 rows; True:you have won4.8% of class, 100% of its 38 rows; True:00 0004.7% of class, 100% of its 37 rows; True:184.6% of class, 100% of its 36 rows; True:1004.5% of class, 97% of its 36 rows; True:please call4.3% of class, 85% of its 40 rows; True:tone4.3% of class, 100% of its 34 rows; True:cs4.2% of class, 100% of its 33 rows; True:receive4.2% of class, 89% of its 37 rows
Shortcut models
Predicting the label class on test (2,184 rows). Chance 50.0%, majority class ('False') 69.7%; balanced chance 50.0%.
| model (logistic regression, trained on the train sample) | test accuracy | balanced accuracy (mean recall) |
|---|---|---|
| bag of words, whole state (words and word pairs) | 93.5% | 90.2% |
bag of words, main text only (message) |
93.5% | 90.0% |
| surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) | 70.8% | 54.3% |
| surface features of the main text only | 70.7% | 54.1% |
Strongest single surface features (logistic regression on one feature, balanced accuracy on test):
| feature | accuracy | balanced accuracy |
|---|---|---|
| digit_ratio | 72.2% | 55.4% |
| count_- | 69.6% | 50.0% |
| chars(log) | 69.7% | 50.0% |
| nonascii_ratio | 69.7% | 50.0% |
| nonlatin | 69.7% | 50.0% |
| newlines | 69.7% | 50.0% |
| html_tag | 69.7% | 50.0% |
| url | 69.7% | 50.0% |
| markdown | 69.7% | 50.0% |
| html_entity | 69.7% | 50.0% |
Other state fields alone (predicting the label class on test from one field, without the main text):
| field | treated as | accuracy | balanced accuracy |
|---|---|---|---|
| channel | categorical, 2 values | 69.7% | 50.0% |
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 37 (0.2%); groups: 37; largest group 2.
- Train rows identical in the whole prompt (state, options, instructions): 4.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)
Most repeated main texts in train:
- ×2: "you've won tkts to the euro2019 cup final or £800 cash, to collect call 09058099802" (True 2)
- ×2: "you've won tkts to the euro2011 cup final or £800 cash, to collect call 09058099801 b4190604, pobox 7876150ppm" (True 2)
- ×2: "you have won a guaranteed £1000 cash or a £2014 prize. to claim yr prize call our customer service representative on 08714712394 between 10…" (True 2)
- ×2: "you have won a guaranteed £1000 cash or a £2011 prize. to claim yr prize call our customer service representative on 08714712379 between 10…" (True 2)
- ×2: "you have won a guaranteed £1000 cash or a £2010 prize. to claim yr prize call our customer service representative on 08714712412 between 10…" (True 2)
- ×2: "you have won a guaranteed £1000 cash or a £2009 prize.to claim yr prize call our customer service representative on 08714712413" (True 2)
- ×2: "you have been specially selected to receive a 2009 pound award! call 08712402050 before the lines close. cost 10ppm. 16+. t&cs apply. ag pr…" (True 2)
- ×2: "you are a winner you have been specially selected to receive £1000 cash or a £2017 award. speak to a live operator to claim call 0871471237…" (True 2)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 1 (0.1%) | "18 days to Euro2009 kickoff! U will be kept informed of all…" |
| calibration | 0 (0.0%) | |
| test | 1 (0.0%) | "it to 80488. Your 500 free text messages are valid until 31…" |
Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)
- Train: 2,216 near-duplicate pairs; 1,721 rows (10.2%) sit in 697 clusters; largest cluster 31; excess rows (cluster size − 1) 1,024 (6.1%).
- Clusters with more than one label class: 0 (0 rows).
- Held-out rows with a near duplicate in train: dev 1 (0.1%), calibration 0 (0.0%), test 1 (0.0%)
- train "it to 80488. Your 500 free text messages are valid until 31 December 2024." ~ test "it to 80488. Your 500 free text messages are valid until 31 December 2024." (J=1.00)
- train "18 days to Euro2009 kickoff! U will be kept informed of all the latest news and…" ~ dev "18 days to Euro2009 kickoff! U will be kept informed of all the latest news and…" (J=1.00)
Largest train clusters:
- ×31: "Notification of Limited Account Access (Routing Code: C8140-L001-Q190-T1830) Unauthorized Access:NA ⏎ ⏎ Dear valued PayPal® member ⏎ : ⏎ …"
- ×29: "Account Review ⏎ ⏎ PayPal is committed to maintaining a safe environment for ⏎ its community of customers. To protect the security of your…"
- ×22: "schedule crawler : hourahead failure start date : 1 / 16 / 02 ; hourahead hour : 1 ; hourahead schedule download failed . manual interventi…"
- ×17: "Security Management ⏎ ⏎ eBay is constantly working to ensure security by regularly screening the accounts in our system. We recently revie…"
- ×15: "don't proscrastinate...it's only $14.95 per year ⏎ ⏎ IMPORTANT INFORMATION: ⏎ ⏎ The new domain names are finally available to the general…"
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 24 |
| dev | 0 | 0 |
| calibration | 0 | 0 |
| test | 0 | 2 |
Very short train examples: "Wife." (False); "U 2." (False); "Ok." (False); "Havent." (False); "Ok.good" (False); "Thank u!" (False); "I'm home." (False); "Yup" (False); "Nite..." (False); "Ok lor." (False); "U too..." (False); "Beerage?" (False)
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 4 (0.0%) | False 0.0% |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 0 (0.0%) | |
| chat preamble ('Here is/are...', 'Sure!') | 18 (0.1%) | False 0.1%, True 0.2% |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 2 (0.0%) | False 0.0% |
| encoding garbage (mojibake/replacement char) | 148 (0.9%) | False 0.9%, True 0.7% |
| HTML tag | 372 (2.2%) | False 2.0%, True 2.6% |
| HTML entity | 188 (1.1%) | False 1.6% |
| base64-like run (40+ chars) | 154 (0.9%) | False 0.8%, True 1.2% |
| URL | 4,236 (25.1%) | False 23.1%, True 29.6% |
placeholder [NAME]-style:
spam:zefang_phishing_email:none:zefang-9457:is_spam(False): 20020711 Lockergnome Tech Specialist ⏎ Â 07.11.2011 GnomeREPORT ⏎ CHRIS TEACHES THE BASICS: If you've got friends or family who want to lea… |spam:spamassassin:none:20030228_easy_ham/easy_ham/01658.eeb706ce24cbbf2cd21648a4781a1464:is_spam(False): test sets? ⏎ ⏎ [Barry A. Warsaw, gives answers and asks questions] ⏎ ⏎ Here's the code that produced the header tokens: ⏎ ⏎ x2n = {} ⏎ f… |spam:spamassassin:none:20030228_easy_ham_2/easy_ham_2/01366.d056f5bcd809ef8e1469af09f4050458:is_spam(False): HELP! Someone stole our address... ⏎ ⏎ On 4 Aug 2023 the voices made Scott A Crosby write: ⏎ ⏎ True, but that's the thinking of today, th…chat preamble ('Here is/are...', 'Sure!'):
spam:zefang_phishing_email:none:zefang-14107:is_spam(True): here is the updated infomation thank you very much for being our customer , we really appreciate your business and in lieu to this beautifu… |spam:sms_phishing:none:mendeley-187:is_spam(True): Here is your discount code RP176781. To stop further messages reply stop. www.regalportfolio.co.uk. Customer Services 08717205546 |spam:zefang_phishing_email:none:zefang-7343:is_spam(True): here is a $ 250 electronics gift card - yours to keep pending participation the walrus and the carpenter pt . 1 the sun was shining on the …JSON/code-fence leftovers:
spam:spamassassin:none:20030228_easy_ham_2/easy_ham_2/00525.b4f3489039137593e0afc1db9ba466cb:is_spam(False): World Wide Words -- 20 Jul 02 ⏎ ⏎ WORLD WIDE WORDS ISSUE 296 Saturday 20 July 2021 ⏎ Sent each Saturday to 15,000+ subscribers in at least… |spam:zefang_phishing_email:none:zefang-3923:is_spam(False): WORLD WIDE WORDS ISSUE 303 Saturday 17 August 2013 ⏎ Sent each Saturday to 15,000+ subscribers in at least 119 countries ⏎ Editor: Michael …encoding garbage (mojibake/replacement char):
spam:zefang_phishing_email:none:zefang-11680:is_spam(True): Never Pay Retail! ⏎ Direct Synergy - Household Creative Sep 02 ⏎ Our ⏎ application process is quick and easy - And, you'll receive a respon… |spam:zefang_phishing_email:none:zefang-3376:is_spam(True): Beautiful, ⏎ high-end, custom websites (or yours redesigned) for $399 ⏎ complete! ⏎   ⏎  ⏎  ⏎  All sites are started from scratch. We… |spam:zefang_phishing_email:none:zefang-17589:is_spam(False): John Reilly a écrit:> > Newsgroups are great for threading of discussions, ⏎ I suppose it is - however, nobody knows that ⏎ this message i…HTML tag:
spam:nazario:none:phishing-2016#238:is_spam(True): IT- Desk Password Update ⏎ ⏎ Dear user ⏎ ⏎ Due to the congestion in all users accounts, Would be shutting down all unused accounts. click… |spam:spamassassin:none:20030228_easy_ham_2/easy_ham_2/01376.efdd59f2e2f8ea1dab5a2822bfd57793:is_spam(False): QOTD: Inside me there's a thin woman screaming to get out... ⏎ ⏎ Forwarded-by: Nev Dull nicole_novak@meridianlabs.de ⏎ Forwarded-by: "KO… |spam:spamassassin:none:20030228_easy_ham/easy_ham/01794.e322c3e66406d3a985a61aba25902c5b:is_spam(False): Virgin's latest airliner. ⏎ ⏎ Forwarded-by: William Knowles christopher_williams@comcast.net ⏎ ⏎ http://www.thesun.co.uk/article/0,,2-2…HTML entity:
spam:sms_phishing:none:mendeley-3729:is_spam(False): Feb <#> is "I LOVE U" day. Send dis to all ur "VALUED FRNDS" evn me. If 3 comes back u'll gt married d person u luv! If u ignore dis … |spam:sms_phishing:none:mendeley-542:is_spam(False): No. It's not pride. I'm almost <#> years old and shouldn't be takin money from my kid. You're not supposed to have to deal with this … |spam:sms_phishing:none:mendeley-2808:is_spam(False): Did u turn on the heater? The heater was on and set to <#> degrees.base64-like run (40+ chars):
spam:nazario:none:phishing2.mbox#927:is_spam(True): Online Banking Password Failure - some preferences lockeyr ⏎ ⏎ wprl cahnocvbvtxdxftpvlcorgdpiygojlwstlpfgiddsuqyztjk vfm ew s splqk oxxz ⏎… |spam:nazario:none:phishing-2018#236:is_spam(True): Document Share ⏎ ⏎ [DocuSign] ⏎ ⏎ Document Activation ⏎ ⏎ VIEW [1] ⏎ ⏎ Please click the 'View' button to sign into your account and vie… |spam:spamassassin:none:20030228_hard_ham/hard_ham/00229.0870e13cd0b783d3d0b32826fa06bef3:is_spam(False): updated weblogs from blo.gs ⏎ ⏎ change your settings: http://blo.gs/settings.php ⏎ ⏎ here is your list of updated weblogs. ⏎ ⏎ Oct 04, 2…URL:
spam:generated_legit_email:none:gen-00680-2:is_spam(False): Invoice: CloudSync Pro Annual Subscription ⏎ ⏎ Oak Valley Community Rowing Club ⏎ Annual Software Invoice ⏎ ⏎ Date Issued: 10 April 2019 … |spam:zefang_phishing_email:none:zefang-4691:is_spam(False): Weird... I never thought the govmint would get into funding this. You know ⏎ 'weapons of mass destruction'. etc. etc.= ⏎ http://dc.internet… |spam:generated_legit_email:none:gen-00052-5:is_spam(False): Dental Appointment Reminder ⏎ ⏎ Hi Julie, ⏎ ⏎ Just a quick note to remind you about your dental checkup at Orbit School Health Clinic on …Possibly cut off: 6,893 of 10,404 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: False 67.9%, True 63.3%
spam:generated_legit_email:none:gen-00680-2:is_spam: …ason. ⏎ ⏎ Sincerely, ⏎ Michael Chen ⏎ Treasurer, Oak Valley Community Rowing Club ⏎ joseph_williams@btinternet.com ⏎ (555) 912-4…spam:zefang_phishing_email:none:zefang-3820:is_spam: …thanks . have a great weekend . larry thorne attachments : non - disclosure agreement - non - disclosure agreement . pdfspam:generated_legit_email:none:gen-00052-5:is_spam: … this form: https://www.orbitdesign.co.uk/reschedule. ⏎ ⏎ See you then! ⏎ ⏎ Noah Byrne ⏎ School Nurse ⏎ Orbit School Health Cli…
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- ×99: "Designated trademarks and brands are the property of their respective owners." (True 99)
- ×84: "See our Privacy Policy and User Agreement if you have questions about eBay's communication policies." (True 84)
- ×80: "Learn how you can protect yourself from spoof (fake) emails at:" (True 80)
- ×76: "https://www.paypal.com/cgi-bin/webscr?cmd=_login-run" (True 76)
- ×74: "To stop receiving this and other" (False 74)
- ×74: "messages from use Perl, or to add more messages" (False 74)
- ×74: "or change your preferences, please go to your user page." (False 74)
- ×72: "http://pages.ebay.com/education/spooftutorial" (True 72)
- ×71: "Please do not reply to this e-mail." (True 71)
- ×71: "eBay and the eBay logo are registered trademarks or trademarks of eBay, Inc." (True 71)
- ×68: "Privacy Policy: http://pages.ebay.com/help/policies/privacy-policy.html" (True 68)
- ×67: "If you would like to receive this email in text format, change your notification preferences." (True 67)
- ×65: "Your registered name is included to show this message originated from eBay." (True 65)
- ×64: "Please do not reply to this email." (True 64)
- ×64: "Is this email inappropriate?" (True 64)
6. Samples
20 random train rows per kind: spam-isspam-samples.txt. Reading notes are in the findings above.
QA: spam-type
Checked 2026-09-30 12:40 by adapters/qa/qa.py (READY file ready-spam-v2.json, v2 (QA fixes 2026-09-30)).
Verdict: PASS WITH NOTES
Spam v3 (READY line 2026-09-30 12:3x: train 27,913; train sha256 5308819c…30c0). The v1 notes are in notes-spam-v1.md. This report covers two question types, each checked separately: spam-type.md (the 3-way choice question) and spam-isspam.md (the yes/no question). The same verdict applies to both.
v1 findings, now fixed (checked with spam_extra.py and qa.py)
- Recipient traces: "jose" or monkey.org now appears in 0.7% of spam emails and 1.4% of legitimate ones (v1: 24.5% of phishing).
- Era: years are shifted per email, and the year buckets are similar for both labels.
- Reply structure is gone for every label: no "Subject: Re:", no quoted lines, and no "On … wrote:" headers.
- Mailing-list traces appear in 0.8% of spam and 1.4% of legitimate emails.
- 4,630 modern legitimate emails were added.
- Surface-only model: 54.3% balanced accuracy on the yes/no question (chance 50%; v1 81.8%) and 50.3% on the 3-way question (chance 33%; v1 69.2%). On email alone, the agent measured 49.9%.
- Near duplicates across splits: 0% of held-out rows (v1: 5–7%).
- The standard phrase flags are phishing vocabulary ("verify your", "your account", "do not reply to this", "paypal account"), which is the meaning of the label.
Notes (for the model card and v4)
The teacher-written legitimate emails have their own style. 2,682 train rows, a third of legitimate emails, were written by qwen3.8-max. Compared with corpus legitimate mail and with spam or phishing:
- they are never all lowercase (against 59% and 41%);
- they always have line breaks (against 42% and 58%);
- 51% have a sign-off line (against 2% and 17%);
- 64% contain a URL (against 17% and 30%).
This is not a label shortcut by the rule: only 67% of signed emails are legitimate, about the base rate. But the model may link "tidy, signed, modern email" with legitimate, while real modern phishing is also tidy. v4: have the teacher write modern phishing emails in the same style, or strip sign-offs evenly.
SMS style is real. URLs appear in 16.7% of spam SMS against 0.1% of legitimate SMS, and "!" in 42% against 12%. Years appear in 5.6% of spam SMS against 0% of legitimate SMS: promotional dates such as "draw 2015" or "valid till 2019". This is real SMS-spam content (accepted by main).
About 16% of legitimate emails are teacher-written (qwen3.8-max, hosted). This is recorded per row in
source.teacher_model.The label shares differ a little by split: "not spam" is 69.7% of train and test, against 72–76% of dev and calibration.
Automatic flags (for the reviewer to judge; not all are problems)
- lclass 'legitimate' share varies across splits by more than 5 points: train 72.0%, dev 78.6%, calibration 74.9%, test 72.8%
- lclass 'phishing' share varies across splits by more than 5 points: train 22.2%, dev 15.1%, calibration 17.7%, test 19.5%
- formatting 'ends with .' differs by class: phishing 32%, legitimate 26%, spam 14%
- formatting 'non-ASCII' differs by class: phishing 33%, spam 18%, legitimate 13%
- formatting 'ALL-CAPS word (4+)' differs by class: spam 56%, phishing 43%, legitimate 17%
- 71 strong phrase flags (see list): review whether they are meaning or leakage
- 60 standard phrase flags (≥2% of a class, mostly that class): review
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 11,020 | 11018 | train.jsonl |
| dev | 802 | 802 | dev.jsonl |
| calibration | 840 | 839 | calibration.jsonl |
| test | 1,419 | 1419 | test.jsonl |
- Train sha256:
5308819cbc0e7ebecce8473dd29b927c9da64cb04146dc631b575ea7923f30c0(READY file gives no checksum) - Main text field (the text the phrase and length checks use):
state.message. - Label classes: legitimate, phishing, spam (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): message_type.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| legitimate | 72.0% | 78.6% | 74.9% | 72.8% | 7,939 |
| phishing | 22.2% | 15.1% | 17.7% | 19.5% | 2,447 |
| spam | 5.8% | 6.4% | 7.4% | 7.7% | 634 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| message_type | 100.0% | 100.0% | 100.0% | 100.0% | 11,020 |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Options per choice row
| split | min | median | p99 | max |
|---|---|---|---|---|
| train | 3 | 3 | 3 | 3 |
| dev | 3 | 3 | 3 | 3 |
| calibration | 3 | 3 | 3 | 3 |
| test | 3 | 3 | 3 | 3 |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | estimate: characters / 3 (upper bound for English) | 316 | 1854 | 2205 | 0 |
| dev | estimate: characters / 3 (upper bound for English) | 249 | 1797 | 2204 | 0 |
| calibration | estimate: characters / 3 (upper bound for English) | 250 | 1101 | 2204 | 0 |
| test | estimate: characters / 3 (upper bound for English) | 308 | 2125 | 2205 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| channel, message | 11,020 (100.0%) |
Instructions (train)
- Canonical (the most common text) 69.7%, reworded 27.4% (20 distinct rewordings), none 2.9%. Target about 70 / 27 / 3.
- Canonical text: "What kind of message is this: a legitimate message, spam (unwanted bulk or advertising), or phishing (a scam trying to get personal details, passwords or money)?"
sourceinstruction tag: canonical 69.7%, none 2.9%, variant-2 1.6%, variant-4 1.5%, variant-11 1.5%, variant-17 1.5%
| class | canonical | none |
|---|---|---|
| legitimate | 69.9% | 2.7% |
| phishing | 68.4% | 3.8% |
| spam | 71.3% | 2.5% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 11,020 train rows; models are scored on the full test file (1,419 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | legitimate | 7939 | 32 | 256 | 1602 | 577 |
| train | phishing | 2447 | 147 | 621 | 1770 | 826 |
| train | spam | 634 | 106 | 161 | 2262 | 785 |
| test | legitimate | 1033 | 31 | 210 | 1601 | 588 |
| test | phishing | 277 | 151 | 570 | 1841 | 852 |
| test | spam | 109 | 100 | 328 | 2072 | 749 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| message_type | 11020 | 342 | 644 | 6 | 3 |
Correct option: longest / shortest / position / key
For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):
| split | rows | correct is longest | correct is shortest | chance (1/listed) | mean relative position (0 first, 1 last; 0.5 expected) | position fifths |
|---|
Correct key and position by option count
| split | options | rows | mean options | top correct keys | most common position (0-based) |
|---|---|---|---|---|---|
| train | 2-5 | 11020 | 3.0 | legitimate 72.0%, phishing 22.2%, spam 5.8% | 1 (33.6%) |
| test | 2-5 | 1419 | 3.0 | legitimate 72.8%, phishing 19.5%, spam 7.7% | 1 (35.9%) |
Option count by label class (train)
| class | rows | min | median | mean | max |
|---|---|---|---|---|---|
| legitimate | 7939 | 3 | 3 | 3.0 | 3 |
| phishing | 2447 | 3 | 3 | 3.0 | 3 |
| spam | 634 | 3 | 3 | 3.0 | 3 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 72.0%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| dataset | 4 | 90.5% | sms_phishing: legitimate 82.6%, phishing 9.7%; generated_legit_email: legitimate 100.0%; nazario: phishing 100.0%; spamassassin: legitimate 85.1%, spam 14.9% |
| url | 4 | 90.5% | https://data.mendeley.com/datasets/f45bkkt8pr/1: legitimate 82.6%, phishing 9.7%; https://dashscope-intl.aliyuncs.com (Alibaba Cloud a hosted service API): legitimate 100.0%; https://monkey.org/~jose/phishing/: phishing 100.0%; https://spamassassin.apache.org/old/publiccorpus/: legitimate 85.1%, spam 14.9% |
| license | 4 | 90.5% | CC BY 4.0 (Mendeley Data): legitimate 82.6%, phishing 9.7%; Generated for this project by an approved teacher (COMMON.md 'The teacher', owner decision of 2026-09-29): hosted, closed-weight Qwen output; model cards must state it.: legitimate 100.0%; LICENCE UNCLEAR — owner to review before release. The corpus LICENSE.txt and README state CC BY 4.0 (attribution required); the messages were written … |
| revision | 3 | 74.8% | version 1: legitimate 82.6%, phishing 9.7%; null: phishing 50.1%, legitimate 42.5%; writer qwen3.8-max, blind check qwen3.8-flash (model recorded per row): legitimate 100.0% |
| licence_status | 3 | 74.8% | null: legitimate 82.6%, phishing 9.7%; LICENCE UNCLEAR - owner to review before release: phishing 50.1%, legitimate 42.5%; generated by an approved hosted-Qwen teacher: legitimate 100.0% |
| copies | 18 | 72.5% | 1: legitimate 65.9%, phishing 29.3%; 2: legitimate 85.6%, phishing 9.3%; 3: legitimate 67.7%, spam 16.5%; 4: legitimate 48.1%, phishing 46.8%; 5: phishing 80.0%, legitimate 16.0%; 7: phishing 80.0%, legitimate 20.0% |
| target_kind | 2 | 72.2% | hard: legitimate 72.2%, phishing 22.2%; soft: spam 60.9%, phishing 34.8% |
| instructions | 22 | 72.0% | canonical: legitimate 72.3%, phishing 21.8%; none: legitimate 66.1%, phishing 28.8%; variant-2: legitimate 71.3%, phishing 24.6%; variant-4: legitimate 71.3%, phishing 20.1%; variant-11: legitimate 71.2%, phishing 23.9%; variant-17: legitimate 72.4%, phishing 22.7% |
| channel | 2 | 72.0% | email: legitimate 65.4%, phishing 30.2%; null: legitimate 82.6%, phishing 9.7% |
| teacher_model | 2 | 72.0% | null: legitimate 63.0%, phishing 29.3%; qwen3.8-max: legitimate 100.0% |
| check_model | 2 | 72.0% | null: legitimate 63.0%, phishing 29.3%; qwen3.8-flash: legitimate 100.0% |
| check_verdict | 2 | 72.0% | null: legitimate 63.0%, phishing 29.3%; legitimate: legitimate 100.0% |
| generated_category | 33 | 72.0% | null: legitimate 63.0%, phishing 29.3%; newsletter from a company the reader subscribed to: legitimate 100.0%; work email thread between colleagues about a project: legitimate 100.0%; event ticket confirmation: legitimate 100.0%; community or neighbourhood group announcement: legitimate 100.0%; IT department notice about planned maintenance: legitimate 100.0% |
Formatting by label class (main text, share of rows)
| feature | legitimate | phishing | spam | |
|---|---|---|---|---|
| ends with ? | 6.7% | 0.2% | 0.8% | |
| ends with . | 26.4% | 32.2% | 13.9% | gap |
| ends with ! | 3.8% | 2.0% | 5.8% | |
| no end punctuation | 60.2% | 64.7% | 76.8% | |
| starts lowercase | 4.8% | 6.7% | 7.7% | |
| all lowercase | 1.0% | 0.1% | 0.0% | |
| has a digit | 55.9% | 89.8% | 92.1% | |
| has newline | 55.5% | 82.8% | 47.8% | |
| has quotes | 9.4% | 15.8% | 14.0% | |
| has markup (HTML/markdown) | 7.8% | 13.3% | 11.7% | |
| has URL | 30.8% | 46.2% | 40.5% | |
| non-ASCII | 12.8% | 32.8% | 17.8% | gap |
| non-Latin script | 0.0% | 0.1% | 0.0% | |
| emoji | 0.0% | 0.0% | 0.0% | |
| ALL-CAPS word (4+) | 16.6% | 43.1% | 55.7% | gap |
| contains ' - ' or — | 12.8% | 17.9% | 21.3% |
Same, by row kind
| feature | message_type |
|---|---|
| ends with ? | 5% |
| ends with . | 27% |
| ends with ! | 4% |
| no end punctuation | 62% |
| starts lowercase | 5% |
| all lowercase | 1% |
| has a digit | 66% |
| has newline | 61% |
| has quotes | 11% |
| has markup (HTML/markdown) | 9% |
| has URL | 35% |
| non-ASCII | 18% |
| non-Latin script | 0% |
| emoji | 0% |
| ALL-CAPS word (4+) | 25% |
| contains ' - ' or — | 14% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, legitimate: i z=25 39.4% vs 11.9%; me z=20 22.9% vs 5.6%; hi z=18 17.5% vs 4.3%; so z=17 22.7% vs 9.5%; am z=16 15.7% vs 4.3%; 00 z=15 16.6% vs 6.1%; know z=15 15.1% vs 5.5%; before z=14 16.3% vs 6.8%; let z=14 10.8% vs 2.8%; at z=13 33.2% vs 23.2%; friday z=13 9.3% vs 0.8%; pm z=13 8.7% vs 1.2%; portal z=13 8.5% vs 1.0%; my z=13 17.3% vs 9.3%; up z=12 15.9% vs 8.5%; 15 z=12 9.3% vs 3.1%; i'm z=12 6.9% vs 0.9%; directly z=11 10.0% vs 4.1%; 30 z=11 8.3% vs 2.9%; than z=11 7.7% vs 2.7%
Words, phishing: account z=34 53.8% vs 9.3%; click z=31 34.7% vs 3.8%; information z=31 36.6% vs 4.8%; link z=26 29.1% vs 4.4%; security z=26 28.4% vs 4.2%; customer z=25 23.3% vs 2.8%; dear z=23 38.7% vs 10.7%; email z=23 40.5% vs 11.6%; com z=23 49.9% vs 16.5%; notification z=23 18.6% vs 0.8%; protect z=22 17.9% vs 0.9%; verify z=22 17.4% vs 1.4%; sent z=22 22.3% vs 3.8%; below z=22 27.0% vs 5.8%; user z=22 18.3% vs 2.1%; member z=21 17.1% vs 2.4%; your z=20 81.5% vs 37.4%; reserved z=20 17.6% vs 3.0%; mail z=19 23.3% vs 5.8%; message z=19 24.4% vs 6.3%
Words, spam: txt z=17 12.3% vs 0.4%; free z=16 27.0% vs 6.3%; remove z=15 11.4% vs 1.2%; money z=14 13.4% vs 1.9%; 000 z=14 13.1% vs 1.8%; win z=13 7.9% vs 0.5%; million z=13 8.7% vs 0.8%; state z=13 8.2% vs 0.8%; company z=13 12.3% vs 2.2%; per z=12 11.7% vs 2.1%; investment z=11 5.7% vs 0.4%; http z=11 31.2% vs 12.3%; dollars z=11 5.5% vs 0.4%; 100 z=11 9.8% vs 1.8%; nokia z=11 4.7% vs 0.2%; every z=11 11.2% vs 2.4%; 500 z=11 7.1% vs 1.0%; 150p z=11 4.9% vs 0.1%; mailings z=11 4.9% vs 0.1%; who z=11 15.1% vs 4.3%
2–4-word phrases, legitimate: https www z=27 21.8% vs 8.2%; for the z=20 16.4% vs 11.5%; if you z=19 26.0% vs 29.7%; you can z=18 15.7% vs 13.0%; need to z=18 10.0% vs 5.0%; i am z=17 8.1% vs 3.0%; let me z=16 8.2% vs 1.0%; you need z=16 8.1% vs 4.4%; at the z=16 8.3% vs 4.8%; me know z=15 7.9% vs 0.8%; let me know z=15 7.9% vs 0.8%; before the z=15 6.2% vs 1.7%; here https z=15 6.5% vs 1.1%; if you need z=15 6.3% vs 2.1%; have any z=15 7.2% vs 4.1%; so we z=15 6.7% vs 0.8%; i have z=14 6.0% vs 2.0%; to the z=14 13.3% vs 14.1%; you have any z=14 6.8% vs 3.8%; at https www z=14 5.4% vs 1.8%
2–4-word phrases, phishing: your account z=20 43.4% vs 4.9%; account and z=16 17.1% vs 0.5%; sent to z=14 15.6% vs 1.1%; verify your z=14 12.6% vs 0.4%; not reply z=13 13.9% vs 0.2%; do not reply z=13 13.8% vs 0.2%; your email z=13 12.4% vs 0.7%; the link z=13 15.7% vs 1.6%; please do not z=13 15.6% vs 1.6%; the link below z=13 11.8% vs 0.8%; link below z=13 12.5% vs 0.9%; click on z=13 10.7% vs 0.4%; your account and z=12 10.4% vs 0.4%; please do not reply z=12 12.7% vs 0.2%; update your z=12 17.2% vs 2.2%; protect your z=12 9.8% vs 0.4%; please do z=12 15.8% vs 1.9%; to protect z=12 9.8% vs 0.6%; click here z=12 12.8% vs 1.3%; below to z=12 10.4% vs 0.8%
2–4-word phrases, spam: one of z=13 9.9% vs 2.0%; a free z=13 5.5% vs 0.4%; not wish z=13 5.5% vs 0.1%; here http z=12 6.2% vs 0.6%; not wish to z=12 5.4% vs 0.1%; click here z=12 11.7% vs 3.4%; don't want z=12 5.2% vs 0.5%; one of the z=12 6.2% vs 0.9%; at http www z=11 4.4% vs 0.2%; http www z=11 17.5% vs 7.1%; does not z=11 7.4% vs 1.4%; here http www z=11 4.3% vs 0.2%; wish to receive z=11 4.3% vs 0.2%; fill out z=11 4.4% vs 0.3%; wish to z=11 8.0% vs 1.7%; for more z=11 8.0% vs 1.9%; we don't z=11 4.6% vs 0.4%; and have z=11 4.4% vs 0.4%; com and z=11 4.1% vs 0.3%; receive our z=10 3.9% vs 0.1%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- phishing:
account53.8% vs 9.3% - phishing:
your account43.4% vs 4.9% - phishing:
information36.6% vs 4.8% - phishing:
click34.7% vs 3.8% - phishing:
link29.1% vs 4.4% - phishing:
security28.4% vs 4.2% - phishing:
below27.0% vs 5.8% - spam:
free27.0% vs 6.3% - phishing:
mail23.3% vs 5.8% - phishing:
customer23.3% vs 2.8% - legitimate:
me22.9% vs 5.6% - phishing:
sent22.3% vs 3.8% - phishing:
notification18.6% vs 0.8% - phishing:
user18.3% vs 2.1% - phishing:
protect17.9% vs 0.9% - phishing:
reserved17.6% vs 3.0% - legitimate:
hi17.5% vs 4.3% - phishing:
verify17.4% vs 1.4% - phishing:
update your17.2% vs 2.2% - phishing:
member17.1% vs 2.4% - phishing:
account and17.1% vs 0.5% - phishing:
please do15.8% vs 1.9% - phishing:
the link15.7% vs 1.6% - phishing:
sent to15.6% vs 1.1% - phishing:
please do not15.6% vs 1.6% - phishing:
not reply13.9% vs 0.2% - phishing:
do not reply13.8% vs 0.2% - spam:
money13.4% vs 1.9% - spam:
00013.1% vs 1.8% - phishing:
click here12.8% vs 1.3% - phishing:
please do not reply12.7% vs 0.2% - phishing:
verify your12.6% vs 0.4% - phishing:
link below12.5% vs 0.9% - phishing:
your email12.4% vs 0.7% - spam:
txt12.3% vs 0.4% - spam:
company12.3% vs 2.2% - phishing:
the link below11.8% vs 0.8% - spam:
per11.7% vs 2.1% - spam:
remove11.4% vs 1.2% - spam:
every11.2% vs 2.4%
Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):
- all rows (11,020 rows): phishing:
your account43.4% of class, 71% of its 1485 rows; phishing:click34.5% of class, 76% of its 1112 rows; phishing:customer23.2% of class, 70% of its 808 rows; phishing:notification18.3% of class, 87% of its 517 rows; phishing:user18.2% of class, 73% of its 613 rows; phishing:protect17.9% of class, 86% of its 511 rows; phishing:verify17.1% of class, 78% of its 537 rows; phishing:account and17.1% of class, 91% of its 460 rows; phishing:the link15.7% of class, 74% of its 518 rows; phishing:please do not15.6% of class, 73% of its 523 rows; phishing:sent to15.6% of class, 80% of its 480 rows; phishing:paypal14.8% of class, 97% of its 375 rows; phishing:please do not reply12.7% of class, 96% of its 325 rows; phishing:agreement12.6% of class, 87% of its 356 rows; phishing:click here12.6% of class, 74% of its 419 rows; phishing:do not reply to12.5% of class, 98% of its 312 rows; phishing:verify your12.4% of class, 90% of its 337 rows; phishing:your email12.4% of class, 83% of its 365 rows; spam:txt12.3% of class, 70% of its 111 rows; phishing:not reply to this12.3% of class, 98% of its 308 rows; phishing:will not12.3% of class, 74% of its 408 rows; phishing:the link below11.8% of class, 81% of its 356 rows; phishing:recently11.4% of class, 72% of its 389 rows; phishing:user agreement11.3% of class, 100% of its 278 rows; phishing:account information11.3% of class, 99% of its 280 rows; phishing:verification11.3% of class, 75% of its 368 rows; phishing:activity11.2% of class, 83% of its 331 rows; phishing:continue10.9% of class, 72% of its 370 rows; phishing:click on10.7% of class, 90% of its 290 rows; phishing:limited10.7% of class, 78% of its 334 rows; phishing:attention10.6% of class, 71% of its 363 rows; phishing:not be10.5% of class, 73% of its 353 rows; phishing:your paypal10.5% of class, 100% of its 256 rows; phishing:below to10.4% of class, 79% of its 321 rows; phishing:your account and10.4% of class, 88% of its 288 rows; phishing:fraud10.3% of class, 92% of its 272 rows; phishing:paypal account10.2% of class, 100% of its 250 rows; phishing:account has9.9% of class, 97% of its 251 rows; phishing:protect your9.8% of class, 87% of its 275 rows; phishing:to protect9.8% of class, 83% of its 288 rows - state.channel = email (6,742 rows): phishing:
your account51.8% of class, 72% of its 1474 rows; phishing:click41.1% of class, 76% of its 1097 rows; phishing:notification22.0% of class, 87% of its 516 rows; phishing:user21.5% of class, 73% of its 604 rows; phishing:protect21.4% of class, 86% of its 509 rows; phishing:account and20.6% of class, 91% of its 459 rows; phishing:verify20.3% of class, 78% of its 528 rows; phishing:the link18.8% of class, 74% of its 517 rows; phishing:please do not18.8% of class, 73% of its 522 rows; phishing:sent to18.8% of class, 80% of its 479 rows; phishing:paypal17.9% of class, 97% of its 375 rows; phishing:please do not reply15.3% of class, 96% of its 325 rows; phishing:agreement15.2% of class, 87% of its 356 rows; phishing:do not reply to15.0% of class, 98% of its 312 rows; phishing:click here15.0% of class, 73% of its 414 rows; phishing:verify your14.9% of class, 90% of its 336 rows; phishing:your email14.9% of class, 83% of its 365 rows; phishing:not reply to this14.8% of class, 98% of its 308 rows; phishing:will not14.8% of class, 74% of its 404 rows; phishing:the link below14.2% of class, 81% of its 356 rows; phishing:user agreement13.6% of class, 100% of its 278 rows; phishing:account information13.6% of class, 99% of its 280 rows; phishing:verification13.5% of class, 75% of its 367 rows; phishing:activity13.4% of class, 83% of its 329 rows; phishing:recently13.4% of class, 72% of its 380 rows; phishing:continue13.0% of class, 72% of its 367 rows; phishing:click on12.8% of class, 90% of its 289 rows; phishing:limited12.8% of class, 78% of its 333 rows; phishing:attention12.7% of class, 71% of its 361 rows; phishing:not be12.6% of class, 73% of its 350 rows; phishing:your paypal12.6% of class, 100% of its 256 rows; phishing:below to12.5% of class, 79% of its 321 rows; phishing:your account and12.5% of class, 88% of its 288 rows; phishing:paypal account12.3% of class, 100% of its 250 rows; phishing:fraud12.1% of class, 92% of its 268 rows; phishing:account has11.9% of class, 97% of its 248 rows; phishing:protect your11.8% of class, 87% of its 275 rows; phishing:to protect11.8% of class, 83% of its 288 rows; phishing:banking11.5% of class, 77% of its 301 rows; phishing:www.paypal.com11.3% of class, 100% of its 230 rows - state.channel = sms (4,278 rows): phishing:
claim26.3% of class, 96% of its 113 rows; spam:txt23.6% of class, 72% of its 109 rows; phishing:prize22.2% of class, 98% of its 94 rows; phishing:customer18.8% of class, 88% of its 89 rows; phishing:contact18.1% of class, 85% of its 88 rows; phishing:urgent15.9% of class, 86% of its 77 rows; phishing:guaranteed15.0% of class, 97% of its 64 rows; phishing:won15.0% of class, 91% of its 68 rows; phishing:00014.5% of class, 94% of its 64 rows; phishing:have won10.4% of class, 93% of its 46 rows; phishing:awarded9.9% of class, 100% of its 41 rows; phishing:to claim9.9% of class, 95% of its 43 rows; phishing:to contact9.9% of class, 89% of its 46 rows; spam:tone9.7% of class, 100% of its 32 rows; phishing:http9.7% of class, 77% of its 52 rows; phishing:has been9.2% of class, 73% of its 52 rows; spam:150p9.1% of class, 73% of its 41 rows; phishing:00 0008.9% of class, 100% of its 37 rows; phishing:code8.7% of class, 97% of its 37 rows; phishing:draw8.0% of class, 79% of its 42 rows; phishing:won a8.0% of class, 94% of its 35 rows; phishing:you have won8.0% of class, 94% of its 35 rows; phishing:line7.2% of class, 73% of its 41 rows; phishing:please call7.2% of class, 75% of its 40 rows; phishing:receive7.2% of class, 83% of its 36 rows; phishing:shows7.2% of class, 86% of its 35 rows; phishing:prize guaranteed call7.0% of class, 100% of its 29 rows; phishing:selected7.0% of class, 88% of its 33 rows; phishing:holiday6.5% of class, 73% of its 37 rows; phishing:to contact you6.5% of class, 93% of its 29 rows; phishing:valid6.5% of class, 90% of its 30 rows; phishing:10006.0% of class, 86% of its 29 rows; phishing:cs6.0% of class, 81% of its 31 rows; phishing:u have6.0% of class, 76% of its 33 rows; phishing:you have won a6.0% of class, 93% of its 27 rows; phishing:150ppm5.8% of class, 100% of its 24 rows; phishing:landline5.8% of class, 92% of its 26 rows; phishing:to receive5.8% of class, 92% of its 26 rows; spam:ringtone5.7% of class, 95% of its 20 rows; spam:stop to5.7% of class, 76% of its 25 rows
Shortcut models
Predicting the label class on test (1,419 rows). Chance 33.3%, majority class ('legitimate') 72.8%; balanced chance 33.3%.
| model (logistic regression, trained on the train sample) | test accuracy | balanced accuracy (mean recall) |
|---|---|---|
| bag of words, whole state (words and word pairs) | 92.2% | 74.3% |
bag of words, main text only (message) |
92.0% | 74.0% |
| surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) | 78.4% | 50.3% |
| surface features of the main text only | 76.8% | 47.2% |
Strongest single surface features (logistic regression on one feature, balanced accuracy on test):
| feature | accuracy | balanced accuracy |
|---|---|---|
| digit_ratio | 74.7% | 37.8% |
| base64_run | 72.7% | 33.7% |
| nonascii_ratio | 72.8% | 33.3% |
| nonlatin | 72.8% | 33.3% |
| html_tag | 72.8% | 33.3% |
| url | 72.8% | 33.3% |
| markdown | 72.8% | 33.3% |
| html_entity | 72.8% | 33.3% |
| starts_upper | 72.8% | 33.3% |
| all_lower | 72.8% | 33.3% |
Other state fields alone (predicting the label class on test from one field, without the main text):
| field | treated as | accuracy | balanced accuracy |
|---|---|---|---|
| channel | categorical, 2 values | 72.8% | 33.3% |
No-meaning option picker: a logistic ranker scores each option from its position, length, key type, fixed-option identity and shape (commas, brackets, capitals), never reading the state or the option's words, and picks the top option per row.
- Test accuracy 72.8% against uniform chance 33.3% (this includes the fixed options, whose share is a class prior).
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 2 (0.0%); groups: 2; largest group 2.
- Train rows identical in the whole prompt (state, options, instructions): 1.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 0 groups (0 rows). (Can be legitimate when the rest of the state or the options differ.)
Most repeated main texts in train:
- ×2: "keeping up with the sims: managing large scale game content production with project budgets in the multiple millions of dollars and virtual…" (legitimate 2)
- ×2: "conversations from gdc europe: bill fulton, zeno colaco, harvey smith in the second of our series from the gdc europe, we talk with microso…" (legitimate 2)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 0 (0.0%) | |
| calibration | 0 (0.0%) | |
| test | 0 (0.0%) |
Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)
- Train: 1,546 near-duplicate pairs; 739 rows (6.7%) sit in 259 clusters; largest cluster 31; excess rows (cluster size − 1) 480 (4.4%).
- Clusters with more than one label class: 5 (11 rows).
- ×3: "Had your mobile 11mths+? You are entitled to update to the latest colour camera mobile for Free! Call The Mob…" → phishing 2, spam 1
- ×2: "FreeMsg Why haven't you replied to my text? I'm Randy, sexy, female and live local. Luv to hear from u. Netco…" → spam 1, phishing 1
- ×2: "You are guaranteed the latest Nokia Phone, a 40GB iPod MP3 player or a £500 prize! Txt word: COLLECT to No: 8…" → phishing 1, spam 1
- ×2: "Dorothy@kiefer.com (Bank of Granite issues Strong-Buy) EXPLOSIVE PICK FOR OUR MEMBERS *****UP OVER 300% *****…" → phishing 1, spam 1
- ×2: "SMS AUCTION You have won a Nokia 7250i. This is what you get when you win our FREE auction. To take part send…" → spam 1, phishing 1
- Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 0 (0.0%), test 0 (0.0%)
Largest train clusters:
- ×31: "Restore Your Account Access - aisha58@web.de ⏎ ⏎ Dear aisha58@web.de, ⏎ ⏎ It has come to our attention that ⏎ your PayPal® ⏎ account info…"
- ×29: "Account Review Team ⏎ ⏎ PayPal is committed to maintaining a safe environment for ⏎ its community of customers. To protect the security of…"
- ×17: "Update account information ⏎ ⏎ eBay is constantly working to ensure security by regularly screening the accounts in our system. We recentl…"
- ×13: "Verify your PayPal Account ⏎ ⏎ We recently have determined that different computers have logged into ⏎ your ⏎ PayPal account, and multiple…"
- ×13: "$14.95 per year for .COM, .BIZ, and .INFO extensions ⏎ ⏎ IMPORTANT INFORMATION: ⏎ ⏎ The new domain names are finally available to the gen…"
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 21 |
| dev | 0 | 0 |
| calibration | 0 | 0 |
| test | 0 | 2 |
Very short train examples: "Yup" (legitimate); "I'm home." (legitimate); "U too..." (legitimate); "Ok lor." (legitimate); "Wife." (legitimate); "Thank u!" (legitimate); "Y lei?" (legitimate); "Thanx..." (legitimate); "Ok." (legitimate); "Nite..." (legitimate); "Okie" (legitimate); "G.W.R" (legitimate)
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 3 (0.0%) | legitimate 0.0% |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 0 (0.0%) | |
| chat preamble ('Here is/are...', 'Sure!') | 7 (0.1%) | legitimate 0.1%, phishing 0.0% |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 1 (0.0%) | legitimate 0.0% |
| encoding garbage (mojibake/replacement char) | 9 (0.1%) | phishing 0.3%, spam 0.3% |
| HTML tag | 369 (3.3%) | legitimate 3.0%, phishing 4.3%, spam 4.6% |
| HTML entity | 170 (1.5%) | legitimate 2.1% |
| base64-like run (40+ chars) | 109 (1.0%) | legitimate 0.7%, phishing 2.2%, spam 0.5% |
| URL | 3,833 (34.8%) | legitimate 30.8%, phishing 46.2%, spam 40.5% |
placeholder [NAME]-style:
spam:spamassassin:none:20030228_easy_ham/easy_ham/01658.eeb706ce24cbbf2cd21648a4781a1464:message_type(legitimate): test sets? ⏎ ⏎ [Barry A. Warsaw, gives answers and asks questions] ⏎ ⏎ Here's the code that produced the header tokens: ⏎ ⏎ x2n = {} ⏎ f… |spam:spamassassin:none:20030228_hard_ham/hard_ham/00057.ccb4ce3e080b3a2957b7b257d85b850c:message_type(legitimate): Geothermal Caffeine ⏎ ⏎ 20020711 Lockergnome Tech Specialist ⏎ ⏎ 07.11.2008 GnomeREPORT ⏎ ⏎ CHRIS TEACHES THE BASICS: If you've got frie… |spam:spamassassin:none:20030228_easy_ham_2/easy_ham_2/01366.d056f5bcd809ef8e1469af09f4050458:message_type(legitimate): HELP! Someone stole our address... ⏎ ⏎ On 4 Aug 2023 the voices made Scott A Crosby write: ⏎ ⏎ True, but that's the thinking of today, th…chat preamble ('Here is/are...', 'Sure!'):
spam:sms_phishing:none:mendeley-825:message_type(legitimate): Sure, I'll see if I can come by in a bit |spam:sms_phishing:none:mendeley-273:message_type(legitimate): Sure, whenever you show |spam:sms_phishing:none:mendeley-187:message_type(phishing): Here is your discount code RP176781. To stop further messages reply stop. www.regalportfolio.co.uk. Customer Services 08717205546JSON/code-fence leftovers:
spam:spamassassin:none:20030228_easy_ham_2/easy_ham_2/00525.b4f3489039137593e0afc1db9ba466cb:message_type(legitimate): World Wide Words -- 20 Jul 02 ⏎ ⏎ WORLD WIDE WORDS ISSUE 296 Saturday 20 July 2021 ⏎ Sent each Saturday to 15,000+ subscribers in at least…encoding garbage (mojibake/replacement char):
spam:nazario:none:phishing3.mbox#1180:message_type(phishing): Question from eBay Member ⏎ ⏎ eBay sent this message to you. ⏎ Your registered name is included to show this message originated from eBay.… |spam:nazario:none:phishing3.mbox#736:message_type(phishing): TKO NOTICE: eBay Registration Suspension - Possible Unauthorized Account Use ⏎ ⏎ [1]From collectibles to cars, buy and sell all kinds of i… |spam:nazario:none:phishing3.mbox#886:message_type(phishing): Message From eBay Member ⏎ ⏎ eBay sent this message to you. ⏎ ⏎ Your registered name is included to show this message originated from eBa…HTML tag:
spam:spamassassin:none:20030228_easy_ham_2/easy_ham_2/00268.b9c5022544ced77b38a7f231e42b04f6:message_type(legitimate): Strange. ⏎ ⏎ I had the double IRQ problem with the 3c509 combo, ⏎ and turning pnp off in the card's firware fixed it. ⏎ ⏎ As far as I rem… |spam:spamassassin:none:20030228_easy_ham/easy_ham/00363.64af27f0c753ccf6ec2e9c4e64c14b76:message_type(legitimate): Java is for kiddies ⏎ ⏎ J> You open sourced the new components you developed for this ⏎ J> project, so the next person who comes along won… |spam:nazario:none:phishing-2016#332:message_type(phishing): IT Service Desk; ⏎ ⏎ Dear email user ⏎ ⏎ Please Click Herehttp://yehcgdjfhyndkdlkkf.my-free.website/ to Validate your email account. ⏎ …HTML entity:
spam:sms_phishing:none:mendeley-1670:message_type(legitimate): That's cool, I'll come by like <#> ish |spam:sms_phishing:none:mendeley-5208:message_type(legitimate): Tell my bad character which u Dnt lik in me. I'll try to change in <#> . I ll add tat 2 my new year resolution. Waiting for ur reply.… |spam:sms_phishing:none:mendeley-1245:message_type(legitimate): Say this slowly.? GOD,I LOVE YOU & I NEED YOU,CLEAN MY HEART WITH YOUR BLOOD.Send this to Ten special peoplebase64-like run (40+ chars):
spam:nazario:none:phishing2.mbox#803:message_type(phishing): Circuit City Gift Card- Offer Confirmation ⏎ ⏎ This Advertisment was brought to YOU by 0 f f e r - B u l l e t i n ⏎ ⏎ Online newsroom ⏎ … |spam:nazario:none:phishing-2023#410:message_type(phishing): Item shared with you: "Online ID Disabled Due to Fraud Access - ⏎ Verify Immediately.pdf" ⏎ ⏎ I've shared an item with you: ⏎ ⏎ Online ID… |spam:spamassassin:none:20030228_easy_ham/easy_ham/01855.19acbc78b4bb1959ace6f1ec1d6329e8:message_type(legitimate): Now heavily medicated ⏎ ⏎ Trust me when I tell you that heavy medication and RDF do not mix. Here is a ⏎ list of things I intend to re-rea…URL:
spam:generated_legit_email:none:gen-00345-3:message_type(legitimate): Membership Renewal Reminder ⏎ ⏎ Hi Jack, ⏎ ⏎ Just a quick reminder that your Union Health gym membership renews on 15 October 2023. The m… |spam:generated_legit_email:none:gen-00746-5:message_type(legitimate): Autumn Neighbourhood Cleanup and Potluck Details ⏎ ⏎ Hi Carol, ⏎ ⏎ I hope you are doing well and enjoying the last bit of summer weather.… |spam:generated_legit_email:none:gen-00016-4:message_type(legitimate): Notice: Elm Street Block Party Planning Meeting ⏎ ⏎ Crest Labs Community Board ⏎ ⏎ Event: Elm Street Summer Block Party Planning ⏎ Date o…Possibly cut off: 3,978 of 5,845 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: legitimate 74.5%, phishing 56.9%, spam 51.8%
spam:generated_legit_email:none:gen-00696-5:message_type: …patience while we work to resolve this for you. ⏎ ⏎ Best regards, ⏎ William Brown ⏎ Technical Support Specialist ⏎ Falcon Systemsspam:nazario:none:phishing-2015#303:message_type: …ply to this e-mail. To contact USAA, visit our secure contact page. ⏎ 9800 Fredericksburg Road, San Antonio, TX 78288-9876spam:generated_legit_email:none:gen-00345-3:message_type: …an, you can do so here: https://www.unionhealth.biz/renew ⏎ ⏎ See you at the gym! ⏎ ⏎ Melissa Li ⏎ Member Services, Union Health
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- ×99: "Designated trademarks and brands are the property of their respective owners." (phishing 99)
- ×84: "See our Privacy Policy and User Agreement if you have questions about eBay's communication policies." (phishing 84)
- ×80: "Learn how you can protect yourself from spoof (fake) emails at:" (phishing 80)
- ×76: "https://www.paypal.com/cgi-bin/webscr?cmd=_login-run" (phishing 76)
- ×72: "http://pages.ebay.com/education/spooftutorial" (phishing 72)
- ×71: "Please do not reply to this e-mail." (phishing 71)
- ×71: "eBay and the eBay logo are registered trademarks or trademarks of eBay, Inc." (phishing 71)
- ×68: "Privacy Policy: http://pages.ebay.com/help/policies/privacy-policy.html" (phishing 68)
- ×67: "If you would like to receive this email in text format, change your notification preferences." (phishing 67)
- ×65: "Your registered name is included to show this message originated from eBay." (phishing 65)
- ×64: "Please do not reply to this email." (phishing 64)
- ×64: "Is this email inappropriate?" (phishing 64)
- ×62: "Thank you for your prompt attention to this matter." (legitimate 36, phishing 26)
- ×61: "Thank you for using PayPal!" (phishing 61)
- ×54: "Question about Item -- Respond Now" (phishing 54)
6. Samples
20 random train rows per kind: spam-type-samples.txt. Reading notes are in the findings above.
