QA report: ground
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.
QA: ground (v4)
Verdict: PASS WITH NOTES. See ground-grounding.md (verdict and notes) and ground-rerank.md. QA-OK-ground written.
QA: ground-grounding
Checked 2026-09-30 13:26 by adapters/qa/qa.py (READY file READY-ground, v4 (held-out second opinion; grounding answers balanced on surface features, 2026-09-30)).
Verdict: PASS WITH NOTES
Ground v4 (READY-ground 2026-09-30 13:19, train sha256 c8b38f85…15fb2; train 50,760 / dev 1,800 / calibration 1,440 / test 4,160). The v3 check is in ground-v3-summary.md. This report covers both question types: grounding (ground-grounding.md) and re-ranking (ground-rerank.md).
The v3 finding is fixed (checked with ground_extra.py and qa.py)
- Answer length and shape are equal across labels:
- median answer is 28–29 words for every label in train, and 32 for every label in test;
- ", " appears in 68.2% of answers and " and " in 41.5%, identical for every label (test: 79.8% and 45.2%).
- Answer-only models are now near chance:
- surface-only model: 28.4% balanced accuracy (was 35.7%; chance 25%), and 37.1% accuracy against a 35% majority baseline;
- answer bag of words: 33.7% (was 40.7%);
- whole state (answer plus sources): 28.4%.
- The standard phrase check finds no flags for either question type.
- Near duplicates across splits: 0–0.1% of held-out rows.
Notes (for the model card; meaning or small)
Remaining answer-word signal (bag of words 33.7% against 25% chance):
- negation ("not") appears in 8.0% of partly-supported answers against 15–16% of the others;
- digits appear in 60% of contradicted answers against 47% of supported ones (number-changed rows);
- hedged added details appear in partly-supported answers.
These follow from how each label is made; the grounding agent tried adding negation and hedge cells, and it did not help. No single cue meets the 2%-and-mostly-one-label rule.
Sources are about 20% shorter for partly-supported and unsupported rows (median 1,511–1,533 characters, against 1,834–1,899), because a gold passage is removed. The length of the sources alone gives 27.3% balanced accuracy, near chance.
Re-ranking is unchanged and passes. The gold passage is the longest 7.8% of the time against 9.1% chance, its position is uniform, "none" rows have the same option count as the others, and a surface-only model reaches 7.8% balanced against 4.2% chance (from option count). 7,977 train rows are second pools (accepted earlier).
The grounding agent labelled "supported + a claim whose passage was removed" as partly_supported, which is right by the label definitions; I had proposed labelling it unsupported.
Share-alike licences (SQuAD, HotpotQA, FEVER) and hosted-teacher text are covered in the data report.
Automatic flags (for the reviewer to judge; not all are problems)
- 146 strong phrase flags (see list): review whether they are meaning or leakage
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 25,380 | 4250 | train.jsonl |
| dev | 900 | 128 | dev.jsonl |
| calibration | 720 | 94 | calibration.jsonl |
| test | 2,080 | 310 | test.jsonl |
- Train sha256:
c8b38f85a728f5aa916c2b4b2cf18a8cd63567fde55572cc4b0f16b155915fb2(READY saysc8b38f85a728f5aa916c2b4b2cf18a8cd63567fde55572cc4b0f16b155915fb2: match) - Main text field (the text the phrase and length checks use):
state.answer. - Label classes: contradicted, partly_supported, supported, unsupported (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): grounding-docs-contradicted, grounding-docs-gold-removed, grounding-docs-joined-contradicted, grounding-docs-joined-supported, grounding-docs-joined-unsupported, grounding-docs-number-changed, grounding-docs-partly, grounding-docs-supported, grounding-docs-unanswerable, grounding-fever-combined, grounding-fever-contradicted, grounding-fever-joined-contradicted, grounding-fever-joined-supported, grounding-fever-joined-unsupported, grounding-fever-supported, grounding-fever-unsupported, grounding-hotpot-contradicted, grounding-hotpot-gold-removed, grounding-hotpot-number-changed, grounding-hotpot-one-removed, grounding-hotpot-partly, grounding-hotpot-supported, grounding-squad-combined, grounding-squad-contradicted, grounding-squad-gold-removed, grounding-squad-joined-contradicted, grounding-squad-joined-supported, grounding-squad-joined-unsupported, grounding-squad-number-changed, grounding-squad-partly, grounding-squad-supported, grounding-squad-unanswerable.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| contradicted | 20.0% | 20.0% | 20.0% | 20.0% | 5,076 |
| partly_supported | 25.0% | 25.0% | 25.0% | 25.0% | 6,345 |
| supported | 35.0% | 35.0% | 35.0% | 35.0% | 8,883 |
| unsupported | 20.0% | 20.0% | 20.0% | 20.0% | 5,076 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| grounding-docs-contradicted | 3.7% | 1.9% | 2.1% | 3.0% | 932 |
| grounding-docs-gold-removed | 3.8% | 3.9% | 2.9% | 4.6% | 965 |
| grounding-docs-joined-contradicted | 2.3% | 3.6% | 3.3% | 3.3% | 579 |
| grounding-docs-joined-supported | 5.2% | 4.8% | 5.6% | 6.2% | 1,328 |
| grounding-docs-joined-unsupported | 2.8% | 2.4% | 3.1% | 3.2% | 707 |
| grounding-docs-number-changed | 1.8% | 1.9% | 2.1% | 3.0% | 452 |
| grounding-docs-partly | 5.8% | 5.8% | 6.1% | 7.5% | 1,470 |
| grounding-docs-supported | 7.7% | 7.9% | 5.8% | 8.4% | 1,959 |
| grounding-docs-unanswerable | 1.7% | 1.1% | 1.7% | 1.9% | 432 |
| grounding-fever-combined | 5.0% | 2.8% | 2.8% | 2.9% | 1,260 |
| grounding-fever-contradicted | 2.5% | 1.3% | 1.2% | 1.1% | 626 |
| grounding-fever-joined-contradicted | 1.5% | 0.9% | 1.0% | 1.2% | 382 |
| grounding-fever-joined-supported | 3.0% | 2.1% | 2.6% | 2.3% | 772 |
| grounding-fever-joined-unsupported | 1.7% | 1.7% | 1.7% | 1.1% | 429 |
| grounding-fever-supported | 3.9% | 1.8% | 1.2% | 1.8% | 992 |
| grounding-fever-unsupported | 2.3% | 0.6% | 0.6% | 1.2% | 579 |
| grounding-hotpot-contradicted | 1.5% | 0.2% | 0.3% | 0.6% | 379 |
| grounding-hotpot-gold-removed | 2.1% | 0.2% | 0.4% | 0.3% | 541 |
| grounding-hotpot-number-changed | 1.2% | 0.3% | 0.6% | 0.6% | 307 |
| grounding-hotpot-one-removed | 1.7% | 0.6% | 0.3% | 0.7% | 434 |
| grounding-hotpot-partly | 3.3% | 0.6% | 0.7% | 0.9% | 839 |
| grounding-hotpot-supported | 4.2% | 1.3% | 0.7% | 1.8% | 1,074 |
| grounding-squad-combined | 3.5% | 5.7% | 4.7% | 4.1% | 900 |
| grounding-squad-contradicted | 2.6% | 3.3% | 2.9% | 2.4% | 672 |
| grounding-squad-gold-removed | 3.0% | 6.2% | 4.2% | 4.2% | 767 |
| grounding-squad-joined-contradicted | 1.7% | 2.7% | 3.5% | 2.3% | 424 |
| grounding-squad-joined-supported | 4.3% | 7.3% | 8.3% | 6.3% | 1,097 |
| grounding-squad-joined-unsupported | 2.4% | 3.8% | 5.6% | 3.3% | 605 |
| grounding-squad-number-changed | 1.3% | 3.9% | 3.1% | 2.5% | 323 |
| grounding-squad-partly | 5.7% | 9.7% | 10.4% | 9.0% | 1,442 |
| grounding-squad-supported | 6.5% | 9.8% | 10.7% | 8.3% | 1,661 |
| grounding-squad-unanswerable | 0.2% | 0.1% | 0.0% | 0.2% | 51 |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Options per choice row
| split | min | median | p99 | max |
|---|---|---|---|---|
| train | 4 | 4 | 4 | 4 |
| dev | 4 | 4 | 4 | 4 |
| calibration | 4 | 4 | 4 | 4 |
| test | 4 | 4 | 4 | 4 |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | source.input_tokens (25380/25380 rows) | 593 | 1017 | 1301 | 0 |
| dev | source.input_tokens (900/900 rows) | 576 | 972 | 1140 | 0 |
| calibration | source.input_tokens (720/720 rows) | 593 | 991 | 1045 | 0 |
| test | source.input_tokens (2080/2080 rows) | 596 | 1026 | 1219 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| answer, sources | 25,380 (100.0%) |
Instructions (train)
- Canonical (the most common text) 69.6%, reworded 27.5% (80 distinct rewordings), none 2.9%. Target about 70 / 27 / 3.
- Canonical text: "Is the answer supported by the sources? Judge only from the sources, not from outside knowledge."
sourceinstruction tag: canonical 69.6%, variant 27.5%, none 2.9%
| class | canonical | none |
|---|---|---|
| contradicted | 68.7% | 2.8% |
| partly_supported | 70.6% | 3.1% |
| supported | 69.3% | 3.0% |
| unsupported | 69.8% | 2.7% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 25,380 train rows; models are scored on the full test file (2,080 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | contradicted | 5076 | 74 | 176 | 332 | 191 |
| train | partly_supported | 6345 | 81 | 177 | 311 | 188 |
| train | supported | 8883 | 74 | 178 | 343 | 194 |
| train | unsupported | 5076 | 76 | 178 | 342 | 194 |
| test | contradicted | 416 | 96 | 200 | 360 | 214 |
| test | partly_supported | 520 | 99 | 199 | 316 | 208 |
| test | supported | 728 | 96 | 202 | 364 | 217 |
| test | unsupported | 416 | 97 | 202 | 370 | 218 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| grounding-docs-contradicted | 932 | 178 | 187 | 2175 | 4 |
| grounding-docs-gold-removed | 965 | 173 | 178 | 2138 | 4 |
| grounding-docs-joined-contradicted | 579 | 327 | 336 | 2549 | 4 |
| grounding-docs-joined-supported | 1328 | 297 | 314 | 2734 | 4 |
| grounding-docs-joined-unsupported | 707 | 300 | 316 | 2210 | 4 |
| grounding-docs-number-changed | 452 | 190 | 197 | 2144 | 4 |
| grounding-docs-partly | 1470 | 184 | 203 | 2160 | 4 |
| grounding-docs-supported | 1959 | 172 | 179 | 2164 | 4 |
| grounding-docs-unanswerable | 432 | 200 | 208 | 2111 | 4 |
| grounding-fever-combined | 1260 | 81 | 85 | 1523 | 4 |
| grounding-fever-contradicted | 626 | 66 | 65 | 1319 | 4 |
| grounding-fever-joined-contradicted | 382 | 93 | 97 | 1141 | 4 |
| grounding-fever-joined-supported | 772 | 91 | 94 | 1158 | 4 |
| grounding-fever-joined-unsupported | 429 | 93 | 97 | 1093 | 4 |
| grounding-fever-supported | 992 | 66 | 66 | 1388 | 4 |
| grounding-fever-unsupported | 579 | 67 | 66 | 1333 | 4 |
| grounding-hotpot-contradicted | 379 | 160 | 164 | 1460 | 4 |
| grounding-hotpot-gold-removed | 541 | 161 | 162 | 1473 | 4 |
| grounding-hotpot-number-changed | 307 | 160 | 165 | 1470 | 4 |
| grounding-hotpot-one-removed | 434 | 137 | 145 | 849 | 4 |
| grounding-hotpot-partly | 839 | 184 | 190 | 1484 | 4 |
| grounding-hotpot-supported | 1074 | 157 | 160 | 1470 | 4 |
| grounding-squad-combined | 900 | 307 | 302 | 1411 | 4 |
| grounding-squad-contradicted | 672 | 183 | 183 | 1514 | 4 |
| grounding-squad-gold-removed | 767 | 180 | 179 | 1424 | 4 |
| grounding-squad-joined-contradicted | 424 | 327 | 331 | 1850 | 4 |
| grounding-squad-joined-supported | 1097 | 315 | 317 | 1880 | 4 |
| grounding-squad-joined-unsupported | 605 | 312 | 316 | 1378 | 4 |
| grounding-squad-number-changed | 323 | 178 | 180 | 1544 | 4 |
| grounding-squad-partly | 1442 | 202 | 203 | 1441 | 4 |
| grounding-squad-supported | 1661 | 179 | 179 | 1488 | 4 |
| grounding-squad-unanswerable | 51 | 129 | 126 | 1408 | 4 |
Correct option: longest / shortest / position / key
For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):
| split | rows | correct is longest | correct is shortest | chance (1/listed) | mean relative position (0 first, 1 last; 0.5 expected) | position fifths |
|---|
Correct key and position by option count
| split | options | rows | mean options | top correct keys | most common position (0-based) |
|---|---|---|---|---|---|
| train | 2-5 | 25380 | 4.0 | supported 35.0%, partly_supported 25.0%, contradicted 20.0%, unsupported 20.0% | 2 (25.1%) |
| test | 2-5 | 2080 | 4.0 | supported 35.0%, partly_supported 25.0%, unsupported 20.0%, contradicted 20.0% | 1 (25.5%) |
Option count by label class (train)
| class | rows | min | median | mean | max |
|---|---|---|---|---|---|
| contradicted | 5076 | 4 | 4 | 4.0 | 4 |
| partly_supported | 6345 | 4 | 4 | 4.0 | 4 |
| supported | 8883 | 4 | 4 | 4.0 | 4 |
| unsupported | 5076 | 4 | 4 | 4.0 | 4 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 35.0%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| kind | 32 | 100.0% | grounding-docs-supported: supported 100.0%; grounding-squad-supported: supported 100.0%; grounding-docs-partly: partly_supported 100.0%; grounding-squad-partly: partly_supported 100.0%; grounding-docs-joined-supported: supported 100.0%; grounding-fever-combined: partly_supported 100.0% |
| second_opinion | 2 | 48.3% | null: supported 45.1%, partly_supported 32.2%; deepseek-v4-flash: contradicted 59.2%, unsupported 40.8% |
| dataset | 4 | 35.8% | teacher-docs: supported 37.3%, unsupported 23.8%; squad2: supported 34.7%, partly_supported 29.5%; fever: supported 35.0%, partly_supported 25.0%; hotpotqa: partly_supported 35.6%, supported 30.1% |
| url | 4 | 35.8% | : supported 37.3%, unsupported 23.8%; https://huggingface.co/datasets/rajpurkar/squad_v2: supported 34.7%, partly_supported 29.5%; https://fever.ai/dataset/fever.html: supported 35.0%, partly_supported 25.0%; https://huggingface.co/datasets/hotpotqa/hotpot_qa: partly_supported 35.6%, supported 30.1% |
| revision | 4 | 35.8% | : supported 37.3%, unsupported 23.8%; 3ffb306f725f7d2ce8394bc1873b24868140c412: supported 34.7%, partly_supported 29.5%; sha256 train.jsonl eba7e8f8..., shared_task_dev.jsonl and wiki-pages.zip 4b06d95d... (see raw/manifest.json): supported 35.0%, partly_supported 25.0%; 1908d6afbbead072334abe2965f91bd2709910ab: partly_supported 35.6%, supported 30.1% |
| check_teacher | 2 | 35.2% | hosted: qwen3.8-flash: supported 35.2%, partly_supported 24.7%; local: qwen3.8-flash-next: partly_supported 38.2%, supported 26.4% |
| text_teacher | 4 | 35.1% | hosted: qwen3.8-max: supported 34.3%, partly_supported 26.4%; local: qwen3.8-flash-next: supported 36.9%, unsupported 24.8%; none (public data and code): supported 35.0%, partly_supported 25.0%; mixed: hosted: qwen3.8-max + local: qwen3.8-flash-next: partly_supported 37.5%, supported 34.4% |
| license | 3 | 35.0% | CC BY-SA 4.0: supported 33.3%, partly_supported 31.4%; generated by the local teacher (qwen3.8-flash-next): supported 37.3%, unsupported 23.8%; CC BY-SA 3.0: supported 35.0%, partly_supported 25.0% |
| source_style | 3 | 35.0% | bracket: supported 34.8%, partly_supported 25.2%; document: supported 35.5%, partly_supported 24.8%; source: supported 34.8%, partly_supported 25.0% |
Formatting by label class (main text, share of rows)
| feature | contradicted | partly_supported | supported | unsupported | |
|---|---|---|---|---|---|
| ends with ? | 0.0% | 0.0% | 0.0% | 0.0% | |
| ends with . | 99.6% | 99.5% | 99.6% | 99.6% | |
| ends with ! | 0.0% | 0.0% | 0.0% | 0.0% | |
| no end punctuation | 0.1% | 0.0% | 0.0% | 0.0% | |
| starts lowercase | 0.0% | 0.0% | 0.0% | 0.0% | |
| all lowercase | 0.0% | 0.0% | 0.0% | 0.0% | |
| has a digit | 59.9% | 57.5% | 47.1% | 46.5% | |
| has newline | 0.0% | 0.0% | 0.0% | 0.0% | |
| has quotes | 4.6% | 6.8% | 4.7% | 4.4% | |
| has markup (HTML/markdown) | 0.0% | 0.0% | 0.0% | 0.0% | |
| has URL | 0.0% | 0.0% | 0.0% | 0.0% | |
| non-ASCII | 4.3% | 5.5% | 4.5% | 4.3% | |
| non-Latin script | 0.0% | 0.0% | 0.0% | 0.0% | |
| emoji | 0.0% | 0.0% | 0.0% | 0.0% | |
| ALL-CAPS word (4+) | 2.9% | 2.7% | 2.8% | 2.9% | |
| contains ' - ' or — | 0.1% | 0.1% | 0.1% | 0.1% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, contradicted: yes z=7 6.0% vs 3.7%; did z=6 1.8% vs 0.8%; if z=5 12.5% vs 9.6%; incapable z=5 0.9% vs 0.0%; not z=5 9.5% vs 7.1%; only z=5 6.1% vs 4.2%; ever z=5 0.7% vs 0.3%; unable z=5 0.4% vs 0.1%; 60 z=4 1.4% vs 0.7%; 1 z=4 5.0% vs 3.7%; exactly z=4 2.7% vs 1.8%; yet z=4 0.3% vs 0.1%; refused z=4 0.3% vs 0.1%; you z=4 22.7% vs 19.6%; zero z=4 0.9% vs 0.5%; 8 z=4 2.0% vs 1.3%; avoided z=4 0.2% vs 0.0%; any z=4 5.6% vs 4.3%; 17 z=4 1.0% vs 0.6%; except z=4 0.4% vs 0.1%
Words, partly_supported: which z=15 15.8% vs 9.1%; in z=14 57.0% vs 46.0%; was z=14 32.9% vs 24.3%; born z=11 4.5% vs 1.9%; over z=10 5.4% vs 2.8%; he z=9 9.3% vs 6.1%; who z=9 8.5% vs 5.5%; provided z=9 3.6% vs 1.7%; his z=9 6.8% vs 4.2%; university z=8 2.6% vs 1.2%; particularly z=8 0.9% vs 0.1%; million z=8 2.8% vs 1.3%; roughly z=8 0.9% vs 0.1%; founded z=8 2.2% vs 1.0%; primarily z=8 1.2% vs 0.3%; having z=7 1.1% vs 0.3%; filmed z=7 0.7% vs 0.1%; originally z=7 1.7% vs 0.7%; with z=7 15.0% vs 12.1%; near z=7 1.5% vs 0.6%
Words, supported: no z=8 8.9% vs 6.1%; if z=5 11.7% vs 9.4%; or z=4 10.6% vs 8.8%; cannot z=4 1.9% vs 1.2%; must z=4 13.4% vs 11.5%; date z=4 2.8% vs 2.0%; twelve z=3 1.9% vs 1.4%; are z=3 17.2% vs 15.3%; within z=3 9.4% vs 8.1%; you z=3 21.6% vs 19.5%; fourteen z=3 1.4% vs 0.9%; prohibited z=3 1.2% vs 0.8%; four z=3 3.6% vs 2.9%; strictly z=3 1.5% vs 1.1%; maximum z=3 2.2% vs 1.7%; forty z=3 1.7% vs 1.3%; be z=3 11.5% vs 10.2%; winner z=3 0.2% vs 0.1%; ltd z=3 0.5% vs 0.2%; product z=3 0.9% vs 0.6%
Words, unsupported: yes z=10 6.9% vs 3.5%; automatically z=6 2.8% vs 1.5%; your z=6 12.4% vs 9.3%; parking z=5 0.5% vs 0.1%; app z=5 0.9% vs 0.3%; discount z=5 0.7% vs 0.2%; available z=5 1.7% vs 0.9%; supports z=5 0.6% vs 0.2%; no z=5 8.7% vs 6.7%; mobile z=4 0.6% vs 0.2%; ios z=4 0.3% vs 0.0%; all z=4 6.0% vs 4.5%; android z=4 0.4% vs 0.1%; includes z=4 1.8% vs 1.1%; can z=4 7.4% vs 5.8%; strictly z=4 1.8% vs 1.1%; without z=4 3.0% vs 2.1%; you z=4 22.7% vs 19.6%; fury z=4 0.2% vs 0.0%; device z=4 1.4% vs 0.8%
2–4-word phrases, contradicted: not a z=6 0.6% vs 0.1%; did not z=6 1.5% vs 0.5%; is not a z=6 0.5% vs 0.0%; was not z=5 0.8% vs 0.2%; incapable of z=5 0.9% vs 0.0%; of being z=5 0.4% vs 0.1%; was only z=5 0.3% vs 0.0%; has only z=5 0.3% vs 0.0%; only ever z=4 0.3% vs 0.0%; long as z=4 0.6% vs 0.2%; as long as z=4 0.6% vs 0.2%; as long z=4 0.6% vs 0.2%; incapable of being z=4 0.4% vs 0.0%; unable to z=4 0.4% vs 0.1%; was incapable z=4 0.4% vs 0.0%; was incapable of z=4 0.4% vs 0.0%; 60 days z=4 0.4% vs 0.1%; not an z=4 0.3% vs 0.0%; long as they z=4 0.2% vs 0.0%; as long as they z=4 0.2% vs 0.0%
2–4-word phrases, partly_supported: and you z=15 2.7% vs 0.5%; was born z=12 3.7% vs 1.5%; which was z=12 2.9% vs 1.1%; is a z=11 8.4% vs 5.8%; and you must z=11 1.2% vs 0.1%; born in z=10 2.1% vs 0.8%; in the z=9 16.7% vs 14.3%; provided the z=9 1.0% vs 0.2%; was born in z=9 1.7% vs 0.6%; and it z=9 1.3% vs 0.4%; provided you z=8 1.2% vs 0.4%; in his z=8 1.1% vs 0.3%; was a z=8 3.1% vs 1.8%; is an z=8 3.0% vs 1.7%; and you will z=8 0.7% vs 0.1%; was born on z=8 1.5% vs 0.6%; he was z=8 2.8% vs 1.6%; founded in z=8 1.1% vs 0.3%; is the z=7 6.0% vs 4.5%; born on z=7 1.7% vs 0.8%
2–4-word phrases, supported: is a film z=4 0.5% vs 0.2%; if you z=4 5.5% vs 4.3%; you must z=4 8.0% vs 6.7%; no you z=4 0.9% vs 0.5%; a film z=4 0.8% vs 0.4%; academy awards z=3 0.2% vs 0.1%; no the z=3 0.9% vs 0.6%; are strictly z=3 0.8% vs 0.5%; are not z=3 1.2% vs 0.8%; maximum of z=3 1.1% vs 0.7%; a maximum of z=3 1.1% vs 0.7%; there is z=3 1.3% vs 0.9%; strictly prohibited z=3 0.8% vs 0.5%; are strictly prohibited z=3 0.6% vs 0.3%; proceed to z=3 0.2% vs 0.0%; you cannot z=3 0.7% vs 0.4%; at least one z=3 0.4% vs 0.2%; least one z=3 0.4% vs 0.2%; a maximum z=3 1.5% vs 1.1%; if the z=3 2.6% vs 2.0%
2–4-word phrases, unsupported: yes the z=7 1.6% vs 0.6%; discount on z=5 0.3% vs 0.0%; a dedicated z=5 0.5% vs 0.1%; mobile app z=4 0.3% vs 0.1%; percent discount z=4 0.3% vs 0.1%; available on z=4 0.3% vs 0.0%; stars an z=4 0.2% vs 0.0%; percent discount on z=4 0.2% vs 0.0%; yes there z=4 0.2% vs 0.0%; yes you z=4 1.3% vs 0.7%; gift of the night z=4 0.2% vs 0.0%; the night fury z=4 0.2% vs 0.0%; night fury z=4 0.2% vs 0.0%; gift of the z=4 0.2% vs 0.0%; gift of z=4 0.2% vs 0.0%; of the night fury z=4 0.2% vs 0.0%; yes you are z=4 0.2% vs 0.0%; the night fury stars z=4 0.2% vs 0.0%; fury stars z=4 0.2% vs 0.0%; night fury stars z=4 0.2% vs 0.0%
By row kind, 1–3-word phrases (top 10)
- grounding-docs-contradicted:
you48.0% vs 19.2%;if28.4% vs 9.5%;yes15.3% vs 3.8%;your24.0% vs 9.4%;must26.3% vs 11.6%;if you13.1% vs 4.4%;days17.2% vs 6.8%;within19.1% vs 8.2%;yes you3.6% vs 0.7%;pounds7.9% vs 2.6% - grounding-docs-gold-removed:
you46.7% vs 19.2%;if28.4% vs 9.5%;your24.7% vs 9.3%;must27.9% vs 11.5%;will15.5% vs 5.7%;no17.1% vs 6.7%;if you12.6% vs 4.4%;you must17.0% vs 6.8%;days16.1% vs 6.9%;hours12.0% vs 4.7% - grounding-docs-joined-contradicted:
must45.3% vs 11.4%;your38.2% vs 9.3%;yes19.7% vs 3.8%;you64.4% vs 19.2%;if37.5% vs 9.6%;you must27.8% vs 6.7%;will23.8% vs 5.7%;days26.6% vs 6.8%;percent19.2% vs 4.4%;any18.3% vs 4.2% - grounding-docs-joined-supported:
must44.8% vs 10.3%;you62.7% vs 17.9%;you must28.8% vs 6.0%;no28.2% vs 5.9%;days28.1% vs 6.1%;if35.7% vs 8.8%;your34.4% vs 8.6%;within30.5% vs 7.4%;hours19.6% vs 4.2%;percent17.4% vs 4.1% - grounding-docs-joined-unsupported:
days30.1% vs 6.6%;must43.7% vs 11.2%;you62.4% vs 19.0%;you must28.3% vs 6.6%;if36.6% vs 9.4%;your35.6% vs 9.2%;no27.7% vs 6.5%;within30.0% vs 7.9%;five19.1% vs 4.4%;percent18.7% vs 4.4% - grounding-docs-number-changed:
you49.6% vs 19.7%;if29.6% vs 9.8%;must33.0% vs 11.8%;you must21.2% vs 6.9%;362.2% vs 0.1%;009.5% vs 2.1%;if you15.0% vs 4.5%;113.3% vs 3.8%;753.3% vs 0.3%;seconds6.4% vs 1.3% - grounding-docs-partly:
and you11.8% vs 0.4%;you53.4% vs 18.2%;provided13.9% vs 1.5%;must36.2% vs 10.7%;you must22.1% vs 6.3%;within23.7% vs 7.6%;provided you5.4% vs 0.3%;and you must5.4% vs 0.1%;your25.3% vs 9.0%;days19.8% vs 6.4% - grounding-docs-supported:
you47.8% vs 17.9%;if26.7% vs 8.8%;must28.3% vs 10.8%;your23.7% vs 8.8%;no18.0% vs 6.2%;if you12.9% vs 4.0%;you must16.8% vs 6.4%;will14.5% vs 5.4%;within18.3% vs 7.7%;days15.8% vs 6.5% - grounding-docs-unanswerable:
yes44.9% vs 3.5%;yes the14.1% vs 0.5%;yes you9.3% vs 0.7%;can23.8% vs 5.8%;your30.6% vs 9.6%;app6.2% vs 0.3%;yes you can6.2% vs 0.3%;available8.8% vs 0.9%;parking5.1% vs 0.1%;you can14.4% vs 2.9% - grounding-fever-combined:
is a22.9% vs 5.6%;a55.0% vs 46.9%;is46.3% vs 39.7%;has17.4% vs 6.3%;was32.9% vs 26.1%;an18.6% vs 11.6%;in40.2% vs 49.2%;is an7.2% vs 1.8%;was in4.8% vs 0.6%;was a6.8% vs 1.9% - grounding-fever-contradicted:
did not6.1% vs 0.6%;incapable5.8% vs 0.1%;incapable of5.8% vs 0.1%;did6.2% vs 0.9%;not14.9% vs 7.4%;film11.0% vs 4.3%;only11.0% vs 4.4%;was27.5% vs 26.4%;the54.8% vs 81.4%;of being2.7% vs 0.1% - grounding-fever-joined-contradicted:
is a21.5% vs 6.2%;not20.4% vs 7.4%;is not9.4% vs 1.2%;is50.3% vs 39.9%;was36.1% vs 26.3%;a50.0% vs 47.2%;not a4.2% vs 0.1%;is not a3.4% vs 0.1%;incapable3.1% vs 0.1%;incapable of3.1% vs 0.1% - grounding-fever-joined-supported:
is a25.9% vs 5.8%;is50.1% vs 39.8%;a53.4% vs 47.1%;was35.9% vs 26.2%;film13.5% vs 4.2%;a film5.1% vs 0.4%;is an7.9% vs 1.8%;has13.2% vs 6.6%;in41.3% vs 49.0%;movie5.1% vs 0.9% - grounding-fever-joined-unsupported:
has19.8% vs 6.6%;a53.8% vs 47.2%;was in7.0% vs 0.7%;in a11.7% vs 3.0%;in49.2% vs 48.8%;award7.0% vs 1.0%;was33.3% vs 26.3%;starred in4.9% vs 0.7%;is a12.4% vs 6.3%;was in a3.3% vs 0.2% - grounding-fever-supported:
the58.0% vs 81.6%;film10.7% vs 4.2%;was26.1% vs 26.5%;of39.7% vs 54.1%;a36.1% vs 47.7%;is a10.6% vs 6.3%;in35.3% vs 49.3%;is30.6% vs 40.5%;a film2.8% vs 0.5%;has9.0% vs 6.7% - grounding-fever-unsupported:
has14.2% vs 6.6%;in a9.3% vs 3.0%;in39.7% vs 49.0%;was25.9% vs 26.5%;a34.5% vs 47.6%;the49.4% vs 81.4%;of35.2% vs 54.0%;nominated for1.9% vs 0.2%;film7.3% vs 4.4%;series4.0% vs 1.5% - grounding-hotpot-contradicted:
who23.2% vs 6.0%;film16.6% vs 4.3%;was46.4% vs 26.2%;american13.5% vs 3.7%;he17.7% vs 6.8%;is57.3% vs 39.8%;looking for3.7% vs 0.3%;is the13.7% vs 4.8%;directed7.1% vs 1.5%;known10.3% vs 3.0% - grounding-hotpot-gold-removed:
who25.5% vs 5.8%;american15.0% vs 3.6%;film15.7% vs 4.2%;was46.2% vs 26.0%;is59.1% vs 39.7%;which is10.5% vs 2.6%;which22.6% vs 10.5%;an american7.0% vs 1.2%;is the13.9% vs 4.7%;who was6.1% vs 1.0% - grounding-hotpot-number-changed:
film23.1% vs 4.2%;was60.9% vs 26.0%;born on9.4% vs 0.9%;you are asking6.5% vs 0.4%;are asking6.5% vs 0.4%;are asking about6.5% vs 0.4%;asking6.8% vs 0.5%;which was10.7% vs 1.5%;asking about6.5% vs 0.4%;directed10.4% vs 1.4% - grounding-hotpot-one-removed:
born15.9% vs 2.3%;was born13.8% vs 1.8%;was52.1% vs 26.0%;who22.1% vs 5.9%;born on8.5% vs 0.9%;was born on7.8% vs 0.7%;american14.3% vs 3.7%;he was born5.5% vs 0.4%;he18.9% vs 6.7%;film14.5% vs 4.3% - grounding-hotpot-partly:
which34.4% vs 10.0%;who25.1% vs 5.6%;was55.7% vs 25.5%;born14.2% vs 2.2%;which was10.3% vs 1.3%;in78.9% vs 47.8%;was born11.3% vs 1.7%;film15.3% vs 4.1%;born in7.3% vs 0.9%;he19.1% vs 6.5% - grounding-hotpot-supported:
who22.6% vs 5.5%;was48.2% vs 25.5%;film15.1% vs 4.0%;american13.7% vs 3.4%;is57.3% vs 39.3%;which23.0% vs 10.2%;he16.8% vs 6.5%;in60.0% vs 48.3%;which is9.0% vs 2.5%;is the12.4% vs 4.6% - grounding-squad-combined:
were16.2% vs 4.7%;government6.6% vs 1.5%;that32.2% vs 12.9%;other9.1% vs 2.5%;they18.4% vs 6.5%;these11.6% vs 3.5%;used9.1% vs 2.6%;this37.7% vs 15.9%;it was9.0% vs 2.7%;this happened2.0% vs 0.2% - grounding-squad-contradicted:
were10.9% vs 5.0%;this26.8% vs 16.4%;its9.2% vs 4.1%;that21.7% vs 13.3%;these8.2% vs 3.7%;according to3.1% vs 0.9%;according3.1% vs 0.9%;as24.1% vs 15.5%;their9.8% vs 5.0%;act2.2% vs 0.5% - grounding-squad-gold-removed:
were11.9% vs 4.9%;this25.9% vs 16.4%;these7.7% vs 3.7%;and53.2% vs 41.2%;they11.9% vs 6.8%;the city3.1% vs 1.1%;difficult0.9% vs 0.1%;had been1.6% vs 0.3%;as22.4% vs 15.5%;that19.7% vs 13.4% - grounding-squad-joined-contradicted:
this47.2% vs 16.1%;were18.4% vs 4.9%;that37.3% vs 13.2%;it was11.1% vs 2.7%;this was4.0% vs 0.6%;would5.0% vs 0.8%;their16.3% vs 5.0%;they20.5% vs 6.7%;it38.4% vs 14.7%;they were3.8% vs 0.6% - grounding-squad-joined-supported:
were18.0% vs 4.6%;they21.3% vs 6.3%;this39.6% vs 15.6%;these12.8% vs 3.4%;that33.1% vs 12.7%;it35.4% vs 14.2%;as35.3% vs 14.9%;their14.3% vs 4.7%;many6.3% vs 1.5%;he17.8% vs 6.4% - grounding-squad-joined-unsupported:
these14.2% vs 3.6%;were17.5% vs 4.8%;this42.8% vs 16.0%;they21.0% vs 6.6%;most10.1% vs 2.7%;often4.5% vs 0.8%;as36.7% vs 15.2%;that the4.5% vs 0.9%;because9.6% vs 2.8%;that31.7% vs 13.1% - grounding-squad-number-changed:
had11.1% vs 3.7%;in 19371.2% vs 0.0%;were12.7% vs 5.0%;dates back1.2% vs 0.1%;18801.2% vs 0.1%;as of3.1% vs 0.5%;in 19760.9% vs 0.0%;in 18800.9% vs 0.0%;19260.9% vs 0.0%;hyderabad1.2% vs 0.1% - grounding-squad-partly:
in75.8% vs 47.2%;over10.5% vs 3.1%;particularly3.2% vs 0.2%;in his3.1% vs 0.3%;roughly2.1% vs 0.2%;that22.4% vs 13.0%;his10.3% vs 4.5%;around3.5% vs 0.8%;especially1.9% vs 0.2%;were10.5% vs 4.8% - grounding-squad-supported:
this26.3% vs 16.0%;were10.6% vs 4.8%;that21.4% vs 13.0%;they12.3% vs 6.6%;and52.0% vs 40.9%;as22.8% vs 15.3%;these7.6% vs 3.6%;it21.7% vs 14.6%;their9.0% vs 4.9%;to49.1% vs 40.3% - grounding-squad-unanswerable:
allowing7.8% vs 0.1%;it is called3.9% vs 0.0%;lacking3.9% vs 0.0%;philosophical3.9% vs 0.0%;have no3.9% vs 0.0%;possess3.9% vs 0.0%;forming3.9% vs 0.1%;history of3.9% vs 0.1%;knowledge3.9% vs 0.1%;stopped3.9% vs 0.1%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- grounding-docs-unanswerable:
yes44.9% vs 3.5% - grounding-docs-joined-supported:
must44.8% vs 10.3% - grounding-docs-joined-contradicted:
your38.2% vs 9.3% - grounding-docs-joined-supported:
if35.7% vs 8.8% - grounding-docs-joined-supported:
your34.4% vs 8.6% - grounding-docs-joined-supported:
within30.5% vs 7.4% - grounding-docs-joined-unsupported:
days30.1% vs 6.6% - grounding-docs-joined-supported:
you must28.8% vs 6.0% - grounding-docs-joined-unsupported:
you must28.3% vs 6.6% - grounding-docs-joined-supported:
no28.2% vs 5.9% - grounding-docs-joined-supported:
days28.1% vs 6.1% - grounding-docs-joined-contradicted:
you must27.8% vs 6.7% - grounding-docs-joined-unsupported:
no27.7% vs 6.5% - grounding-fever-joined-supported:
is a25.9% vs 5.8% - grounding-hotpot-gold-removed:
who25.5% vs 5.8% - grounding-hotpot-partly:
who25.1% vs 5.6% - grounding-docs-unanswerable:
can23.8% vs 5.8% - grounding-docs-joined-contradicted:
will23.8% vs 5.7% - grounding-hotpot-number-changed:
film23.1% vs 4.2% - grounding-fever-combined:
is a22.9% vs 5.6% - grounding-hotpot-supported:
who22.6% vs 5.5% - grounding-docs-joined-contradicted:
yes19.7% vs 3.8% - grounding-docs-joined-supported:
hours19.6% vs 4.2% - grounding-docs-joined-contradicted:
percent19.2% vs 4.4% - grounding-docs-joined-contradicted:
hours19.2% vs 4.6% - grounding-docs-joined-unsupported:
five19.1% vs 4.4% - grounding-docs-joined-unsupported:
hours18.8% vs 4.6% - grounding-docs-joined-unsupported:
percent18.7% vs 4.4% - grounding-docs-joined-contradicted:
any18.3% vs 4.2% - grounding-docs-joined-contradicted:
if you18.3% vs 4.4% - grounding-docs-joined-unsupported:
if you17.7% vs 4.3% - grounding-docs-joined-supported:
percent17.4% vs 4.1% - grounding-docs-joined-supported:
if you17.2% vs 4.0% - grounding-hotpot-one-removed:
born15.9% vs 2.3% - grounding-docs-joined-supported:
per15.4% vs 3.7% - grounding-docs-contradicted:
yes15.3% vs 3.8% - grounding-hotpot-gold-removed:
american15.0% vs 3.6% - grounding-docs-unanswerable:
you can14.4% vs 2.9% - grounding-hotpot-partly:
born14.2% vs 2.2% - grounding-docs-unanswerable:
yes the14.1% vs 0.5%
Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):
- all rows (25,380 rows): none
Shortcut models
Predicting the label class on test (2,080 rows). Chance 25.0%, majority class ('supported') 35.0%; balanced chance 25.0%.
| model (logistic regression, trained on the train sample) | test accuracy | balanced accuracy (mean recall) |
|---|---|---|
| bag of words, whole state (words and word pairs) | 34.8% | 28.4% |
bag of words, main text only (answer) |
39.0% | 33.7% |
| surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) | 37.1% | 28.4% |
| surface features of the main text only | 35.7% | 26.3% |
Strongest single surface features (logistic regression on one feature, balanced accuracy on test):
| feature | accuracy | balanced accuracy |
|---|---|---|
| other_state_chars(log) | 36.2% | 27.2% |
| state_chars(log) | 35.3% | 26.3% |
| digit_ratio | 35.0% | 25.3% |
| nonascii_ratio | 35.0% | 25.1% |
| count_" | 35.0% | 25.1% |
| chars(log) | 35.0% | 25.0% |
| words(log) | 35.0% | 25.0% |
| nonlatin | 35.0% | 25.0% |
| starts_upper | 35.0% | 25.0% |
| all_lower | 35.0% | 25.0% |
- Predicting the row kind instead (32 kinds, balanced chance 3.1%): surface features balanced accuracy 20.3% (accuracy 27.6%); bag of words of the state 14.8% (accuracy 23.5%).
Other state fields alone (predicting the label class on test from one field, without the main text):
| field | treated as | accuracy | balanced accuracy |
|---|---|---|---|
| sources | text: bag of words / length+empty | 33.5% / 36.2% | 27.3% / 27.3% |
No-meaning option picker: a logistic ranker scores each option from its position, length, key type, fixed-option identity and shape (commas, brackets, capitals), never reading the state or the option's words, and picks the top option per row.
- Test accuracy 35.0% against uniform chance 25.0% (this includes the fixed options, whose share is a class prior).
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 2,291 (9.0%); groups: 2,176; largest group 4.
- Train rows identical in the whole prompt (state, options, instructions): 0.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 2170 groups (4455 rows). (Can be legitimate when the rest of the state or the options differ.)
- "Cosmopolitan as of 2011 contains content which includes articles on home decor." ×4: unsupported 3, contradicted 1
- "Navin Kumar served as the first chairman of the Goods and Services Tax Network (GSTN). This indirect tax was introduced…" ×3: supported 1, unsupported 1, partly_supported 1
- "The bassist of Rusted Root, the band that released their fifth studio album "Welcome to My Party", is Patrick Norman. T…" ×3: unsupported 1, partly_supported 1, supported 1
- "John Leguizamo is the Colombian-American actor who starred in the 2000 drama film "King of the Jungle." He is also well…" ×3: partly_supported 1, supported 1, unsupported 1
- "Santana Row is located across Stevens Creek Boulevard from Westfield Valley Fair, which is commonly known as Valley Fai…" ×3: unsupported 1, partly_supported 1, supported 1
- "Governor-elect Alejandro García Padilla fought against statehood by asking President Barack Obama to reject the referen…" ×3: unsupported 1, supported 1, partly_supported 1
- "The LA Galaxy, whose president is Chris Klein and which is based in Carson, California, began to play in 1996." ×3: supported 1, partly_supported 1, unsupported 1
- "Alex Cox directed Sid and Nancy, the 1986 British biopic starring Gary Oldman and Chloe Webb that was parodied in The S…" ×3: supported 1, unsupported 1, partly_supported 1
Most repeated main texts in train:
- ×4: "cosmopolitan as of 2011 contains content which includes articles on home decor." (unsupported 3, contradicted 1)
- ×3: "zoé is the older band. they initially formed in mexico city in 1994, while flyleaf was formed in texas in 2002." (supported 1, partly_supported 1, unsupported 1)
- ×3: "zazie beetz has been cast as neena thurman in "deadpool 2", which is directed by david leitch. the film is an upcoming american superhero m…" (unsupported 1, partly_supported 1, supported 1)
- ×3: "vidushi shashikala dani is the only all india radio graded female exponent of the jaltarang, which is a percussion instrument." (supported 1, partly_supported 1, unsupported 1)
- ×3: "vch was reportedly the most popular programming on qube, a cable television system that was launched on december 1, 1977." (partly_supported 1, supported 1, unsupported 1)
- ×3: "tom rolt was a prolific english writer and the biographer of major civil engineering figures including isambard kingdom brunel and thomas t…" (supported 1, partly_supported 1, unsupported 1)
- ×3: "tlc, the american girl group that released "girl talk" in 2002, received the million certification from the recording industry association …" (unsupported 1, supported 1, partly_supported 1)
- ×3: "thomas bartley officiated in home tests against pakistan. the pakistan national cricket team is popularly referred to as the shaheens, men …" (supported 1, partly_supported 1, unsupported 1)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 0 (0.0%) | |
| calibration | 0 (0.0%) | |
| test | 0 (0.0%) |
Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)
- Train: 5,705 near-duplicate pairs; 7,898 rows (31.1%) sit in 3,333 clusters; largest cluster 6; excess rows (cluster size − 1) 4,565 (18.0%).
- Clusters with more than one label class: 3,288 (7,808 rows).
- ×6: "Garrison Hearst, who won the NFL Comeback Player of the Year Award in 2001, was featured on the cover of Madd…" → partly_supported 2, contradicted 2, supported 1, unsupported 1
- ×6: "The screenwriter you are asking about is Marc Silverstein, who co-wrote the film Valentine's Day directed by …" → contradicted 2, partly_supported 2, supported 1, unsupported 1
- ×6: "The LIGO Scientific Collaboration (LSC) was established in 1997 under the leadership of Barry Barish, an Amer…" → partly_supported 2, contradicted 2, supported 1, unsupported 1
- ×6: ""Këmisha e zezë" was the organ of the Albanian Fascist Party. This party held nominal power over Albania from…" → contradicted 2, partly_supported 2, supported 1, unsupported 1
- ×5: "Samuel Fraunces provided for prisoners during the Civil War. He was the owner and operator of Fraunces Tavern…" → contradicted 2, unsupported 1, partly_supported 1, supported 1
- Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 0 (0.0%), test 0 (0.0%)
Largest train clusters:
- ×6: "Garrison Hearst, who won the NFL Comeback Player of the Year Award in 2001, was featured on the cover of Madden NFL 99. Specifically, the E…"
- ×6: "The screenwriter you are asking about is Marc Silverstein, who co-wrote the film Valentine's Day directed by Garry Marshall. He was born on…"
- ×6: "The LIGO Scientific Collaboration (LSC) was established in 1997 under the leadership of Barry Barish, an American experimental physicist. I…"
- ×6: ""Këmisha e zezë" was the organ of the Albanian Fascist Party. This party held nominal power over Albania from 1939, when the country was co…"
- ×5: "Samuel Fraunces provided for prisoners during the Civil War. He was the owner and operator of Fraunces Tavern, which is situated at 54 Pear…"
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 0 |
| dev | 0 | 0 |
| calibration | 0 | 0 |
| test | 0 | 0 |
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 0 (0.0%) | |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 0 (0.0%) | |
| chat preamble ('Here is/are...', 'Sure!') | 0 (0.0%) | |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 0 (0.0%) | |
| encoding garbage (mojibake/replacement char) | 0 (0.0%) | |
| HTML tag | 0 (0.0%) | |
| HTML entity | 0 (0.0%) | |
| base64-like run (40+ chars) | 0 (0.0%) | |
| URL | 0 (0.0%) |
- Possibly cut off: 0 of 3,620 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: contradicted 0.0%, partly_supported 0.0%, supported 0.0%, unsupported 0.0%
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- none
6. Samples
20 random train rows per kind: ground-grounding-samples.txt. Reading notes are in the findings above.
QA: ground-rerank
Checked 2026-09-30 13:20 by adapters/qa/qa.py (READY file READY-ground, v4 (held-out second opinion; grounding answers balanced on surface features, 2026-09-30)).
Verdict: PASS WITH NOTES (see ground-grounding.md for the full notes)
Automatic flags (for the reviewer to judge; not all are problems)
- kind 'rerank-docs' share varies across splits by more than 5 points: train 11.5%, dev 6.1%, calibration 8.2%, test 9.7%
- kind 'rerank-squad' share varies across splits by more than 5 points: train 59.1%, dev 66.0%, calibration 60.7%, test 59.1%
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 25,380 | 5141 | train.jsonl |
| dev | 900 | 205 | dev.jsonl |
| calibration | 720 | 190 | calibration.jsonl |
| test | 2,080 | 565 | test.jsonl |
- Train sha256:
c8b38f85a728f5aa916c2b4b2cf18a8cd63567fde55572cc4b0f16b155915fb2(READY saysc8b38f85a728f5aa916c2b4b2cf18a8cd63567fde55572cc4b0f16b155915fb2: match) - Main text field (the text the phrase and length checks use):
state.question. - Label classes: <listed option>, none, p1, p10, p11, p12, p13, p14, p15, p16, p17, p18, p19, p2, p20, p21, p22, p3, p4, p5, p6, p7, p8, p9 (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): rerank-docs, rerank-docs-unanswerable, rerank-hotpot, rerank-hotpot-removed, rerank-squad, rerank-squad-nearmiss, rerank-squad-nearmiss-none, rerank-squad-removed, rerank-squad-unanswerable.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| <listed option> | 4.5% | 5.9% | 5.6% | 5.9% | 1,136 |
| none | 15.0% | 15.0% | 15.0% | 15.0% | 3,807 |
| p1 | 7.9% | 8.9% | 8.6% | 8.1% | 1,994 |
| p10 | 3.2% | 3.4% | 2.5% | 3.7% | 804 |
| p11 | 2.6% | 3.1% | 2.6% | 2.7% | 669 |
| p12 | 2.4% | 2.9% | 2.1% | 2.4% | 603 |
| p13 | 2.0% | 1.8% | 1.9% | 1.7% | 495 |
| p14 | 1.9% | 1.7% | 2.2% | 2.1% | 492 |
| p15 | 1.7% | 1.2% | 1.7% | 1.2% | 421 |
| p16 | 1.5% | 1.8% | 1.8% | 1.6% | 377 |
| p17 | 1.2% | 1.0% | 1.2% | 1.6% | 314 |
| p18 | 1.3% | 1.0% | 1.7% | 0.9% | 320 |
| p19 | 1.0% | 0.7% | 0.7% | 1.2% | 245 |
| p2 | 7.8% | 9.6% | 8.2% | 7.3% | 1,991 |
| p20 | 0.8% | 0.4% | 0.8% | 1.1% | 210 |
| p21 | 0.6% | 0.4% | 0.3% | 0.6% | 151 |
| p22 | 0.6% | 0.4% | 0.3% | 0.7% | 145 |
| p3 | 7.7% | 7.8% | 9.2% | 7.5% | 1,960 |
| p4 | 7.8% | 7.1% | 6.1% | 6.8% | 1,977 |
| p5 | 7.6% | 7.4% | 8.6% | 7.5% | 1,931 |
| p6 | 6.9% | 5.7% | 6.2% | 6.3% | 1,740 |
| p7 | 5.6% | 5.8% | 5.3% | 5.8% | 1,421 |
| p8 | 4.6% | 3.9% | 4.3% | 4.0% | 1,166 |
| p9 | 4.0% | 3.1% | 3.1% | 4.4% | 1,011 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| rerank-docs | 11.5% | 6.1% | 8.2% | 9.7% | 2,923 |
| rerank-docs-unanswerable | 2.5% | 1.1% | 1.1% | 1.5% | 639 |
| rerank-hotpot | 13.9% | 11.8% | 15.6% | 15.5% | 3,518 |
| rerank-hotpot-removed | 6.8% | 6.3% | 7.5% | 8.8% | 1,714 |
| rerank-squad | 59.1% | 66.0% | 60.7% | 59.1% | 15,007 |
| rerank-squad-nearmiss | 0.5% | 1.1% | 0.6% | 0.7% | 125 |
| rerank-squad-nearmiss-none | 0.3% | 0.6% | 0.3% | 0.2% | 74 |
| rerank-squad-removed | 4.4% | 5.9% | 5.6% | 3.3% | 1,124 |
| rerank-squad-unanswerable | 1.0% | 1.1% | 0.6% | 1.2% | 256 |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Options per choice row
| split | min | median | p99 | max |
|---|---|---|---|---|
| train | 6 | 13 | 40 | 41 |
| dev | 6 | 13 | 40 | 41 |
| calibration | 6 | 14 | 40 | 41 |
| test | 6 | 14 | 41 | 41 |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | source.input_tokens (25380/25380 rows) | 2165 | 7222 | 8126 | 0 |
| dev | source.input_tokens (900/900 rows) | 2228 | 7084 | 8065 | 0 |
| calibration | source.input_tokens (720/720 rows) | 2184 | 7233 | 7574 | 0 |
| test | source.input_tokens (2080/2080 rows) | 2290 | 7282 | 8025 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| question | 25,380 (100.0%) |
Instructions (train)
- Canonical (the most common text) 70.3%, reworded 26.7% (80 distinct rewordings), none 3.0%. Target about 70 / 27 / 3.
- Canonical text: "Which passage answers the question? If none of them does, choose none."
sourceinstruction tag: canonical 70.3%, variant 26.7%, none 3.0%
| class | canonical | none |
|---|---|---|
| <listed option> | 71.4% | 2.5% |
| none | 70.3% | 2.8% |
| p1 | 69.0% | 3.2% |
| p10 | 72.0% | 2.9% |
| p11 | 72.2% | 3.4% |
| p12 | 67.0% | 3.8% |
| p13 | 70.7% | 3.6% |
| p14 | 69.9% | 2.4% |
| p15 | 71.0% | 2.9% |
| p16 | 72.1% | 3.2% |
| p17 | 63.7% | 3.8% |
| p18 | 72.2% | 1.9% |
| p19 | 66.9% | 3.7% |
| p2 | 69.9% | 2.8% |
| p20 | 66.7% | 3.3% |
| p21 | 67.5% | 0.7% |
| p22 | 66.2% | 4.8% |
| p3 | 71.8% | 3.4% |
| p4 | 69.0% | 3.0% |
| p5 | 70.7% | 3.3% |
| p6 | 70.6% | 3.1% |
| p7 | 71.9% | 3.3% |
| p8 | 69.3% | 2.7% |
| p9 | 70.7% | 3.2% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 25,380 train rows; models are scored on the full test file (2,080 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | <listed option> | 1136 | 39 | 61 | 114 | 72 |
| train | none | 3807 | 42 | 71 | 139 | 84 |
| train | p1 | 1994 | 38 | 60 | 108 | 70 |
| train | p10 | 804 | 38 | 61 | 116 | 72 |
| train | p11 | 669 | 37 | 58 | 105 | 67 |
| train | p12 | 603 | 37 | 60 | 112 | 69 |
| train | p13 | 495 | 39 | 59 | 108 | 69 |
| train | p14 | 492 | 37 | 59 | 112 | 68 |
| train | p15 | 421 | 39 | 59 | 113 | 69 |
| train | p16 | 377 | 37 | 60 | 105 | 68 |
| train | p17 | 314 | 39 | 60 | 115 | 71 |
| train | p18 | 320 | 38 | 61 | 111 | 70 |
| train | p19 | 245 | 36 | 56 | 96 | 63 |
| train | p2 | 1991 | 37 | 60 | 111 | 70 |
| train | p20 | 210 | 40 | 58 | 100 | 68 |
| train | p21 | 151 | 35 | 55 | 97 | 64 |
| train | p22 | 145 | 34 | 62 | 108 | 67 |
| train | p3 | 1960 | 37 | 60 | 108 | 69 |
| train | p4 | 1977 | 37 | 60 | 111 | 69 |
| train | p5 | 1931 | 39 | 61 | 110 | 71 |
| train | p6 | 1740 | 38 | 59 | 107 | 69 |
| train | p7 | 1421 | 37 | 59 | 111 | 69 |
| train | p8 | 1166 | 37 | 59 | 110 | 69 |
| train | p9 | 1011 | 37 | 62 | 113 | 71 |
| test | <listed option> | 123 | 40 | 61 | 128 | 74 |
| test | none | 312 | 46 | 83 | 147 | 92 |
| test | p1 | 169 | 37 | 61 | 108 | 68 |
| test | p10 | 77 | 40 | 65 | 122 | 75 |
| test | p11 | 56 | 45 | 67 | 114 | 76 |
| test | p12 | 49 | 31 | 64 | 117 | 70 |
| test | p13 | 36 | 39 | 57 | 93 | 63 |
| test | p14 | 43 | 33 | 64 | 110 | 73 |
| test | p15 | 24 | 40 | 75 | 173 | 100 |
| test | p16 | 33 | 37 | 71 | 108 | 80 |
| test | p17 | 33 | 39 | 67 | 93 | 68 |
| test | p18 | 19 | 32 | 45 | 98 | 57 |
| test | p19 | 24 | 38 | 57 | 94 | 62 |
| test | p2 | 151 | 37 | 65 | 121 | 73 |
| test | p20 | 22 | 43 | 64 | 120 | 75 |
| test | p21 | 13 | 50 | 68 | 161 | 89 |
| test | p22 | 14 | 36 | 71 | 111 | 75 |
| test | p3 | 157 | 41 | 66 | 119 | 74 |
| test | p4 | 142 | 37 | 59 | 109 | 69 |
| test | p5 | 155 | 36 | 63 | 114 | 72 |
| test | p6 | 132 | 38 | 63 | 100 | 69 |
| test | p7 | 121 | 37 | 64 | 119 | 77 |
| test | p8 | 84 | 35 | 59 | 101 | 67 |
| test | p9 | 91 | 34 | 57 | 98 | 64 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| rerank-docs | 2923 | 54 | 55 | 0 | 14 |
| rerank-docs-unanswerable | 639 | 62 | 63 | 0 | 13 |
| rerank-hotpot | 3518 | 111 | 128 | 0 | 13 |
| rerank-hotpot-removed | 1714 | 99 | 115 | 0 | 14 |
| rerank-squad | 15007 | 56 | 59 | 0 | 13 |
| rerank-squad-nearmiss | 125 | 59 | 59 | 0 | 11 |
| rerank-squad-nearmiss-none | 74 | 58 | 60 | 0 | 13 |
| rerank-squad-removed | 1124 | 57 | 58 | 0 | 13 |
| rerank-squad-unanswerable | 256 | 48 | 50 | 0 | 13 |
Correct option: longest / shortest / position / key
For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):
| split | rows | correct is longest | correct is shortest | chance (1/listed) | mean relative position (0 first, 1 last; 0.5 expected) | position fifths |
|---|---|---|---|---|---|---|
| train | 1136 | 9.9% | 10.3% | 11.8% | 0.493 | 23% / 17% / 19% / 18% / 23% |
| dev | 53 | 7.5% | 7.5% | 11.8% | 0.554 | 15% / 23% / 11% / 25% / 26% |
| calibration | 40 | 7.5% | 12.5% | 12.4% | 0.493 | 20% / 15% / 28% / 12% / 25% |
| test | 123 | 17.1% | 10.6% | 11.9% | 0.466 | 29% / 17% / 20% / 11% / 24% |
Correct key and position by option count
| split | options | rows | mean options | top correct keys | most common position (0-based) |
|---|---|---|---|---|---|
| train | 6-10 | 9239 | 8.1 | none 15.3%, p2 12.8%, p4 12.8%, p3 12.5%, p1 12.2% | 1 (13.1%) |
| train | 11-30 | 13162 | 17.5 | none 14.6%, p6 6.1%, p1 6.0%, p5 5.7%, p2 5.7% | 0 (6.6%) |
| train | 31-80 | 2979 | 35.7 | none 15.9%, p18 3.0%, p15 2.9%, p12 2.8%, p16 2.7% | 14 (3.4%) |
| test | 6-10 | 698 | 8.1 | none 14.2%, p3 13.3%, p1 13.0%, p2 12.6%, p5 11.6% | 3 (14.0%) |
| test | 11-30 | 1097 | 17.5 | none 16.0%, p1 6.6%, p10 6.5%, p9 6.4%, p5 5.9% | 10 (8.0%) |
| test | 31-80 | 285 | 35.9 | none 13.3%, p24 5.3%, p23 4.9%, p20 4.2%, p26 3.9% | 26 (5.6%) |
- Train rows whose correct listed key is the first listed key (t1/o1): 12.1%, chance 11.8%.
Option count by label class (train)
| class | rows | min | median | mean | max |
|---|---|---|---|---|---|
| <listed option> | 1136 | 24 | 35 | 34.4 | 41 |
| none | 3807 | 6 | 13 | 16.3 | 41 |
| p1 | 1994 | 6 | 10 | 12.2 | 41 |
| p10 | 804 | 11 | 15 | 17.4 | 41 |
| p11 | 669 | 12 | 18 | 19.7 | 41 |
| p12 | 603 | 13 | 18 | 20.5 | 41 |
| p13 | 495 | 14 | 19 | 21.5 | 41 |
| p14 | 492 | 15 | 20 | 22.2 | 41 |
| p15 | 421 | 16 | 21 | 23.5 | 41 |
| p16 | 377 | 17 | 21 | 24.3 | 41 |
| p17 | 314 | 18 | 21 | 24.1 | 41 |
| p18 | 320 | 19 | 23 | 25.9 | 41 |
| p19 | 245 | 20 | 26 | 26.8 | 41 |
| p2 | 1991 | 6 | 9 | 11.6 | 41 |
| p20 | 210 | 21 | 27 | 28.3 | 40 |
| p21 | 151 | 22 | 28 | 29.5 | 41 |
| p22 | 145 | 23 | 31 | 31.0 | 41 |
| p3 | 1960 | 6 | 9 | 11.8 | 41 |
| p4 | 1977 | 6 | 9 | 11.8 | 41 |
| p5 | 1931 | 6 | 10 | 12.0 | 41 |
| p6 | 1740 | 7 | 11 | 13.0 | 41 |
| p7 | 1421 | 8 | 11 | 13.9 | 41 |
| p8 | 1166 | 9 | 12 | 14.8 | 41 |
| p9 | 1011 | 10 | 14 | 16.1 | 41 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 15.0%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| kind | 9 | 23.0% | rerank-squad: p1 9.4%, p2 9.4%; rerank-hotpot: p4 9.4%, p5 9.2%; rerank-docs: p6 9.2%, p7 8.6%; rerank-hotpot-removed: none 100.0%; rerank-squad-removed: none 100.0%; rerank-docs-unanswerable: none 100.0% |
| second_opinion | 2 | 22.9% | null: p1 9.2%, p2 9.2%; deepseek-v4-flash: none 100.0% |
| second_pool | 2 | 15.7% | null: none 18.5%, p2 7.6%; true: p4 9.4%, p3 8.9% |
| dataset | 3 | 15.0% | squad2: none 8.8%, p1 8.6%; hotpotqa: none 32.8%, p4 6.3%; teacher-docs: none 17.9%, p6 7.5% |
| license | 2 | 15.0% | CC BY-SA 4.0: none 14.5%, p4 8.0%; generated by the local teacher (qwen3.8-flash-next): none 17.9%, p6 7.5% |
| url | 3 | 15.0% | https://huggingface.co/datasets/rajpurkar/squad_v2: none 8.8%, p1 8.6%; https://huggingface.co/datasets/hotpotqa/hotpot_qa: none 32.8%, p4 6.3%; : none 17.9%, p6 7.5% |
| revision | 3 | 15.0% | 3ffb306f725f7d2ce8394bc1873b24868140c412: none 8.8%, p1 8.6%; 1908d6afbbead072334abe2965f91bd2709910ab: none 32.8%, p4 6.3%; : none 17.9%, p6 7.5% |
| titles_shown | 2 | 15.0% | true: none 14.9%, p2 8.1%; false: none 15.1%, p1 8.1% |
| text_teacher | 3 | 15.0% | none (public data and code): none 14.3%, p1 8.0%; local: qwen3.8-flash-next: none 18.1%, p6 7.5%; hosted: qwen3.8-max: none 21.0%, p5 8.1% |
Formatting by label class (main text, share of rows)
| feature | lowest classes | highest classes | |
|---|---|---|---|
| ends with ? | p14 98%, p16 98%, p2 98% | p21 100%, p19 99%, p18 99% | |
| ends with . | p11 0%, p13 0%, p14 0% | p15 1%, p22 1%, p7 1% | |
| ends with ! | <listed option> 0%, none 0%, p1 0% | p9 0%, p8 0%, p7 0% | |
| no end punctuation | p21 0%, p22 0%, p19 0% | p16 2%, p14 2%, p17 2% | |
| starts lowercase | p20 0%, p22 0%, p12 0% | p17 2%, p6 1%, p21 1% | |
| all lowercase | p20 0%, p22 0%, p11 0% | p17 1%, p10 1%, p19 1% | |
| has a digit | p11 13%, p8 13%, p20 13% | none 21%, p17 19%, <listed option> 18% | |
| has newline | <listed option> 0%, none 0%, p1 0% | p9 0%, p8 0%, p7 0% | |
| has quotes | p19 2%, p21 2%, p11 2% | none 7%, p16 7%, p18 7% | |
| has markup (HTML/markdown) | <listed option> 0%, none 0%, p10 0% | p14 0%, p4 0%, p1 0% | |
| has URL | <listed option> 0%, none 0%, p1 0% | p9 0%, p8 0%, p7 0% | |
| non-ASCII | p12 0%, p17 1%, p22 1% | p21 3%, none 2%, p18 2% | |
| non-Latin script | <listed option> 0%, p1 0%, p10 0% | p22 1%, p9 0%, p5 0% | |
| emoji | <listed option> 0%, none 0%, p1 0% | p6 0%, p9 0%, p8 0% | |
| ALL-CAPS word (4+) | p11 1%, p6 1%, p5 1% | p13 3%, p1 3%, p10 2% | |
| contains ' - ' or — | p10 0%, p12 0%, p13 0% | p14 0%, p1 0%, p11 0% |
Same, by row kind
| feature | rerank-docs | rerank-docs-unanswerable | rerank-hotpot | rerank-hotpot-removed | rerank-squad | rerank-squad-nearmiss | rerank-squad-nearmiss-none | rerank-squad-removed | rerank-squad-unanswerable |
|---|---|---|---|---|---|---|---|---|---|
| ends with ? | 99% | 99% | 97% | 97% | 99% | 100% | 100% | 98% | 98% |
| ends with . | 0% | 0% | 1% | 1% | 0% | 0% | 0% | 0% | 0% |
| ends with ! | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| no end punctuation | 1% | 1% | 2% | 1% | 1% | 0% | 0% | 1% | 2% |
| starts lowercase | 3% | 3% | 1% | 1% | 1% | 0% | 1% | 0% | 0% |
| all lowercase | 3% | 2% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| has a digit | 4% | 9% | 40% | 35% | 13% | 16% | 11% | 11% | 12% |
| has newline | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| has quotes | 0% | 0% | 19% | 15% | 2% | 0% | 1% | 1% | 1% |
| has markup (HTML/markdown) | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| has URL | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| non-ASCII | 0% | 0% | 4% | 5% | 1% | 0% | 0% | 1% | 0% |
| non-Latin script | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| emoji | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
| ALL-CAPS word (4+) | 0% | 0% | 3% | 3% | 2% | 2% | 3% | 2% | 3% |
| contains ' - ' or — | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% | 0% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, <listed option>: nc z=4 0.4% vs 0.0%; albert z=4 0.5% vs 0.1%; n z=4 0.4% vs 0.0%; meetings z=4 0.3% vs 0.0%; regulators z=4 0.3% vs 0.0%; adolescent z=4 0.3% vs 0.0%; communist z=4 0.4% vs 0.0%; earliest z=4 0.6% vs 0.1%; rings z=4 0.3% vs 0.0%; philippine z=4 0.3% vs 0.0%; tertiary z=4 0.3% vs 0.0%; difficult z=3 0.4% vs 0.0%; tram z=3 0.3% vs 0.0%; specifically z=3 0.3% vs 0.0%; surviving z=3 0.3% vs 0.0%; heir z=3 0.3% vs 0.0%; hunting z=3 0.5% vs 0.1%; princess z=3 0.4% vs 0.0%; helps z=3 0.3% vs 0.0%; quake z=3 0.3% vs 0.0%
Words, none: there z=13 4.6% vs 1.1%; born z=10 4.1% vs 1.2%; available z=10 1.7% vs 0.2%; offer z=9 1.4% vs 0.2%; specific z=7 1.1% vs 0.2%; parking z=7 0.8% vs 0.0%; a z=7 27.5% vs 18.5%; discount z=7 0.8% vs 0.0%; for z=7 18.2% vs 11.6%; that z=7 10.7% vs 6.2%; team z=7 2.7% vs 1.0%; which z=6 15.8% vs 10.2%; album z=6 2.2% vs 0.9%; film z=6 5.4% vs 2.8%; founded z=6 1.8% vs 0.7%; dress z=6 0.5% vs 0.0%; county z=6 1.2% vs 0.4%; won z=5 1.5% vs 0.5%; by z=5 12.0% vs 7.8%; mobile z=5 0.5% vs 0.1%
Words, p1: lies z=4 0.3% vs 0.0%; cultures z=4 0.4% vs 0.1%; philosophical z=4 0.2% vs 0.0%; coastal z=3 0.2% vs 0.0%; foundation z=3 0.4% vs 0.1%; universities z=3 0.5% vs 0.1%; 1948 z=3 0.3% vs 0.0%; cause z=3 0.7% vs 0.2%; hands z=3 0.2% vs 0.0%; industries z=3 0.2% vs 0.0%; society z=3 0.5% vs 0.2%; vegas z=3 0.2% vs 0.0%; basis z=3 0.2% vs 0.0%; officer z=3 0.5% vs 0.1%; absorb z=3 0.2% vs 0.0%; commanding z=3 0.2% vs 0.0%; descended z=3 0.2% vs 0.0%; adding z=3 0.2% vs 0.0%; acre z=3 0.3% vs 0.0%; destroy z=3 0.3% vs 0.0%
Words, p10: initially z=5 0.7% vs 0.1%; planning z=4 0.5% vs 0.0%; courses z=4 0.4% vs 0.0%; bermuda z=4 0.6% vs 0.1%; edwards z=4 0.4% vs 0.0%; vi z=4 0.6% vs 0.1%; impact z=4 0.5% vs 0.1%; recognized z=4 0.4% vs 0.0%; attendance z=4 0.4% vs 0.0%; paul z=4 1.0% vs 0.2%; fly z=3 0.4% vs 0.0%; upload z=3 0.4% vs 0.0%; sciences z=3 0.4% vs 0.0%; 1990s z=3 0.5% vs 0.1%; america's z=3 0.4% vs 0.0%; boys z=3 0.5% vs 0.1%; studying z=3 0.4% vs 0.0%; rated z=3 0.4% vs 0.0%; highest z=3 1.0% vs 0.3%; thompson z=3 0.2% vs 0.0%
Words, p11: 1975 z=4 0.6% vs 0.1%; religion z=4 1.2% vs 0.3%; floor z=4 0.6% vs 0.1%; cartoon z=4 0.4% vs 0.0%; multiple z=4 0.7% vs 0.1%; peninsula z=4 0.4% vs 0.0%; emotion z=4 0.4% vs 0.0%; quickly z=3 0.9% vs 0.2%; graduate z=3 0.4% vs 0.0%; immigrant z=3 0.3% vs 0.0%; adoption z=3 0.3% vs 0.0%; pigment z=3 0.3% vs 0.0%; samoan z=3 0.3% vs 0.0%; doctrine z=3 0.4% vs 0.1%; 17th z=3 0.3% vs 0.0%; traits z=3 0.3% vs 0.0%; agreed z=3 0.3% vs 0.0%; mice z=3 0.3% vs 0.0%; mammals z=3 0.3% vs 0.0%; long z=3 3.7% vs 2.0%
Words, p12: mans z=4 0.5% vs 0.0%; crew z=4 0.5% vs 0.0%; ge z=4 0.5% vs 0.0%; cultural z=4 0.8% vs 0.1%; opposed z=4 0.5% vs 0.0%; largest z=4 2.2% vs 0.7%; 1899 z=4 0.3% vs 0.0%; stays z=4 0.3% vs 0.0%; purposes z=4 0.3% vs 0.0%; container z=4 0.3% vs 0.0%; dangerous z=4 0.3% vs 0.0%; inch z=4 0.3% vs 0.0%; macarthur z=4 0.3% vs 0.0%; penalty z=4 0.8% vs 0.1%; laserdisc z=3 0.5% vs 0.0%; whitehead's z=3 0.5% vs 0.0%; explain z=3 0.5% vs 0.0%; whether z=3 0.3% vs 0.0%; romanian z=3 0.3% vs 0.0%; capita z=3 0.3% vs 0.0%
Words, p13: serving z=4 0.6% vs 0.0%; connection z=4 0.6% vs 0.0%; destinations z=4 0.4% vs 0.0%; rodney z=4 0.4% vs 0.0%; montgomery z=4 0.4% vs 0.0%; millennium z=4 0.4% vs 0.0%; sentences z=4 0.4% vs 0.0%; patriot z=4 0.4% vs 0.0%; tito z=4 0.8% vs 0.1%; proto z=4 0.4% vs 0.0%; significantly z=4 0.4% vs 0.0%; link's z=4 0.4% vs 0.0%; pavilion z=4 0.4% vs 0.0%; concern z=4 0.4% vs 0.0%; ghana z=4 0.4% vs 0.0%; szlachta z=3 0.4% vs 0.0%; jordan z=3 0.4% vs 0.0%; split z=3 0.6% vs 0.1%; tuberculosis z=3 0.4% vs 0.0%; distinguished z=3 0.4% vs 0.0%
Words, p14: prior z=5 1.4% vs 0.2%; listed z=4 1.0% vs 0.1%; schwarzenegger z=4 1.0% vs 0.2%; variation z=4 0.6% vs 0.0%; oldest z=4 1.0% vs 0.2%; unicef z=4 0.4% vs 0.0%; federation z=4 0.4% vs 0.0%; input z=4 0.4% vs 0.0%; skater z=4 0.4% vs 0.0%; poultry z=3 0.4% vs 0.0%; plateau z=3 0.4% vs 0.0%; biology z=3 0.4% vs 0.0%; bills z=3 0.4% vs 0.0%; dave z=3 0.4% vs 0.0%; eu z=3 0.4% vs 0.0%; metro z=3 0.6% vs 0.1%; bell z=3 1.0% vs 0.2%; night z=3 1.0% vs 0.2%; 1860 z=3 0.4% vs 0.0%; galicia z=3 0.4% vs 0.0%
Words, p15: classification z=5 1.0% vs 0.1%; midna z=4 0.5% vs 0.0%; veto z=4 0.5% vs 0.0%; hackers z=4 0.5% vs 0.0%; climate z=4 1.0% vs 0.1%; 1952 z=4 0.5% vs 0.0%; assention z=4 0.5% vs 0.0%; banking z=4 0.5% vs 0.0%; facet z=4 0.5% vs 0.0%; rulers z=3 0.5% vs 0.0%; financially z=3 0.5% vs 0.0%; taiwan z=3 0.5% vs 0.0%; boards z=3 0.5% vs 0.0%; northwest z=3 0.5% vs 0.0%; square z=3 0.7% vs 0.1%; technologies z=3 0.5% vs 0.0%; login z=3 0.5% vs 0.0%; ask z=3 0.5% vs 0.0%; logo z=3 0.5% vs 0.0%; option z=3 0.5% vs 0.0%
Words, p16: passengers z=5 0.8% vs 0.0%; followed z=4 0.8% vs 0.0%; notice z=4 1.1% vs 0.1%; rejection z=4 0.5% vs 0.0%; expanded z=4 0.5% vs 0.0%; sale z=4 0.5% vs 0.0%; ranks z=4 0.5% vs 0.0%; protest z=4 0.5% vs 0.0%; purple z=4 0.5% vs 0.0%; country z=4 3.4% vs 1.2%; city's z=4 0.8% vs 0.1%; middle z=3 1.3% vs 0.2%; header z=3 0.5% vs 0.0%; lab z=3 0.5% vs 0.0%; luxury z=3 0.5% vs 0.0%; tank z=3 0.5% vs 0.0%; principle z=3 0.5% vs 0.0%; automatic z=3 0.5% vs 0.0%; give z=3 1.3% vs 0.3%; decision z=3 0.8% vs 0.1%
Words, p17: mid z=4 1.3% vs 0.1%; stain z=4 0.6% vs 0.0%; 1987 z=4 1.0% vs 0.1%; barcelona's z=4 0.6% vs 0.0%; indies z=4 0.6% vs 0.0%; seattle's z=4 0.6% vs 0.0%; butler z=4 0.6% vs 0.0%; specialized z=4 0.6% vs 0.0%; declined z=4 0.6% vs 0.0%; societies z=4 0.6% vs 0.0%; thames z=3 0.6% vs 0.0%; performing z=3 0.6% vs 0.0%; same z=3 1.9% vs 0.4%; agencies z=3 0.6% vs 0.0%; heritage z=3 0.6% vs 0.0%; underground z=3 0.6% vs 0.0%; ago z=3 0.6% vs 0.0%; went z=3 1.0% vs 0.1%; horse z=3 0.6% vs 0.1%; tibetan z=3 0.6% vs 0.1%
Words, p18: 33 z=5 0.9% vs 0.0%; reportedly z=4 0.6% vs 0.0%; past z=4 0.9% vs 0.1%; cutoff z=4 0.6% vs 0.0%; enact z=4 0.6% vs 0.0%; healing z=4 0.6% vs 0.0%; wire z=4 0.6% vs 0.0%; 9 z=3 0.9% vs 0.1%; what's z=3 1.6% vs 0.3%; throw z=3 0.6% vs 0.0%; 1937 z=3 0.6% vs 0.0%; parental z=3 0.6% vs 0.0%; tablets z=3 0.6% vs 0.0%; empires z=3 0.6% vs 0.0%; range z=3 1.2% vs 0.2%; july z=3 0.9% vs 0.1%; shape z=3 0.6% vs 0.0%; nickname z=3 1.2% vs 0.2%; guam z=3 0.6% vs 0.0%; devices z=3 0.9% vs 0.1%
Words, p19: jay z=4 0.8% vs 0.0%; xbox z=4 0.8% vs 0.0%; quantum z=4 0.8% vs 0.0%; paperwork z=4 0.8% vs 0.0%; playstation z=4 0.8% vs 0.0%; platform z=4 0.8% vs 0.0%; ps3 z=3 0.8% vs 0.0%; slaves z=3 0.8% vs 0.1%; aspect z=3 0.8% vs 0.1%; account z=3 2.0% vs 0.4%; actors z=3 0.8% vs 0.1%; songs z=3 1.2% vs 0.2%; spectre z=3 0.8% vs 0.1%; joint z=3 0.8% vs 0.1%; jazz z=3 0.8% vs 0.1%; testing z=3 0.8% vs 0.1%; die z=3 1.2% vs 0.2%; city's z=3 0.8% vs 0.1%; 1985 z=3 0.8% vs 0.1%; 2011 z=3 1.6% vs 0.3%
Words, p2: argument z=4 0.3% vs 0.0%; congressional z=4 0.3% vs 0.0%; kitchen z=4 0.2% vs 0.0%; n z=4 0.3% vs 0.0%; obesity z=4 0.2% vs 0.0%; empty z=3 0.3% vs 0.0%; stated z=3 0.4% vs 0.1%; sanskrit z=3 0.3% vs 0.1%; help z=3 0.7% vs 0.2%; transportation z=3 0.3% vs 0.0%; books z=3 0.5% vs 0.2%; 1938 z=3 0.2% vs 0.0%; decides z=3 0.2% vs 0.0%; terminate z=3 0.2% vs 0.0%; refused z=3 0.2% vs 0.0%; kathmandu's z=3 0.2% vs 0.0%; motorcycle z=3 0.2% vs 0.0%; schumann z=3 0.2% vs 0.0%; recession z=3 0.2% vs 0.0%; java z=3 0.2% vs 0.0%
Words, p20: canada z=5 1.9% vs 0.1%; resident z=5 1.4% vs 0.1%; rangers z=4 1.0% vs 0.0%; john's z=4 1.0% vs 0.0%; reserve z=4 1.0% vs 0.0%; equipment z=4 1.0% vs 0.0%; transport z=4 1.0% vs 0.0%; resource z=3 1.0% vs 0.1%; low z=3 1.4% vs 0.2%; assistant z=3 1.0% vs 0.1%; issued z=3 1.0% vs 0.1%; projects z=3 1.0% vs 0.1%; fail z=3 1.0% vs 0.1%; metro z=3 1.0% vs 0.1%; mall z=3 1.0% vs 0.1%; chopin z=3 1.9% vs 0.3%; joined z=3 1.0% vs 0.1%; emergency z=3 1.0% vs 0.1%; defeat z=3 1.0% vs 0.1%; hub z=3 1.0% vs 0.1%
Words, p21: statistics z=5 1.3% vs 0.0%; failed z=4 2.0% vs 0.1%; lawyer z=4 1.3% vs 0.1%; retail z=4 1.3% vs 0.1%; colors z=4 1.3% vs 0.1%; cubs z=3 1.3% vs 0.1%; bowl z=3 1.3% vs 0.1%; heart z=3 1.3% vs 0.1%; francis z=3 1.3% vs 0.1%; defined z=3 1.3% vs 0.1%; faith z=3 1.3% vs 0.1%; records z=3 2.0% vs 0.2%; profession z=3 1.3% vs 0.1%; white z=3 2.0% vs 0.2%; files z=3 1.3% vs 0.1%; 500 z=3 1.3% vs 0.1%; anti z=3 1.3% vs 0.1%; address z=3 1.3% vs 0.1%; jaws z=3 0.7% vs 0.0%; investor z=3 0.7% vs 0.0%
Words, p22: mistake z=5 1.4% vs 0.0%; strategy z=4 1.4% vs 0.0%; regards z=4 1.4% vs 0.1%; coast z=3 1.4% vs 0.1%; 1988 z=3 1.4% vs 0.1%; zone z=3 1.4% vs 0.1%; 1947 z=3 0.7% vs 0.0%; rebel z=3 0.7% vs 0.0%; katrina z=3 0.7% vs 0.0%; sells z=3 0.7% vs 0.0%; everyday z=3 0.7% vs 0.0%; judaism z=3 0.7% vs 0.0%; feeling z=3 0.7% vs 0.0%; shelters z=3 0.7% vs 0.0%; showtime z=3 0.7% vs 0.0%; wit z=3 0.7% vs 0.0%; exhibit z=3 0.7% vs 0.0%; alcoholic z=3 0.7% vs 0.0%; arch z=3 0.7% vs 0.0%; strasbourg z=3 0.7% vs 0.0%
Words, p3: refer z=4 0.4% vs 0.1%; consumer z=4 0.3% vs 0.0%; spot z=4 0.3% vs 0.0%; frequency z=4 0.4% vs 0.1%; allies z=4 0.2% vs 0.0%; boston z=4 0.6% vs 0.2%; liam z=3 0.2% vs 0.0%; pre z=3 0.4% vs 0.1%; consonants z=3 0.2% vs 0.0%; infringement z=3 0.2% vs 0.0%; appear z=3 0.5% vs 0.1%; higher z=3 0.4% vs 0.1%; aggression z=3 0.2% vs 0.0%; kid z=3 0.2% vs 0.0%; aspirated z=3 0.2% vs 0.0%; allowing z=3 0.2% vs 0.0%; acting z=3 0.3% vs 0.0%; spoke z=3 0.2% vs 0.0%; details z=3 0.2% vs 0.0%; 1857 z=3 0.2% vs 0.0%
Words, p4: student z=4 0.6% vs 0.2%; excessive z=4 0.3% vs 0.0%; friars z=3 0.2% vs 0.0%; clubs z=3 0.2% vs 0.0%; wider z=3 0.2% vs 0.0%; emerge z=3 0.2% vs 0.0%; combine z=3 0.2% vs 0.0%; rhine z=3 0.2% vs 0.0%; innovator z=3 0.2% vs 0.0%; 1891 z=3 0.2% vs 0.0%; divisions z=3 0.3% vs 0.0%; isles z=3 0.3% vs 0.0%; sand z=3 0.3% vs 0.0%; torch z=3 0.3% vs 0.1%; bern z=3 0.2% vs 0.0%; attempted z=3 0.2% vs 0.0%; imperial's z=3 0.2% vs 0.0%; collaborate z=3 0.2% vs 0.0%; settlements z=3 0.2% vs 0.0%; coverage z=3 0.3% vs 0.1%
Words, p5: qing z=4 0.4% vs 0.1%; endangered z=4 0.3% vs 0.0%; relationship z=4 0.5% vs 0.1%; vote z=3 0.4% vs 0.1%; needed z=3 0.3% vs 0.1%; throughout z=3 0.3% vs 0.1%; graphics z=3 0.2% vs 0.0%; rule z=3 0.6% vs 0.2%; vertical z=3 0.2% vs 0.0%; relay z=3 0.3% vs 0.0%; torch z=3 0.3% vs 0.1%; pain z=3 0.4% vs 0.1%; increase z=3 0.4% vs 0.1%; sequel z=3 0.2% vs 0.0%; chuck z=3 0.2% vs 0.0%; households z=3 0.2% vs 0.0%; mtv z=3 0.2% vs 0.0%; displayed z=3 0.2% vs 0.0%; applications z=3 0.2% vs 0.0%; subjects z=3 0.3% vs 0.1%
Words, p6: hd z=4 0.3% vs 0.0%; control z=4 0.7% vs 0.2%; 000 z=4 0.5% vs 0.1%; server z=4 0.3% vs 0.0%; montini z=4 0.3% vs 0.0%; rule z=4 0.7% vs 0.2%; half z=4 0.5% vs 0.1%; foreign z=3 0.5% vs 0.1%; returned z=3 0.2% vs 0.0%; can't z=3 0.3% vs 0.1%; brown z=3 0.4% vs 0.1%; germany z=3 0.5% vs 0.1%; need z=3 2.4% vs 1.4%; 24 z=3 0.3% vs 0.1%; hours z=3 0.3% vs 0.1%; forget z=3 0.3% vs 0.1%; eliminate z=3 0.2% vs 0.0%; accredited z=3 0.2% vs 0.0%; kerry's z=3 0.2% vs 0.0%; suburban z=3 0.2% vs 0.0%
Words, p7: it z=4 4.4% vs 2.7%; posses z=4 0.2% vs 0.0%; couple z=3 0.3% vs 0.0%; frequencies z=3 0.2% vs 0.0%; nationalist z=3 0.2% vs 0.0%; handle z=3 0.4% vs 0.1%; breaks z=3 0.3% vs 0.0%; younger z=3 0.3% vs 0.0%; philosopher z=3 0.4% vs 0.1%; photographer z=3 0.2% vs 0.0%; warm z=3 0.2% vs 0.0%; read z=3 0.5% vs 0.1%; biggest z=3 0.5% vs 0.1%; religions z=3 0.3% vs 0.0%; okay z=3 0.2% vs 0.0%; asked z=3 0.2% vs 0.0%; invaded z=3 0.2% vs 0.0%; get z=3 2.4% vs 1.4%; constitutional z=3 0.3% vs 0.0%; concerns z=3 0.3% vs 0.0%
Words, p8: attempt z=5 0.7% vs 0.1%; makes z=4 0.8% vs 0.2%; becoming z=4 0.3% vs 0.0%; concentrated z=4 0.3% vs 0.0%; order z=4 1.0% vs 0.3%; currency z=4 0.4% vs 0.1%; mammals z=4 0.3% vs 0.0%; membership z=3 0.3% vs 0.0%; crimes z=3 0.3% vs 0.0%; silent z=3 0.3% vs 0.0%; neighborhood z=3 0.3% vs 0.0%; southern z=3 0.7% vs 0.2%; humanism z=3 0.3% vs 0.1%; winner z=3 0.4% vs 0.1%; fiction z=3 0.6% vs 0.2%; claim z=3 0.9% vs 0.3%; lands z=3 0.3% vs 0.0%; 1925 z=3 0.3% vs 0.0%; green z=3 0.5% vs 0.1%; islam z=3 0.3% vs 0.1%
Words, p9: intake z=4 0.3% vs 0.0%; northwestern z=4 0.5% vs 0.1%; cricketer z=4 0.3% vs 0.0%; fibers z=4 0.3% vs 0.0%; ad z=4 0.4% vs 0.0%; future z=3 0.5% vs 0.1%; orders z=3 0.3% vs 0.0%; nigeria z=3 0.5% vs 0.1%; firm z=3 0.3% vs 0.0%; organized z=3 0.3% vs 0.0%; let z=3 0.3% vs 0.0%; missing z=3 0.3% vs 0.0%; critical z=3 0.4% vs 0.1%; printing z=3 0.2% vs 0.0%; chapel z=3 0.2% vs 0.0%; unincorporated z=3 0.2% vs 0.0%; protocols z=3 0.2% vs 0.0%; customs z=3 0.2% vs 0.0%; interdisciplinary z=3 0.2% vs 0.0%; 1839 z=3 0.2% vs 0.0%
2–4-word phrases, <listed option>: company based z=4 0.4% vs 0.0%; of solar z=4 0.4% vs 0.0%; was previously z=4 0.4% vs 0.0%; how long has z=4 0.4% vs 0.0%; long has z=4 0.4% vs 0.0%; what person z=4 0.4% vs 0.0%; during a z=4 0.4% vs 0.0%; is the earliest z=4 0.4% vs 0.0%; the earliest z=4 0.6% vs 0.1%; company based in z=4 0.3% vs 0.0%; of prince z=4 0.3% vs 0.0%; the biggest z=4 0.5% vs 0.1%; what property z=4 0.3% vs 0.0%; a report z=4 0.3% vs 0.0%; is it called z=4 0.4% vs 0.0%; what is it called z=4 0.4% vs 0.0%; tertiary education z=4 0.3% vs 0.0%; a video game z=4 0.3% vs 0.0%; game released z=4 0.3% vs 0.0%; how many people in z=4 0.3% vs 0.0%
2–4-word phrases, none: is there z=14 3.3% vs 0.4%; is there a z=14 3.0% vs 0.3%; there a z=13 3.0% vs 0.3%; does the z=9 5.1% vs 2.0%; available for z=8 1.0% vs 0.0%; in which z=7 3.5% vs 1.5%; by an z=6 0.9% vs 0.2%; a discount z=6 0.7% vs 0.0%; was born z=6 1.3% vs 0.3%; policy cover z=6 0.6% vs 0.0%; by a z=6 1.5% vs 0.4%; are there any z=6 0.6% vs 0.0%; the api z=6 0.5% vs 0.0%; are there z=6 1.1% vs 0.3%; there any z=6 0.6% vs 0.1%; is a z=6 4.2% vs 2.0%; in what z=6 6.5% vs 3.6%; discount for z=5 0.5% vs 0.0%; born in z=5 1.3% vs 0.4%; the film z=5 1.4% vs 0.5%
2–4-word phrases, p1: the las vegas z=4 0.2% vs 0.0%; the las z=4 0.2% vs 0.0%; of the federal z=4 0.2% vs 0.0%; to drive z=3 0.2% vs 0.0%; who commanded the z=3 0.2% vs 0.0%; who commanded z=3 0.2% vs 0.0%; can cause z=3 0.2% vs 0.0%; to be z=3 2.1% vs 1.2%; las vegas z=3 0.2% vs 0.0%; become the first z=3 0.2% vs 0.0%; team to win z=3 0.2% vs 0.0%; my plot z=3 0.2% vs 0.0%; chief executive officer of z=3 0.2% vs 0.0%; on my plot z=3 0.2% vs 0.0%; executive officer of z=3 0.2% vs 0.0%; and led z=3 0.2% vs 0.0%; american businessman z=3 0.2% vs 0.0%; has held z=3 0.2% vs 0.0%; of the first single z=3 0.2% vs 0.0%; that isn't z=3 0.2% vs 0.0%
2–4-word phrases, p10: paul vi z=4 0.6% vs 0.1%; what club z=4 0.4% vs 0.0%; maximum file z=4 0.4% vs 0.0%; the self titled z=4 0.4% vs 0.0%; did paul vi z=4 0.4% vs 0.0%; in the 1990s z=4 0.4% vs 0.0%; an issue z=4 0.4% vs 0.0%; of representatives z=3 0.4% vs 0.0%; in the 19th z=3 0.4% vs 0.0%; the self z=3 0.4% vs 0.0%; house of representatives z=3 0.4% vs 0.0%; self titled z=3 0.4% vs 0.0%; i am z=3 0.5% vs 0.1%; what happens to z=3 0.6% vs 0.1%; did paul z=3 0.4% vs 0.0%; the 1990s z=3 0.4% vs 0.0%; of one z=3 0.4% vs 0.0%; and author z=3 0.4% vs 0.0%; happens to z=3 0.6% vs 0.1%; was marked z=3 0.2% vs 0.0%
2–4-word phrases, p11: long can z=5 1.0% vs 0.1%; how long can z=5 1.0% vs 0.1%; decade was z=4 0.4% vs 0.0%; is the number z=4 0.4% vs 0.0%; how quickly z=4 0.9% vs 0.1%; in 1975 z=4 0.4% vs 0.0%; when does the z=4 0.9% vs 0.2%; can a z=4 0.7% vs 0.1%; population in z=3 0.4% vs 0.0%; that was used z=3 0.3% vs 0.0%; in what period z=3 0.3% vs 0.0%; is the number of z=3 0.3% vs 0.0%; which doctrine z=3 0.3% vs 0.0%; to keep the z=3 0.3% vs 0.0%; which spanish z=3 0.3% vs 0.0%; file is z=3 0.3% vs 0.0%; of city z=3 0.3% vs 0.0%; my plot z=3 0.3% vs 0.0%; what is the number z=3 0.3% vs 0.0%; what field z=3 0.3% vs 0.0%
2–4-word phrases, p12: penalty if z=4 0.5% vs 0.0%; what song was z=4 0.5% vs 0.0%; need to pass z=4 0.5% vs 0.0%; on november z=4 0.5% vs 0.0%; released by z=4 0.7% vs 0.1%; to pass z=4 0.5% vs 0.0%; song was z=4 0.5% vs 0.0%; the wife z=4 0.5% vs 0.0%; the wife of z=4 0.5% vs 0.0%; in miami z=4 0.5% vs 0.0%; a limit on how z=4 0.5% vs 0.0%; there a limit on z=4 0.5% vs 0.0%; aircraft that z=3 0.3% vs 0.0%; the dead z=3 0.3% vs 0.0%; i need to pass z=3 0.3% vs 0.0%; in when z=3 0.3% vs 0.0%; in mathematics z=3 0.3% vs 0.0%; the book the z=3 0.3% vs 0.0%; its obligations z=3 0.3% vs 0.0%; a native z=3 0.3% vs 0.0%
2–4-word phrases, p13: of all the z=5 0.6% vs 0.0%; name of the book z=4 0.6% vs 0.0%; the major z=4 0.8% vs 0.1%; was the largest z=4 0.6% vs 0.0%; the pacific z=4 0.4% vs 0.0%; credit for z=4 0.4% vs 0.0%; like a z=4 0.4% vs 0.0%; what is the earliest z=4 0.4% vs 0.0%; i get if my z=4 0.4% vs 0.0%; proto indo z=4 0.4% vs 0.0%; hotel in z=4 0.4% vs 0.0%; about how much of z=4 0.4% vs 0.0%; the cincinnati z=4 0.4% vs 0.0%; at any z=4 0.4% vs 0.0%; get if my z=4 0.4% vs 0.0%; apply for the z=3 0.4% vs 0.0%; about how much z=3 0.4% vs 0.0%; of the major z=3 0.4% vs 0.0%; american drama film z=3 0.4% vs 0.0%; the people in z=3 0.4% vs 0.0%
2–4-word phrases, p14: prior to z=4 1.2% vs 0.2%; result of z=4 1.0% vs 0.1%; the result of z=4 0.6% vs 0.0%; band from z=4 0.6% vs 0.0%; record high z=4 0.4% vs 0.0%; update my z=4 0.4% vs 0.0%; the plant z=4 0.4% vs 0.0%; look at z=4 0.4% vs 0.0%; in 1860 z=4 0.4% vs 0.0%; which season z=4 0.4% vs 0.0%; stage name z=4 0.4% vs 0.0%; was the first president z=4 0.4% vs 0.0%; what colors z=4 0.4% vs 0.0%; will you z=4 0.4% vs 0.0%; into a z=3 0.6% vs 0.1%; on time z=3 0.4% vs 0.0%; the actor born z=3 0.4% vs 0.0%; was the highest z=3 0.4% vs 0.0%; the chicago z=3 0.4% vs 0.0%; the result z=3 0.6% vs 0.1%
2–4-word phrases, p15: a 2005 z=4 0.7% vs 0.0%; the airport z=4 0.7% vs 0.0%; of valencia z=4 0.5% vs 0.0%; type of climate z=4 0.5% vs 0.0%; year at z=4 0.5% vs 0.0%; year at the z=4 0.5% vs 0.0%; what type of climate z=4 0.5% vs 0.0%; also played for z=4 0.5% vs 0.0%; else in z=4 0.5% vs 0.0%; all india z=4 0.5% vs 0.0%; who worked with z=4 0.5% vs 0.0%; decrease in z=4 0.5% vs 0.0%; organization was z=4 0.5% vs 0.0%; be considered z=4 0.7% vs 0.1%; how much in z=4 0.5% vs 0.0%; american film director z=4 0.5% vs 0.0%; what facet z=4 0.5% vs 0.0%; where else z=4 0.5% vs 0.0%; much in z=4 0.5% vs 0.0%; facet of z=4 0.5% vs 0.0%
2–4-word phrases, p16: what country z=5 2.4% vs 0.4%; a singer z=4 0.8% vs 0.0%; the swiss z=4 0.8% vs 0.0%; what role did z=4 0.8% vs 0.0%; role did z=4 0.8% vs 0.1%; how many passengers z=4 0.5% vs 0.0%; high middle z=4 0.5% vs 0.0%; what level z=4 0.5% vs 0.0%; sale of z=4 0.5% vs 0.0%; to and z=4 0.5% vs 0.0%; give for z=4 0.5% vs 0.0%; the sale z=4 0.5% vs 0.0%; what level of z=4 0.5% vs 0.0%; rejection of z=4 0.5% vs 0.0%; many passengers z=4 0.5% vs 0.0%; high middle ages z=4 0.5% vs 0.0%; the sale of z=4 0.5% vs 0.0%; much notice do i z=4 0.8% vs 0.1%; notice do i need z=4 0.8% vs 0.1%; notice do i z=4 0.8% vs 0.1%
2–4-word phrases, p17: start in z=4 0.6% vs 0.0%; which singer z=4 0.6% vs 0.0%; that he z=4 0.6% vs 0.0%; performing arts z=4 0.6% vs 0.0%; does seattle z=4 0.6% vs 0.0%; in 1999 z=4 0.6% vs 0.0%; mid 20th z=4 0.6% vs 0.0%; footballer who plays z=4 0.6% vs 0.0%; mid 20th century z=4 0.6% vs 0.0%; what are these z=4 0.6% vs 0.0%; how many times did z=4 0.6% vs 0.0%; many times did z=4 0.6% vs 0.0%; developed in z=4 0.6% vs 0.0%; a performance z=4 0.6% vs 0.0%; roles as z=4 0.6% vs 0.0%; times did z=4 0.6% vs 0.0%; are more z=4 0.6% vs 0.0%; in the 1990s z=3 0.6% vs 0.0%; went to z=3 0.6% vs 0.0%; the same z=3 1.9% vs 0.4%
2–4-word phrases, p18: in the past z=5 0.9% vs 0.0%; the past z=5 0.9% vs 0.0%; should i throw z=4 0.6% vs 0.0%; it has been z=4 0.6% vs 0.0%; to enact z=4 0.6% vs 0.0%; of northern z=4 0.6% vs 0.0%; originally called z=4 0.6% vs 0.0%; should i throw away z=4 0.6% vs 0.0%; the nickname z=4 1.2% vs 0.1%; were killed z=4 0.9% vs 0.1%; above the z=4 0.6% vs 0.0%; of the 20th century z=4 0.6% vs 0.0%; of healing z=4 0.6% vs 0.0%; book of healing z=4 0.6% vs 0.0%; was george z=4 0.6% vs 0.0%; the nickname of z=4 0.9% vs 0.1%; for when z=4 0.6% vs 0.0%; why has z=4 0.6% vs 0.0%; i throw z=4 0.6% vs 0.0%; throw away z=4 0.6% vs 0.0%
2–4-word phrases, p19: the xbox z=4 0.8% vs 0.0%; population was z=4 0.8% vs 0.0%; into which z=4 0.8% vs 0.0%; featured in the z=4 1.2% vs 0.1%; the most recent z=4 0.8% vs 0.0%; of large z=4 0.8% vs 0.0%; france and z=4 0.8% vs 0.0%; do the z=4 2.4% vs 0.4%; of mary z=4 0.8% vs 0.0%; most recent z=4 0.8% vs 0.0%; the brother of z=4 0.8% vs 0.0%; times was z=4 0.8% vs 0.0%; prime minister of z=4 0.8% vs 0.0%; the brother z=4 0.8% vs 0.0%; at the time of z=3 0.8% vs 0.0%; the city's z=3 0.8% vs 0.0%; is the oldest z=3 0.8% vs 0.0%; the leader of the z=3 0.8% vs 0.0%; how long did the z=3 0.8% vs 0.1%; long did the z=3 0.8% vs 0.1%
2–4-word phrases, p2: to give z=4 0.4% vs 0.1%; take place z=4 0.7% vs 0.2%; the 19th century z=3 0.4% vs 0.1%; are found z=3 0.2% vs 0.0%; the border z=3 0.2% vs 0.0%; get my money z=3 0.4% vs 0.1%; get my money back z=3 0.4% vs 0.1%; came to z=3 0.3% vs 0.0%; who decides z=3 0.2% vs 0.0%; of income z=3 0.2% vs 0.0%; of bermuda's z=3 0.2% vs 0.0%; the main character of z=3 0.2% vs 0.0%; main character of z=3 0.2% vs 0.0%; km north z=3 0.2% vs 0.0%; almost all z=3 0.2% vs 0.0%; replaced the z=3 0.2% vs 0.0%; who stated that z=3 0.2% vs 0.0%; was dedicated to z=3 0.2% vs 0.0%; who stated z=3 0.2% vs 0.0%; who was the father z=3 0.2% vs 0.0%
2–4-word phrases, p20: the notre dame z=4 1.0% vs 0.0%; the assistant z=4 1.0% vs 0.0%; the defeat z=4 1.0% vs 0.0%; take part z=4 1.0% vs 0.0%; the notre z=4 1.0% vs 0.0%; take part in z=4 1.0% vs 0.0%; the defeat of z=4 1.0% vs 0.0%; are in the z=4 1.4% vs 0.1%; defeat of z=4 1.0% vs 0.0%; how large is z=4 1.0% vs 0.0%; of the new york z=4 1.0% vs 0.0%; in canada z=4 1.0% vs 0.0%; large is z=4 1.0% vs 0.0%; what will z=4 1.0% vs 0.0%; was known for z=4 1.0% vs 0.0%; up a z=3 1.0% vs 0.0%; are in z=3 1.9% vs 0.2%; how large z=3 1.0% vs 0.1%; co wrote z=3 1.0% vs 0.1%; the new york z=3 1.4% vs 0.1%
2–4-word phrases, p21: why would z=4 1.3% vs 0.0%; anti aircraft z=4 1.3% vs 0.0%; is written z=4 1.3% vs 0.0%; what in the z=4 1.3% vs 0.0%; who recorded the z=4 1.3% vs 0.0%; who recorded z=4 1.3% vs 0.0%; a failed z=4 1.3% vs 0.0%; recorded the z=4 1.3% vs 0.0%; home of the z=4 1.3% vs 0.0%; what channel z=4 1.3% vs 0.0%; my own z=4 1.3% vs 0.0%; the cubs z=4 1.3% vs 0.1%; the science z=4 1.3% vs 0.1%; home of z=4 1.3% vs 0.1%; how did z=3 2.6% vs 0.4%; why did the z=3 1.3% vs 0.1%; where are the z=3 1.3% vs 0.1%; the ruler of z=3 0.7% vs 0.0%; the restaurant z=3 0.7% vs 0.0%; which location z=3 0.7% vs 0.0%
2–4-word phrases, p22: known as a z=5 1.4% vs 0.0%; has there been z=4 1.4% vs 0.0%; has there z=4 1.4% vs 0.0%; there been z=4 1.4% vs 0.0%; in regards z=4 1.4% vs 0.0%; in regards to z=4 1.4% vs 0.0%; regards to z=4 1.4% vs 0.0%; was part of z=3 1.4% vs 0.1%; was part z=3 1.4% vs 0.1%; charge the battery z=3 0.7% vs 0.0%; served in the z=3 0.7% vs 0.0%; in french z=3 0.7% vs 0.0%; a 1985 z=3 0.7% vs 0.0%; university did z=3 0.7% vs 0.0%; a character named z=3 0.7% vs 0.0%; yongle emperor z=3 0.7% vs 0.0%; show hosted z=3 0.7% vs 0.0%; show hosted by z=3 0.7% vs 0.0%; in taiwan z=3 0.7% vs 0.0%; a response z=3 0.7% vs 0.0%
2–4-word phrases, p3: day was z=4 0.3% vs 0.0%; refer to z=4 0.4% vs 0.1%; start to z=4 0.2% vs 0.0%; child of z=4 0.2% vs 0.0%; what day z=3 0.4% vs 0.1%; what day was z=3 0.2% vs 0.0%; secretary of state z=3 0.2% vs 0.0%; took place on z=3 0.2% vs 0.0%; did the united z=3 0.2% vs 0.0%; a branch z=3 0.2% vs 0.0%; how much was the z=3 0.2% vs 0.0%; cap on z=3 0.2% vs 0.0%; to london z=3 0.2% vs 0.0%; much was the z=3 0.2% vs 0.0%; died from z=3 0.2% vs 0.0%; what tradition z=3 0.2% vs 0.0%; acting in z=3 0.2% vs 0.0%; did general z=3 0.2% vs 0.0%; is another name for z=3 0.4% vs 0.1%; on how much z=3 0.3% vs 0.0%
2–4-word phrases, p4: life of z=4 0.3% vs 0.0%; played at z=3 0.3% vs 0.0%; working on z=3 0.2% vs 0.0%; torch relay z=3 0.2% vs 0.0%; a university z=3 0.3% vs 0.1%; adult contemporary z=3 0.3% vs 0.0%; in western z=3 0.2% vs 0.0%; hit the z=3 0.2% vs 0.0%; was featured on z=3 0.2% vs 0.0%; attempted to z=3 0.2% vs 0.0%; did the u s z=3 0.2% vs 0.0%; did the u z=3 0.2% vs 0.0%; how is z=3 0.6% vs 0.2%; type of z=3 2.6% vs 1.6%; many people were z=3 0.4% vs 0.1%; how many people were z=3 0.4% vs 0.1%; and sand z=3 0.2% vs 0.0%; coverage of z=3 0.2% vs 0.0%; being what z=3 0.2% vs 0.0%; a period of z=3 0.2% vs 0.0%
2–4-word phrases, p5: if an z=4 0.2% vs 0.0%; produced by z=4 0.7% vs 0.2%; of the current z=4 0.2% vs 0.0%; from a z=4 0.7% vs 0.3%; behind the z=3 0.2% vs 0.0%; other artists z=3 0.2% vs 0.0%; activity in z=3 0.2% vs 0.0%; is given to z=3 0.2% vs 0.0%; of plymouth z=3 0.2% vs 0.0%; did the qing z=3 0.2% vs 0.0%; of the roman z=3 0.2% vs 0.0%; if i don't pick z=3 0.2% vs 0.0%; i don't pick z=3 0.2% vs 0.0%; don't pick z=3 0.2% vs 0.0%; carry the z=3 0.2% vs 0.0%; song for z=3 0.2% vs 0.0%; war of the z=3 0.2% vs 0.0%; the qing z=3 0.3% vs 0.0%; the torch z=3 0.3% vs 0.0%; to many z=3 0.2% vs 0.0%
2–4-word phrases, p6: control of z=4 0.4% vs 0.1%; happens if i forget z=4 0.3% vs 0.0%; i need z=4 2.1% vs 1.1%; do i z=4 3.7% vs 2.3%; for creating z=4 0.2% vs 0.0%; the us air z=4 0.2% vs 0.0%; the us air force z=4 0.2% vs 0.0%; forget to z=4 0.3% vs 0.1%; i forget to z=4 0.3% vs 0.1%; if i forget to z=4 0.3% vs 0.1%; if i forget z=3 0.3% vs 0.1%; i forget z=3 0.3% vs 0.1%; host the z=3 0.2% vs 0.0%; i need to z=3 1.7% vs 0.9%; first in z=3 0.2% vs 0.0%; september 2014 z=3 0.2% vs 0.0%; renew my z=3 0.2% vs 0.0%; cost if z=3 0.2% vs 0.0%; giving the z=3 0.2% vs 0.0%; returned for z=3 0.2% vs 0.0%
2–4-word phrases, p7: when does z=4 0.9% vs 0.3%; couple of z=4 0.3% vs 0.0%; a couple of z=4 0.3% vs 0.0%; was the founder z=4 0.3% vs 0.0%; was the founder of z=4 0.3% vs 0.0%; to get z=4 0.9% vs 0.3%; a couple z=4 0.3% vs 0.0%; does my z=4 0.5% vs 0.1%; write in z=3 0.2% vs 0.0%; can my z=3 0.4% vs 0.1%; was the leader z=3 0.2% vs 0.0%; who was the leader z=3 0.2% vs 0.0%; have to use z=3 0.2% vs 0.0%; was the leader of z=3 0.2% vs 0.0%; present day z=3 0.3% vs 0.0%; all the z=3 0.4% vs 0.1%; top level z=3 0.2% vs 0.0%; go into z=3 0.2% vs 0.0%; writer who z=3 0.2% vs 0.0%; who was the founder z=3 0.2% vs 0.0%
2–4-word phrases, p8: the time of the z=4 0.3% vs 0.0%; time of the z=4 0.3% vs 0.0%; at the time z=4 0.4% vs 0.0%; of humanism z=4 0.3% vs 0.0%; the winner of the z=4 0.3% vs 0.0%; at the time of z=4 0.3% vs 0.0%; in india in z=4 0.3% vs 0.0%; winner of z=4 0.4% vs 0.1%; what percent of z=4 0.6% vs 0.1%; what is one of z=4 0.4% vs 0.1%; the time of z=4 0.4% vs 0.1%; how long are z=4 0.3% vs 0.0%; is played by z=4 0.3% vs 0.0%; india in z=4 0.3% vs 0.0%; long are z=4 0.3% vs 0.0%; a show z=3 0.3% vs 0.0%; the winner of z=3 0.3% vs 0.0%; winner of the z=3 0.3% vs 0.0%; the winner z=3 0.3% vs 0.0%; focus of z=3 0.3% vs 0.0%
2–4-word phrases, p9: one country z=4 0.4% vs 0.0%; law review z=4 0.3% vs 0.0%; to cancel my z=4 0.3% vs 0.0%; was the most z=4 0.5% vs 0.1%; which city was z=3 0.3% vs 0.0%; had to z=3 0.3% vs 0.0%; to cancel z=3 0.4% vs 0.0%; in 2017 z=3 0.3% vs 0.0%; what does the term z=3 0.3% vs 0.0%; city was z=3 0.6% vs 0.1%; and two z=3 0.3% vs 0.0%; was william z=3 0.3% vs 0.0%; the book of z=3 0.3% vs 0.0%; does the term z=3 0.3% vs 0.0%; agree to z=3 0.3% vs 0.0%; how can z=3 0.5% vs 0.1%; city was the z=3 0.4% vs 0.1%; loosely based on z=3 0.3% vs 0.0%; the largest city z=3 0.3% vs 0.0%; loosely based z=3 0.3% vs 0.0%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- none
Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):
- all rows (25,380 rows): none
Shortcut models
Predicting the label class on test (2,080 rows). Chance 4.2%, majority class ('none') 15.0%; balanced chance 4.2%.
| model (logistic regression, trained on the train sample) | test accuracy | balanced accuracy (mean recall) |
|---|---|---|
| bag of words, whole state (words and word pairs) | 14.0% | 5.1% |
bag of words, main text only (question) |
14.0% | 5.1% |
| surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) | 18.1% | 7.8% |
| surface features of the main text only | 15.1% | 4.3% |
Strongest single surface features (logistic regression on one feature, balanced accuracy on test):
| feature | accuracy | balanced accuracy |
|---|---|---|
| n_options(log) | 17.9% | 7.6% |
| chars(log) | 15.1% | 4.3% |
| state_chars(log) | 15.1% | 4.3% |
| words(log) | 15.1% | 4.2% |
| count_! | 15.0% | 4.2% |
| upper_ratio | 15.0% | 4.2% |
| digit_ratio | 15.0% | 4.2% |
| nonascii_ratio | 15.0% | 4.2% |
| nonlatin | 15.0% | 4.2% |
| emoji | 15.0% | 4.2% |
Other state fields alone (predicting the label class on test from one field, without the main text):
| field | treated as | accuracy | balanced accuracy |
|---|
No-meaning option picker: a logistic ranker scores each option from its position, length, key type, fixed-option identity and shape (commas, brackets, capitals), never reading the state or the option's words, and picks the top option per row.
- Test accuracy 15.0% against uniform chance 7.9% (this includes the fixed options, whose share is a class prior).
- Among rows whose answer is a listed option (123), picking only among listed options: 10.6% against chance 11.9%.
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 5,123 (20.2%); groups: 5,107; largest group 6.
- Train rows identical in the whole prompt (state, options, instructions): 0.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 4580 groups (9176 rows). (Can be legitimate when the rest of the state or the options differ.)
- "What are the most common side effects?" ×6: p2 2, p7 2, p1 1, p13 1
- "How long does my API key last before it expires?" ×4: p6 1, p10 1, p7 1, p8 1
- "Who published A Theory of Justice?" ×4: p9 1, p1 1, p11 1, p10 1
- "What is the most common side effect?" ×4: p1 2, p4 1, p3 1
- "How long can I keep the bottle after opening it?" ×4: p8 1, p11 1, p4 1, p15 1
- "How much does it cost to print a color page?" ×4: p3 2, p1 1, p9 1
- "How quickly will I get a response if I submit a support ticket?" ×3: p22 1, p1 1, p5 1
- "How long can I keep the bottle once I've opened it?" ×3: p3 1, p12 1, p7 1
Most repeated main texts in train:
- ×6: "what are the most common side effects?" (p2 2, p7 2, p1 1)
- ×4: "who published a theory of justice?" (p9 1, p1 1, p11 1)
- ×4: "what is the most common side effect?" (p1 2, p4 1, p3 1)
- ×4: "how much does it cost to print a color page?" (p3 2, p1 1, p9 1)
- ×4: "how long does my api key last before it expires?" (p6 1, p10 1, p7 1)
- ×4: "how long can i keep the bottle after opening it?" (p8 1, p11 1, p4 1)
- ×3: "how quickly will i get a response if i submit a support ticket?" (p22 1, p1 1, p5 1)
- ×3: "how long can i keep the bottle once i've opened it?" (p3 1, p12 1, p7 1)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 0 (0.0%) | |
| calibration | 0 (0.0%) | |
| test | 0 (0.0%) |
Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)
- Train: 5,169 near-duplicate pairs; 10,245 rows (40.4%) sit in 5,111 clusters; largest cluster 6; excess rows (cluster size − 1) 5,134 (20.2%).
- Clusters with more than one label class: 4,584 (9,190 rows).
- ×6: "What are the most common side effects?" → p2 2, p7 2, p1 1, p13 1
- ×4: "How long does my API key last before it expires?" → p6 1, p10 1, p7 1, p8 1
- ×4: "Who published A Theory of Justice?" → p9 1, p1 1, p11 1, p10 1
- ×4: "What is the most common side effect?" → p1 2, p4 1, p3 1
- ×4: "How long can I keep the bottle after opening it?" → p8 1, p11 1, p4 1, p15 1
- Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 1 (0.1%), test 1 (0.0%)
- train "How often does the software check my license online?" ~ calibration "How often does the software check my license?" (J=0.86)
- train "What should I do if I forget to take my daily dose?" ~ test "What should I do if I forget to take my daily capsule?" (J=0.82)
- train "How often does the software check my license online?" ~ calibration "How often does the software check my license?" (J=0.86)
Largest train clusters:
- ×6: "What are the most common side effects?"
- ×4: "How long does my API key last before it expires?"
- ×4: "Who published A Theory of Justice?"
- ×4: "What is the most common side effect?"
- ×4: "How long can I keep the bottle after opening it?"
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 0 |
| dev | 0 | 0 |
| calibration | 0 | 0 |
| test | 0 | 0 |
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 0 (0.0%) | |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 0 (0.0%) | |
| chat preamble ('Here is/are...', 'Sure!') | 0 (0.0%) | |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 0 (0.0%) | |
| encoding garbage (mojibake/replacement char) | 0 (0.0%) | |
| HTML tag | 0 (0.0%) | |
| HTML entity | 0 (0.0%) | |
| base64-like run (40+ chars) | 0 (0.0%) | |
| URL | 0 (0.0%) |
- Possibly cut off: 0 of 148 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: <listed option> 0.0%, none 0.0%, p1 0.0%, p10 0.0%, p11 0.0%, p12 0.0%, p13 0.0%, p14 0.0%, p15 0.0%, p16 0.0%, p17 0.0%, p18 0.0%, p2 0.0%, p20 0.0%, p3 0.0%, p4 0.0%, p5 0.0%, p6 0.0%, p7 0.0%, p8 0.0%, p9 0.0%
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- none
6. Samples
20 random train rows per kind: ground-rerank-samples.txt. Reading notes are in the findings above.
