Skip to content
JeffHub

QA report: ground

The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.

The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.

QA: ground (v4)

Verdict: PASS WITH NOTES. See ground-grounding.md (verdict and notes) and ground-rerank.md. QA-OK-ground written.

QA: ground-grounding

Checked 2026-09-30 13:26 by adapters/qa/qa.py (READY file READY-ground, v4 (held-out second opinion; grounding answers balanced on surface features, 2026-09-30)).

Verdict: PASS WITH NOTES

Ground v4 (READY-ground 2026-09-30 13:19, train sha256 c8b38f85…15fb2; train 50,760 / dev 1,800 / calibration 1,440 / test 4,160). The v3 check is in ground-v3-summary.md. This report covers both question types: grounding (ground-grounding.md) and re-ranking (ground-rerank.md).

The v3 finding is fixed (checked with ground_extra.py and qa.py)

  • Answer length and shape are equal across labels:
    • median answer is 28–29 words for every label in train, and 32 for every label in test;
    • ", " appears in 68.2% of answers and " and " in 41.5%, identical for every label (test: 79.8% and 45.2%).
  • Answer-only models are now near chance:
    • surface-only model: 28.4% balanced accuracy (was 35.7%; chance 25%), and 37.1% accuracy against a 35% majority baseline;
    • answer bag of words: 33.7% (was 40.7%);
    • whole state (answer plus sources): 28.4%.
  • The standard phrase check finds no flags for either question type.
  • Near duplicates across splits: 0–0.1% of held-out rows.

Notes (for the model card; meaning or small)

  1. Remaining answer-word signal (bag of words 33.7% against 25% chance):

    • negation ("not") appears in 8.0% of partly-supported answers against 15–16% of the others;
    • digits appear in 60% of contradicted answers against 47% of supported ones (number-changed rows);
    • hedged added details appear in partly-supported answers.

    These follow from how each label is made; the grounding agent tried adding negation and hedge cells, and it did not help. No single cue meets the 2%-and-mostly-one-label rule.

  2. Sources are about 20% shorter for partly-supported and unsupported rows (median 1,511–1,533 characters, against 1,834–1,899), because a gold passage is removed. The length of the sources alone gives 27.3% balanced accuracy, near chance.

  3. Re-ranking is unchanged and passes. The gold passage is the longest 7.8% of the time against 9.1% chance, its position is uniform, "none" rows have the same option count as the others, and a surface-only model reaches 7.8% balanced against 4.2% chance (from option count). 7,977 train rows are second pools (accepted earlier).

  4. The grounding agent labelled "supported + a claim whose passage was removed" as partly_supported, which is right by the label definitions; I had proposed labelling it unsupported.

  5. Share-alike licences (SQuAD, HotpotQA, FEVER) and hosted-teacher text are covered in the data report.

Automatic flags (for the reviewer to judge; not all are problems)

  • 146 strong phrase flags (see list): review whether they are meaning or leakage

Data checked

split rows families file
train 25,380 4250 train.jsonl
dev 900 128 dev.jsonl
calibration 720 94 calibration.jsonl
test 2,080 310 test.jsonl
  • Train sha256: c8b38f85a728f5aa916c2b4b2cf18a8cd63567fde55572cc4b0f16b155915fb2 (READY says c8b38f85a728f5aa916c2b4b2cf18a8cd63567fde55572cc4b0f16b155915fb2: match)
  • Main text field (the text the phrase and length checks use): state.answer.
  • Label classes: contradicted, partly_supported, supported, unsupported (<listed option> = one of the per-row listed options such as t3 or o12). Row kinds (source.kind): grounding-docs-contradicted, grounding-docs-gold-removed, grounding-docs-joined-contradicted, grounding-docs-joined-supported, grounding-docs-joined-unsupported, grounding-docs-number-changed, grounding-docs-partly, grounding-docs-supported, grounding-docs-unanswerable, grounding-fever-combined, grounding-fever-contradicted, grounding-fever-joined-contradicted, grounding-fever-joined-supported, grounding-fever-joined-unsupported, grounding-fever-supported, grounding-fever-unsupported, grounding-hotpot-contradicted, grounding-hotpot-gold-removed, grounding-hotpot-number-changed, grounding-hotpot-one-removed, grounding-hotpot-partly, grounding-hotpot-supported, grounding-squad-combined, grounding-squad-contradicted, grounding-squad-gold-removed, grounding-squad-joined-contradicted, grounding-squad-joined-supported, grounding-squad-joined-unsupported, grounding-squad-number-changed, grounding-squad-partly, grounding-squad-supported, grounding-squad-unanswerable.

3. Balance

Label class share per split

lclass train dev calibration test train rows
contradicted 20.0% 20.0% 20.0% 20.0% 5,076
partly_supported 25.0% 25.0% 25.0% 25.0% 6,345
supported 35.0% 35.0% 35.0% 35.0% 8,883
unsupported 20.0% 20.0% 20.0% 20.0% 5,076

Row kind share per split

kind train dev calibration test train rows
grounding-docs-contradicted 3.7% 1.9% 2.1% 3.0% 932
grounding-docs-gold-removed 3.8% 3.9% 2.9% 4.6% 965
grounding-docs-joined-contradicted 2.3% 3.6% 3.3% 3.3% 579
grounding-docs-joined-supported 5.2% 4.8% 5.6% 6.2% 1,328
grounding-docs-joined-unsupported 2.8% 2.4% 3.1% 3.2% 707
grounding-docs-number-changed 1.8% 1.9% 2.1% 3.0% 452
grounding-docs-partly 5.8% 5.8% 6.1% 7.5% 1,470
grounding-docs-supported 7.7% 7.9% 5.8% 8.4% 1,959
grounding-docs-unanswerable 1.7% 1.1% 1.7% 1.9% 432
grounding-fever-combined 5.0% 2.8% 2.8% 2.9% 1,260
grounding-fever-contradicted 2.5% 1.3% 1.2% 1.1% 626
grounding-fever-joined-contradicted 1.5% 0.9% 1.0% 1.2% 382
grounding-fever-joined-supported 3.0% 2.1% 2.6% 2.3% 772
grounding-fever-joined-unsupported 1.7% 1.7% 1.7% 1.1% 429
grounding-fever-supported 3.9% 1.8% 1.2% 1.8% 992
grounding-fever-unsupported 2.3% 0.6% 0.6% 1.2% 579
grounding-hotpot-contradicted 1.5% 0.2% 0.3% 0.6% 379
grounding-hotpot-gold-removed 2.1% 0.2% 0.4% 0.3% 541
grounding-hotpot-number-changed 1.2% 0.3% 0.6% 0.6% 307
grounding-hotpot-one-removed 1.7% 0.6% 0.3% 0.7% 434
grounding-hotpot-partly 3.3% 0.6% 0.7% 0.9% 839
grounding-hotpot-supported 4.2% 1.3% 0.7% 1.8% 1,074
grounding-squad-combined 3.5% 5.7% 4.7% 4.1% 900
grounding-squad-contradicted 2.6% 3.3% 2.9% 2.4% 672
grounding-squad-gold-removed 3.0% 6.2% 4.2% 4.2% 767
grounding-squad-joined-contradicted 1.7% 2.7% 3.5% 2.3% 424
grounding-squad-joined-supported 4.3% 7.3% 8.3% 6.3% 1,097
grounding-squad-joined-unsupported 2.4% 3.8% 5.6% 3.3% 605
grounding-squad-number-changed 1.3% 3.9% 3.1% 2.5% 323
grounding-squad-partly 5.7% 9.7% 10.4% 9.0% 1,442
grounding-squad-supported 6.5% 9.8% 10.7% 8.3% 1,661
grounding-squad-unanswerable 0.2% 0.1% 0.0% 0.2% 51

4. Format

split row-level format problems
train none
dev none
calibration none
test none

Options per choice row

split min median p99 max
train 4 4 4 4
dev 4 4 4 4
calibration 4 4 4 4
test 4 4 4 4

Prompt length in tokens

split measure median p99 max > 8192
train source.input_tokens (25380/25380 rows) 593 1017 1301 0
dev source.input_tokens (900/900 rows) 576 972 1140 0
calibration source.input_tokens (720/720 rows) 593 991 1045 0
test source.input_tokens (2080/2080 rows) 596 1026 1219 0

State key sets (train)

keys rows
answer, sources 25,380 (100.0%)

Instructions (train)

  • Canonical (the most common text) 69.6%, reworded 27.5% (80 distinct rewordings), none 2.9%. Target about 70 / 27 / 3.
  • Canonical text: "Is the answer supported by the sources? Judge only from the sources, not from outside knowledge."
  • source instruction tag: canonical 69.6%, variant 27.5%, none 2.9%
class canonical none
contradicted 68.7% 2.8%
partly_supported 70.6% 3.1%
supported 69.3% 3.0%
unsupported 69.8% 2.7%

1. Shortcuts

Phrase statistics and models use a label-stratified sample of 25,380 train rows; models are scored on the full test file (2,080 rows).

Text length by label class (main text, characters)

split class rows p10 median p90 mean
train contradicted 5076 74 176 332 191
train partly_supported 6345 81 177 311 188
train supported 8883 74 178 343 194
train unsupported 5076 76 178 342 194
test contradicted 416 96 200 360 214
test partly_supported 520 99 199 316 208
test supported 728 96 202 364 217
test unsupported 416 97 202 370 218

By row kind (train): main-text length, length of the rest of the state, options

kind rows median chars mean chars median other-state chars median options
grounding-docs-contradicted 932 178 187 2175 4
grounding-docs-gold-removed 965 173 178 2138 4
grounding-docs-joined-contradicted 579 327 336 2549 4
grounding-docs-joined-supported 1328 297 314 2734 4
grounding-docs-joined-unsupported 707 300 316 2210 4
grounding-docs-number-changed 452 190 197 2144 4
grounding-docs-partly 1470 184 203 2160 4
grounding-docs-supported 1959 172 179 2164 4
grounding-docs-unanswerable 432 200 208 2111 4
grounding-fever-combined 1260 81 85 1523 4
grounding-fever-contradicted 626 66 65 1319 4
grounding-fever-joined-contradicted 382 93 97 1141 4
grounding-fever-joined-supported 772 91 94 1158 4
grounding-fever-joined-unsupported 429 93 97 1093 4
grounding-fever-supported 992 66 66 1388 4
grounding-fever-unsupported 579 67 66 1333 4
grounding-hotpot-contradicted 379 160 164 1460 4
grounding-hotpot-gold-removed 541 161 162 1473 4
grounding-hotpot-number-changed 307 160 165 1470 4
grounding-hotpot-one-removed 434 137 145 849 4
grounding-hotpot-partly 839 184 190 1484 4
grounding-hotpot-supported 1074 157 160 1470 4
grounding-squad-combined 900 307 302 1411 4
grounding-squad-contradicted 672 183 183 1514 4
grounding-squad-gold-removed 767 180 179 1424 4
grounding-squad-joined-contradicted 424 327 331 1850 4
grounding-squad-joined-supported 1097 315 317 1880 4
grounding-squad-joined-unsupported 605 312 316 1378 4
grounding-squad-number-changed 323 178 180 1544 4
grounding-squad-partly 1442 202 203 1441 4
grounding-squad-supported 1661 179 179 1488 4
grounding-squad-unanswerable 51 129 126 1408 4

Correct option: longest / shortest / position / key

For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):

split rows correct is longest correct is shortest chance (1/listed) mean relative position (0 first, 1 last; 0.5 expected) position fifths

Correct key and position by option count

split options rows mean options top correct keys most common position (0-based)
train 2-5 25380 4.0 supported 35.0%, partly_supported 25.0%, contradicted 20.0%, unsupported 20.0% 2 (25.1%)
test 2-5 2080 4.0 supported 35.0%, partly_supported 25.0%, unsupported 20.0%, contradicted 20.0% 1 (25.5%)

Option count by label class (train)

class rows min median mean max
contradicted 5076 4 4 4.0 4
partly_supported 6345 4 4 4.0 4
supported 8883 4 4 4.0 4
unsupported 5076 4 4 4.0 4

Source fields by label class (train)

Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.

Overall majority: 35.0%.

source field values purity top values → classes
kind 32 100.0% grounding-docs-supported: supported 100.0%; grounding-squad-supported: supported 100.0%; grounding-docs-partly: partly_supported 100.0%; grounding-squad-partly: partly_supported 100.0%; grounding-docs-joined-supported: supported 100.0%; grounding-fever-combined: partly_supported 100.0%
second_opinion 2 48.3% null: supported 45.1%, partly_supported 32.2%; deepseek-v4-flash: contradicted 59.2%, unsupported 40.8%
dataset 4 35.8% teacher-docs: supported 37.3%, unsupported 23.8%; squad2: supported 34.7%, partly_supported 29.5%; fever: supported 35.0%, partly_supported 25.0%; hotpotqa: partly_supported 35.6%, supported 30.1%
url 4 35.8% : supported 37.3%, unsupported 23.8%; https://huggingface.co/datasets/rajpurkar/squad_v2: supported 34.7%, partly_supported 29.5%; https://fever.ai/dataset/fever.html: supported 35.0%, partly_supported 25.0%; https://huggingface.co/datasets/hotpotqa/hotpot_qa: partly_supported 35.6%, supported 30.1%
revision 4 35.8% : supported 37.3%, unsupported 23.8%; 3ffb306f725f7d2ce8394bc1873b24868140c412: supported 34.7%, partly_supported 29.5%; sha256 train.jsonl eba7e8f8..., shared_task_dev.jsonl and wiki-pages.zip 4b06d95d... (see raw/manifest.json): supported 35.0%, partly_supported 25.0%; 1908d6afbbead072334abe2965f91bd2709910ab: partly_supported 35.6%, supported 30.1%
check_teacher 2 35.2% hosted: qwen3.8-flash: supported 35.2%, partly_supported 24.7%; local: qwen3.8-flash-next: partly_supported 38.2%, supported 26.4%
text_teacher 4 35.1% hosted: qwen3.8-max: supported 34.3%, partly_supported 26.4%; local: qwen3.8-flash-next: supported 36.9%, unsupported 24.8%; none (public data and code): supported 35.0%, partly_supported 25.0%; mixed: hosted: qwen3.8-max + local: qwen3.8-flash-next: partly_supported 37.5%, supported 34.4%
license 3 35.0% CC BY-SA 4.0: supported 33.3%, partly_supported 31.4%; generated by the local teacher (qwen3.8-flash-next): supported 37.3%, unsupported 23.8%; CC BY-SA 3.0: supported 35.0%, partly_supported 25.0%
source_style 3 35.0% bracket: supported 34.8%, partly_supported 25.2%; document: supported 35.5%, partly_supported 24.8%; source: supported 34.8%, partly_supported 25.0%

Formatting by label class (main text, share of rows)

feature contradicted partly_supported supported unsupported
ends with ? 0.0% 0.0% 0.0% 0.0%
ends with . 99.6% 99.5% 99.6% 99.6%
ends with ! 0.0% 0.0% 0.0% 0.0%
no end punctuation 0.1% 0.0% 0.0% 0.0%
starts lowercase 0.0% 0.0% 0.0% 0.0%
all lowercase 0.0% 0.0% 0.0% 0.0%
has a digit 59.9% 57.5% 47.1% 46.5%
has newline 0.0% 0.0% 0.0% 0.0%
has quotes 4.6% 6.8% 4.7% 4.4%
has markup (HTML/markdown) 0.0% 0.0% 0.0% 0.0%
has URL 0.0% 0.0% 0.0% 0.0%
non-ASCII 4.3% 5.5% 4.5% 4.3%
non-Latin script 0.0% 0.0% 0.0% 0.0%
emoji 0.0% 0.0% 0.0% 0.0%
ALL-CAPS word (4+) 2.9% 2.7% 2.8% 2.9%
contains ' - ' or — 0.1% 0.1% 0.1% 0.1%

Over-represented words and phrases per label class (main text)

Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.

Words, contradicted: yes z=7 6.0% vs 3.7%; did z=6 1.8% vs 0.8%; if z=5 12.5% vs 9.6%; incapable z=5 0.9% vs 0.0%; not z=5 9.5% vs 7.1%; only z=5 6.1% vs 4.2%; ever z=5 0.7% vs 0.3%; unable z=5 0.4% vs 0.1%; 60 z=4 1.4% vs 0.7%; 1 z=4 5.0% vs 3.7%; exactly z=4 2.7% vs 1.8%; yet z=4 0.3% vs 0.1%; refused z=4 0.3% vs 0.1%; you z=4 22.7% vs 19.6%; zero z=4 0.9% vs 0.5%; 8 z=4 2.0% vs 1.3%; avoided z=4 0.2% vs 0.0%; any z=4 5.6% vs 4.3%; 17 z=4 1.0% vs 0.6%; except z=4 0.4% vs 0.1%

Words, partly_supported: which z=15 15.8% vs 9.1%; in z=14 57.0% vs 46.0%; was z=14 32.9% vs 24.3%; born z=11 4.5% vs 1.9%; over z=10 5.4% vs 2.8%; he z=9 9.3% vs 6.1%; who z=9 8.5% vs 5.5%; provided z=9 3.6% vs 1.7%; his z=9 6.8% vs 4.2%; university z=8 2.6% vs 1.2%; particularly z=8 0.9% vs 0.1%; million z=8 2.8% vs 1.3%; roughly z=8 0.9% vs 0.1%; founded z=8 2.2% vs 1.0%; primarily z=8 1.2% vs 0.3%; having z=7 1.1% vs 0.3%; filmed z=7 0.7% vs 0.1%; originally z=7 1.7% vs 0.7%; with z=7 15.0% vs 12.1%; near z=7 1.5% vs 0.6%

Words, supported: no z=8 8.9% vs 6.1%; if z=5 11.7% vs 9.4%; or z=4 10.6% vs 8.8%; cannot z=4 1.9% vs 1.2%; must z=4 13.4% vs 11.5%; date z=4 2.8% vs 2.0%; twelve z=3 1.9% vs 1.4%; are z=3 17.2% vs 15.3%; within z=3 9.4% vs 8.1%; you z=3 21.6% vs 19.5%; fourteen z=3 1.4% vs 0.9%; prohibited z=3 1.2% vs 0.8%; four z=3 3.6% vs 2.9%; strictly z=3 1.5% vs 1.1%; maximum z=3 2.2% vs 1.7%; forty z=3 1.7% vs 1.3%; be z=3 11.5% vs 10.2%; winner z=3 0.2% vs 0.1%; ltd z=3 0.5% vs 0.2%; product z=3 0.9% vs 0.6%

Words, unsupported: yes z=10 6.9% vs 3.5%; automatically z=6 2.8% vs 1.5%; your z=6 12.4% vs 9.3%; parking z=5 0.5% vs 0.1%; app z=5 0.9% vs 0.3%; discount z=5 0.7% vs 0.2%; available z=5 1.7% vs 0.9%; supports z=5 0.6% vs 0.2%; no z=5 8.7% vs 6.7%; mobile z=4 0.6% vs 0.2%; ios z=4 0.3% vs 0.0%; all z=4 6.0% vs 4.5%; android z=4 0.4% vs 0.1%; includes z=4 1.8% vs 1.1%; can z=4 7.4% vs 5.8%; strictly z=4 1.8% vs 1.1%; without z=4 3.0% vs 2.1%; you z=4 22.7% vs 19.6%; fury z=4 0.2% vs 0.0%; device z=4 1.4% vs 0.8%

2–4-word phrases, contradicted: not a z=6 0.6% vs 0.1%; did not z=6 1.5% vs 0.5%; is not a z=6 0.5% vs 0.0%; was not z=5 0.8% vs 0.2%; incapable of z=5 0.9% vs 0.0%; of being z=5 0.4% vs 0.1%; was only z=5 0.3% vs 0.0%; has only z=5 0.3% vs 0.0%; only ever z=4 0.3% vs 0.0%; long as z=4 0.6% vs 0.2%; as long as z=4 0.6% vs 0.2%; as long z=4 0.6% vs 0.2%; incapable of being z=4 0.4% vs 0.0%; unable to z=4 0.4% vs 0.1%; was incapable z=4 0.4% vs 0.0%; was incapable of z=4 0.4% vs 0.0%; 60 days z=4 0.4% vs 0.1%; not an z=4 0.3% vs 0.0%; long as they z=4 0.2% vs 0.0%; as long as they z=4 0.2% vs 0.0%

2–4-word phrases, partly_supported: and you z=15 2.7% vs 0.5%; was born z=12 3.7% vs 1.5%; which was z=12 2.9% vs 1.1%; is a z=11 8.4% vs 5.8%; and you must z=11 1.2% vs 0.1%; born in z=10 2.1% vs 0.8%; in the z=9 16.7% vs 14.3%; provided the z=9 1.0% vs 0.2%; was born in z=9 1.7% vs 0.6%; and it z=9 1.3% vs 0.4%; provided you z=8 1.2% vs 0.4%; in his z=8 1.1% vs 0.3%; was a z=8 3.1% vs 1.8%; is an z=8 3.0% vs 1.7%; and you will z=8 0.7% vs 0.1%; was born on z=8 1.5% vs 0.6%; he was z=8 2.8% vs 1.6%; founded in z=8 1.1% vs 0.3%; is the z=7 6.0% vs 4.5%; born on z=7 1.7% vs 0.8%

2–4-word phrases, supported: is a film z=4 0.5% vs 0.2%; if you z=4 5.5% vs 4.3%; you must z=4 8.0% vs 6.7%; no you z=4 0.9% vs 0.5%; a film z=4 0.8% vs 0.4%; academy awards z=3 0.2% vs 0.1%; no the z=3 0.9% vs 0.6%; are strictly z=3 0.8% vs 0.5%; are not z=3 1.2% vs 0.8%; maximum of z=3 1.1% vs 0.7%; a maximum of z=3 1.1% vs 0.7%; there is z=3 1.3% vs 0.9%; strictly prohibited z=3 0.8% vs 0.5%; are strictly prohibited z=3 0.6% vs 0.3%; proceed to z=3 0.2% vs 0.0%; you cannot z=3 0.7% vs 0.4%; at least one z=3 0.4% vs 0.2%; least one z=3 0.4% vs 0.2%; a maximum z=3 1.5% vs 1.1%; if the z=3 2.6% vs 2.0%

2–4-word phrases, unsupported: yes the z=7 1.6% vs 0.6%; discount on z=5 0.3% vs 0.0%; a dedicated z=5 0.5% vs 0.1%; mobile app z=4 0.3% vs 0.1%; percent discount z=4 0.3% vs 0.1%; available on z=4 0.3% vs 0.0%; stars an z=4 0.2% vs 0.0%; percent discount on z=4 0.2% vs 0.0%; yes there z=4 0.2% vs 0.0%; yes you z=4 1.3% vs 0.7%; gift of the night z=4 0.2% vs 0.0%; the night fury z=4 0.2% vs 0.0%; night fury z=4 0.2% vs 0.0%; gift of the z=4 0.2% vs 0.0%; gift of z=4 0.2% vs 0.0%; of the night fury z=4 0.2% vs 0.0%; yes you are z=4 0.2% vs 0.0%; the night fury stars z=4 0.2% vs 0.0%; fury stars z=4 0.2% vs 0.0%; night fury stars z=4 0.2% vs 0.0%

By row kind, 1–3-word phrases (top 10)

  • grounding-docs-contradicted: you 48.0% vs 19.2%; if 28.4% vs 9.5%; yes 15.3% vs 3.8%; your 24.0% vs 9.4%; must 26.3% vs 11.6%; if you 13.1% vs 4.4%; days 17.2% vs 6.8%; within 19.1% vs 8.2%; yes you 3.6% vs 0.7%; pounds 7.9% vs 2.6%
  • grounding-docs-gold-removed: you 46.7% vs 19.2%; if 28.4% vs 9.5%; your 24.7% vs 9.3%; must 27.9% vs 11.5%; will 15.5% vs 5.7%; no 17.1% vs 6.7%; if you 12.6% vs 4.4%; you must 17.0% vs 6.8%; days 16.1% vs 6.9%; hours 12.0% vs 4.7%
  • grounding-docs-joined-contradicted: must 45.3% vs 11.4%; your 38.2% vs 9.3%; yes 19.7% vs 3.8%; you 64.4% vs 19.2%; if 37.5% vs 9.6%; you must 27.8% vs 6.7%; will 23.8% vs 5.7%; days 26.6% vs 6.8%; percent 19.2% vs 4.4%; any 18.3% vs 4.2%
  • grounding-docs-joined-supported: must 44.8% vs 10.3%; you 62.7% vs 17.9%; you must 28.8% vs 6.0%; no 28.2% vs 5.9%; days 28.1% vs 6.1%; if 35.7% vs 8.8%; your 34.4% vs 8.6%; within 30.5% vs 7.4%; hours 19.6% vs 4.2%; percent 17.4% vs 4.1%
  • grounding-docs-joined-unsupported: days 30.1% vs 6.6%; must 43.7% vs 11.2%; you 62.4% vs 19.0%; you must 28.3% vs 6.6%; if 36.6% vs 9.4%; your 35.6% vs 9.2%; no 27.7% vs 6.5%; within 30.0% vs 7.9%; five 19.1% vs 4.4%; percent 18.7% vs 4.4%
  • grounding-docs-number-changed: you 49.6% vs 19.7%; if 29.6% vs 9.8%; must 33.0% vs 11.8%; you must 21.2% vs 6.9%; 36 2.2% vs 0.1%; 00 9.5% vs 2.1%; if you 15.0% vs 4.5%; 1 13.3% vs 3.8%; 75 3.3% vs 0.3%; seconds 6.4% vs 1.3%
  • grounding-docs-partly: and you 11.8% vs 0.4%; you 53.4% vs 18.2%; provided 13.9% vs 1.5%; must 36.2% vs 10.7%; you must 22.1% vs 6.3%; within 23.7% vs 7.6%; provided you 5.4% vs 0.3%; and you must 5.4% vs 0.1%; your 25.3% vs 9.0%; days 19.8% vs 6.4%
  • grounding-docs-supported: you 47.8% vs 17.9%; if 26.7% vs 8.8%; must 28.3% vs 10.8%; your 23.7% vs 8.8%; no 18.0% vs 6.2%; if you 12.9% vs 4.0%; you must 16.8% vs 6.4%; will 14.5% vs 5.4%; within 18.3% vs 7.7%; days 15.8% vs 6.5%
  • grounding-docs-unanswerable: yes 44.9% vs 3.5%; yes the 14.1% vs 0.5%; yes you 9.3% vs 0.7%; can 23.8% vs 5.8%; your 30.6% vs 9.6%; app 6.2% vs 0.3%; yes you can 6.2% vs 0.3%; available 8.8% vs 0.9%; parking 5.1% vs 0.1%; you can 14.4% vs 2.9%
  • grounding-fever-combined: is a 22.9% vs 5.6%; a 55.0% vs 46.9%; is 46.3% vs 39.7%; has 17.4% vs 6.3%; was 32.9% vs 26.1%; an 18.6% vs 11.6%; in 40.2% vs 49.2%; is an 7.2% vs 1.8%; was in 4.8% vs 0.6%; was a 6.8% vs 1.9%
  • grounding-fever-contradicted: did not 6.1% vs 0.6%; incapable 5.8% vs 0.1%; incapable of 5.8% vs 0.1%; did 6.2% vs 0.9%; not 14.9% vs 7.4%; film 11.0% vs 4.3%; only 11.0% vs 4.4%; was 27.5% vs 26.4%; the 54.8% vs 81.4%; of being 2.7% vs 0.1%
  • grounding-fever-joined-contradicted: is a 21.5% vs 6.2%; not 20.4% vs 7.4%; is not 9.4% vs 1.2%; is 50.3% vs 39.9%; was 36.1% vs 26.3%; a 50.0% vs 47.2%; not a 4.2% vs 0.1%; is not a 3.4% vs 0.1%; incapable 3.1% vs 0.1%; incapable of 3.1% vs 0.1%
  • grounding-fever-joined-supported: is a 25.9% vs 5.8%; is 50.1% vs 39.8%; a 53.4% vs 47.1%; was 35.9% vs 26.2%; film 13.5% vs 4.2%; a film 5.1% vs 0.4%; is an 7.9% vs 1.8%; has 13.2% vs 6.6%; in 41.3% vs 49.0%; movie 5.1% vs 0.9%
  • grounding-fever-joined-unsupported: has 19.8% vs 6.6%; a 53.8% vs 47.2%; was in 7.0% vs 0.7%; in a 11.7% vs 3.0%; in 49.2% vs 48.8%; award 7.0% vs 1.0%; was 33.3% vs 26.3%; starred in 4.9% vs 0.7%; is a 12.4% vs 6.3%; was in a 3.3% vs 0.2%
  • grounding-fever-supported: the 58.0% vs 81.6%; film 10.7% vs 4.2%; was 26.1% vs 26.5%; of 39.7% vs 54.1%; a 36.1% vs 47.7%; is a 10.6% vs 6.3%; in 35.3% vs 49.3%; is 30.6% vs 40.5%; a film 2.8% vs 0.5%; has 9.0% vs 6.7%
  • grounding-fever-unsupported: has 14.2% vs 6.6%; in a 9.3% vs 3.0%; in 39.7% vs 49.0%; was 25.9% vs 26.5%; a 34.5% vs 47.6%; the 49.4% vs 81.4%; of 35.2% vs 54.0%; nominated for 1.9% vs 0.2%; film 7.3% vs 4.4%; series 4.0% vs 1.5%
  • grounding-hotpot-contradicted: who 23.2% vs 6.0%; film 16.6% vs 4.3%; was 46.4% vs 26.2%; american 13.5% vs 3.7%; he 17.7% vs 6.8%; is 57.3% vs 39.8%; looking for 3.7% vs 0.3%; is the 13.7% vs 4.8%; directed 7.1% vs 1.5%; known 10.3% vs 3.0%
  • grounding-hotpot-gold-removed: who 25.5% vs 5.8%; american 15.0% vs 3.6%; film 15.7% vs 4.2%; was 46.2% vs 26.0%; is 59.1% vs 39.7%; which is 10.5% vs 2.6%; which 22.6% vs 10.5%; an american 7.0% vs 1.2%; is the 13.9% vs 4.7%; who was 6.1% vs 1.0%
  • grounding-hotpot-number-changed: film 23.1% vs 4.2%; was 60.9% vs 26.0%; born on 9.4% vs 0.9%; you are asking 6.5% vs 0.4%; are asking 6.5% vs 0.4%; are asking about 6.5% vs 0.4%; asking 6.8% vs 0.5%; which was 10.7% vs 1.5%; asking about 6.5% vs 0.4%; directed 10.4% vs 1.4%
  • grounding-hotpot-one-removed: born 15.9% vs 2.3%; was born 13.8% vs 1.8%; was 52.1% vs 26.0%; who 22.1% vs 5.9%; born on 8.5% vs 0.9%; was born on 7.8% vs 0.7%; american 14.3% vs 3.7%; he was born 5.5% vs 0.4%; he 18.9% vs 6.7%; film 14.5% vs 4.3%
  • grounding-hotpot-partly: which 34.4% vs 10.0%; who 25.1% vs 5.6%; was 55.7% vs 25.5%; born 14.2% vs 2.2%; which was 10.3% vs 1.3%; in 78.9% vs 47.8%; was born 11.3% vs 1.7%; film 15.3% vs 4.1%; born in 7.3% vs 0.9%; he 19.1% vs 6.5%
  • grounding-hotpot-supported: who 22.6% vs 5.5%; was 48.2% vs 25.5%; film 15.1% vs 4.0%; american 13.7% vs 3.4%; is 57.3% vs 39.3%; which 23.0% vs 10.2%; he 16.8% vs 6.5%; in 60.0% vs 48.3%; which is 9.0% vs 2.5%; is the 12.4% vs 4.6%
  • grounding-squad-combined: were 16.2% vs 4.7%; government 6.6% vs 1.5%; that 32.2% vs 12.9%; other 9.1% vs 2.5%; they 18.4% vs 6.5%; these 11.6% vs 3.5%; used 9.1% vs 2.6%; this 37.7% vs 15.9%; it was 9.0% vs 2.7%; this happened 2.0% vs 0.2%
  • grounding-squad-contradicted: were 10.9% vs 5.0%; this 26.8% vs 16.4%; its 9.2% vs 4.1%; that 21.7% vs 13.3%; these 8.2% vs 3.7%; according to 3.1% vs 0.9%; according 3.1% vs 0.9%; as 24.1% vs 15.5%; their 9.8% vs 5.0%; act 2.2% vs 0.5%
  • grounding-squad-gold-removed: were 11.9% vs 4.9%; this 25.9% vs 16.4%; these 7.7% vs 3.7%; and 53.2% vs 41.2%; they 11.9% vs 6.8%; the city 3.1% vs 1.1%; difficult 0.9% vs 0.1%; had been 1.6% vs 0.3%; as 22.4% vs 15.5%; that 19.7% vs 13.4%
  • grounding-squad-joined-contradicted: this 47.2% vs 16.1%; were 18.4% vs 4.9%; that 37.3% vs 13.2%; it was 11.1% vs 2.7%; this was 4.0% vs 0.6%; would 5.0% vs 0.8%; their 16.3% vs 5.0%; they 20.5% vs 6.7%; it 38.4% vs 14.7%; they were 3.8% vs 0.6%
  • grounding-squad-joined-supported: were 18.0% vs 4.6%; they 21.3% vs 6.3%; this 39.6% vs 15.6%; these 12.8% vs 3.4%; that 33.1% vs 12.7%; it 35.4% vs 14.2%; as 35.3% vs 14.9%; their 14.3% vs 4.7%; many 6.3% vs 1.5%; he 17.8% vs 6.4%
  • grounding-squad-joined-unsupported: these 14.2% vs 3.6%; were 17.5% vs 4.8%; this 42.8% vs 16.0%; they 21.0% vs 6.6%; most 10.1% vs 2.7%; often 4.5% vs 0.8%; as 36.7% vs 15.2%; that the 4.5% vs 0.9%; because 9.6% vs 2.8%; that 31.7% vs 13.1%
  • grounding-squad-number-changed: had 11.1% vs 3.7%; in 1937 1.2% vs 0.0%; were 12.7% vs 5.0%; dates back 1.2% vs 0.1%; 1880 1.2% vs 0.1%; as of 3.1% vs 0.5%; in 1976 0.9% vs 0.0%; in 1880 0.9% vs 0.0%; 1926 0.9% vs 0.0%; hyderabad 1.2% vs 0.1%
  • grounding-squad-partly: in 75.8% vs 47.2%; over 10.5% vs 3.1%; particularly 3.2% vs 0.2%; in his 3.1% vs 0.3%; roughly 2.1% vs 0.2%; that 22.4% vs 13.0%; his 10.3% vs 4.5%; around 3.5% vs 0.8%; especially 1.9% vs 0.2%; were 10.5% vs 4.8%
  • grounding-squad-supported: this 26.3% vs 16.0%; were 10.6% vs 4.8%; that 21.4% vs 13.0%; they 12.3% vs 6.6%; and 52.0% vs 40.9%; as 22.8% vs 15.3%; these 7.6% vs 3.6%; it 21.7% vs 14.6%; their 9.0% vs 4.9%; to 49.1% vs 40.3%
  • grounding-squad-unanswerable: allowing 7.8% vs 0.1%; it is called 3.9% vs 0.0%; lacking 3.9% vs 0.0%; philosophical 3.9% vs 0.0%; have no 3.9% vs 0.0%; possess 3.9% vs 0.0%; forming 3.9% vs 0.1%; history of 3.9% vs 0.1%; knowledge 3.9% vs 0.1%; stopped 3.9% vs 0.1%

Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):

  • grounding-docs-unanswerable: yes 44.9% vs 3.5%
  • grounding-docs-joined-supported: must 44.8% vs 10.3%
  • grounding-docs-joined-contradicted: your 38.2% vs 9.3%
  • grounding-docs-joined-supported: if 35.7% vs 8.8%
  • grounding-docs-joined-supported: your 34.4% vs 8.6%
  • grounding-docs-joined-supported: within 30.5% vs 7.4%
  • grounding-docs-joined-unsupported: days 30.1% vs 6.6%
  • grounding-docs-joined-supported: you must 28.8% vs 6.0%
  • grounding-docs-joined-unsupported: you must 28.3% vs 6.6%
  • grounding-docs-joined-supported: no 28.2% vs 5.9%
  • grounding-docs-joined-supported: days 28.1% vs 6.1%
  • grounding-docs-joined-contradicted: you must 27.8% vs 6.7%
  • grounding-docs-joined-unsupported: no 27.7% vs 6.5%
  • grounding-fever-joined-supported: is a 25.9% vs 5.8%
  • grounding-hotpot-gold-removed: who 25.5% vs 5.8%
  • grounding-hotpot-partly: who 25.1% vs 5.6%
  • grounding-docs-unanswerable: can 23.8% vs 5.8%
  • grounding-docs-joined-contradicted: will 23.8% vs 5.7%
  • grounding-hotpot-number-changed: film 23.1% vs 4.2%
  • grounding-fever-combined: is a 22.9% vs 5.6%
  • grounding-hotpot-supported: who 22.6% vs 5.5%
  • grounding-docs-joined-contradicted: yes 19.7% vs 3.8%
  • grounding-docs-joined-supported: hours 19.6% vs 4.2%
  • grounding-docs-joined-contradicted: percent 19.2% vs 4.4%
  • grounding-docs-joined-contradicted: hours 19.2% vs 4.6%
  • grounding-docs-joined-unsupported: five 19.1% vs 4.4%
  • grounding-docs-joined-unsupported: hours 18.8% vs 4.6%
  • grounding-docs-joined-unsupported: percent 18.7% vs 4.4%
  • grounding-docs-joined-contradicted: any 18.3% vs 4.2%
  • grounding-docs-joined-contradicted: if you 18.3% vs 4.4%
  • grounding-docs-joined-unsupported: if you 17.7% vs 4.3%
  • grounding-docs-joined-supported: percent 17.4% vs 4.1%
  • grounding-docs-joined-supported: if you 17.2% vs 4.0%
  • grounding-hotpot-one-removed: born 15.9% vs 2.3%
  • grounding-docs-joined-supported: per 15.4% vs 3.7%
  • grounding-docs-contradicted: yes 15.3% vs 3.8%
  • grounding-hotpot-gold-removed: american 15.0% vs 3.6%
  • grounding-docs-unanswerable: you can 14.4% vs 2.9%
  • grounding-hotpot-partly: born 14.2% vs 2.2%
  • grounding-docs-unanswerable: yes the 14.1% vs 0.5%

Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):

  • all rows (25,380 rows): none

Shortcut models

Predicting the label class on test (2,080 rows). Chance 25.0%, majority class ('supported') 35.0%; balanced chance 25.0%.

model (logistic regression, trained on the train sample) test accuracy balanced accuracy (mean recall)
bag of words, whole state (words and word pairs) 34.8% 28.4%
bag of words, main text only (answer) 39.0% 33.7%
surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) 37.1% 28.4%
surface features of the main text only 35.7% 26.3%

Strongest single surface features (logistic regression on one feature, balanced accuracy on test):

feature accuracy balanced accuracy
other_state_chars(log) 36.2% 27.2%
state_chars(log) 35.3% 26.3%
digit_ratio 35.0% 25.3%
nonascii_ratio 35.0% 25.1%
count_" 35.0% 25.1%
chars(log) 35.0% 25.0%
words(log) 35.0% 25.0%
nonlatin 35.0% 25.0%
starts_upper 35.0% 25.0%
all_lower 35.0% 25.0%
  • Predicting the row kind instead (32 kinds, balanced chance 3.1%): surface features balanced accuracy 20.3% (accuracy 27.6%); bag of words of the state 14.8% (accuracy 23.5%).

Other state fields alone (predicting the label class on test from one field, without the main text):

field treated as accuracy balanced accuracy
sources text: bag of words / length+empty 33.5% / 36.2% 27.3% / 27.3%

No-meaning option picker: a logistic ranker scores each option from its position, length, key type, fixed-option identity and shape (commas, brackets, capitals), never reading the state or the option's words, and picks the top option per row.

  • Test accuracy 35.0% against uniform chance 25.0% (this includes the fixed options, whose share is a class prior).

2. Duplicates and split separation

Families shared between splits

splits shared families examples
train ∩ dev 0
train ∩ calibration 0
train ∩ test 0
dev ∩ calibration 0
dev ∩ test 0
calibration ∩ test 0
  • Train rows whose main text repeats an earlier row's (normalised): 2,291 (9.0%); groups: 2,176; largest group 4.
  • Train rows identical in the whole prompt (state, options, instructions): 0.
  • Identical whole prompt, different answer: 0 groups (0 rows).
  • Identical main text, different label class: 2170 groups (4455 rows). (Can be legitimate when the rest of the state or the options differ.)
    • "Cosmopolitan as of 2011 contains content which includes articles on home decor." ×4: unsupported 3, contradicted 1
    • "Navin Kumar served as the first chairman of the Goods and Services Tax Network (GSTN). This indirect tax was introduced…" ×3: supported 1, unsupported 1, partly_supported 1
    • "The bassist of Rusted Root, the band that released their fifth studio album "Welcome to My Party", is Patrick Norman. T…" ×3: unsupported 1, partly_supported 1, supported 1
    • "John Leguizamo is the Colombian-American actor who starred in the 2000 drama film "King of the Jungle." He is also well…" ×3: partly_supported 1, supported 1, unsupported 1
    • "Santana Row is located across Stevens Creek Boulevard from Westfield Valley Fair, which is commonly known as Valley Fai…" ×3: unsupported 1, partly_supported 1, supported 1
    • "Governor-elect Alejandro García Padilla fought against statehood by asking President Barack Obama to reject the referen…" ×3: unsupported 1, supported 1, partly_supported 1
    • "The LA Galaxy, whose president is Chris Klein and which is based in Carson, California, began to play in 1996." ×3: supported 1, partly_supported 1, unsupported 1
    • "Alex Cox directed Sid and Nancy, the 1986 British biopic starring Gary Oldman and Chloe Webb that was parodied in The S…" ×3: supported 1, unsupported 1, partly_supported 1

Most repeated main texts in train:

  • ×4: "cosmopolitan as of 2011 contains content which includes articles on home decor." (unsupported 3, contradicted 1)
  • ×3: "zoé is the older band. they initially formed in mexico city in 1994, while flyleaf was formed in texas in 2002." (supported 1, partly_supported 1, unsupported 1)
  • ×3: "zazie beetz has been cast as neena thurman in "deadpool 2", which is directed by david leitch. the film is an upcoming american superhero m…" (unsupported 1, partly_supported 1, supported 1)
  • ×3: "vidushi shashikala dani is the only all india radio graded female exponent of the jaltarang, which is a percussion instrument." (supported 1, partly_supported 1, unsupported 1)
  • ×3: "vch was reportedly the most popular programming on qube, a cable television system that was launched on december 1, 1977." (partly_supported 1, supported 1, unsupported 1)
  • ×3: "tom rolt was a prolific english writer and the biographer of major civil engineering figures including isambard kingdom brunel and thomas t…" (supported 1, partly_supported 1, unsupported 1)
  • ×3: "tlc, the american girl group that released "girl talk" in 2002, received the million certification from the recording industry association …" (unsupported 1, supported 1, partly_supported 1)
  • ×3: "thomas bartley officiated in home tests against pakistan. the pakistan national cricket team is popularly referred to as the shaheens, men …" (supported 1, partly_supported 1, unsupported 1)

Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):

split rows examples
dev 0 (0.0%)
calibration 0 (0.0%)
test 0 (0.0%)

Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)

  • Train: 5,705 near-duplicate pairs; 7,898 rows (31.1%) sit in 3,333 clusters; largest cluster 6; excess rows (cluster size − 1) 4,565 (18.0%).
  • Clusters with more than one label class: 3,288 (7,808 rows).
    • ×6: "Garrison Hearst, who won the NFL Comeback Player of the Year Award in 2001, was featured on the cover of Madd…" → partly_supported 2, contradicted 2, supported 1, unsupported 1
    • ×6: "The screenwriter you are asking about is Marc Silverstein, who co-wrote the film Valentine's Day directed by …" → contradicted 2, partly_supported 2, supported 1, unsupported 1
    • ×6: "The LIGO Scientific Collaboration (LSC) was established in 1997 under the leadership of Barry Barish, an Amer…" → partly_supported 2, contradicted 2, supported 1, unsupported 1
    • ×6: ""Këmisha e zezë" was the organ of the Albanian Fascist Party. This party held nominal power over Albania from…" → contradicted 2, partly_supported 2, supported 1, unsupported 1
    • ×5: "Samuel Fraunces provided for prisoners during the Civil War. He was the owner and operator of Fraunces Tavern…" → contradicted 2, unsupported 1, partly_supported 1, supported 1
  • Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 0 (0.0%), test 0 (0.0%)

Largest train clusters:

  • ×6: "Garrison Hearst, who won the NFL Comeback Player of the Year Award in 2001, was featured on the cover of Madden NFL 99. Specifically, the E…"
  • ×6: "The screenwriter you are asking about is Marc Silverstein, who co-wrote the film Valentine's Day directed by Garry Marshall. He was born on…"
  • ×6: "The LIGO Scientific Collaboration (LSC) was established in 1997 under the leadership of Barry Barish, an American experimental physicist. I…"
  • ×6: ""Këmisha e zezë" was the organ of the Albanian Fascist Party. This party held nominal power over Albania from 1939, when the country was co…"
  • ×5: "Samuel Fraunces provided for prisoners during the Civil War. He was the owner and operator of Fraunces Tavern, which is situated at 54 Pear…"

5. Junk

split empty main text main text under 10 characters
train 0 0
dev 0 0
calibration 0 0
test 0 0

Pattern scan of train main texts (count, then the share of each class's rows):

pattern rows by class
placeholder [NAME]-style 0 (0.0%)
lorem ipsum 0 (0.0%)
TODO/TBD/FIXME 0 (0.0%)
'As an AI' / refusal 0 (0.0%)
chat preamble ('Here is/are...', 'Sure!') 0 (0.0%)
meta words (example/variation/message:) 0 (0.0%)
model thinking tags 0 (0.0%)
JSON/code-fence leftovers 0 (0.0%)
encoding garbage (mojibake/replacement char) 0 (0.0%)
HTML tag 0 (0.0%)
HTML entity 0 (0.0%)
base64-like run (40+ chars) 0 (0.0%)
URL 0 (0.0%)
  • Possibly cut off: 0 of 3,620 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: contradicted 0.0%, partly_supported 0.0%, supported 0.0%, unsupported 0.0%

Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):

  • none

6. Samples

20 random train rows per kind: ground-grounding-samples.txt. Reading notes are in the findings above.

QA: ground-rerank

Checked 2026-09-30 13:20 by adapters/qa/qa.py (READY file READY-ground, v4 (held-out second opinion; grounding answers balanced on surface features, 2026-09-30)).

Verdict: PASS WITH NOTES (see ground-grounding.md for the full notes)

Automatic flags (for the reviewer to judge; not all are problems)

  • kind 'rerank-docs' share varies across splits by more than 5 points: train 11.5%, dev 6.1%, calibration 8.2%, test 9.7%
  • kind 'rerank-squad' share varies across splits by more than 5 points: train 59.1%, dev 66.0%, calibration 60.7%, test 59.1%

Data checked

split rows families file
train 25,380 5141 train.jsonl
dev 900 205 dev.jsonl
calibration 720 190 calibration.jsonl
test 2,080 565 test.jsonl
  • Train sha256: c8b38f85a728f5aa916c2b4b2cf18a8cd63567fde55572cc4b0f16b155915fb2 (READY says c8b38f85a728f5aa916c2b4b2cf18a8cd63567fde55572cc4b0f16b155915fb2: match)
  • Main text field (the text the phrase and length checks use): state.question.
  • Label classes: <listed option>, none, p1, p10, p11, p12, p13, p14, p15, p16, p17, p18, p19, p2, p20, p21, p22, p3, p4, p5, p6, p7, p8, p9 (<listed option> = one of the per-row listed options such as t3 or o12). Row kinds (source.kind): rerank-docs, rerank-docs-unanswerable, rerank-hotpot, rerank-hotpot-removed, rerank-squad, rerank-squad-nearmiss, rerank-squad-nearmiss-none, rerank-squad-removed, rerank-squad-unanswerable.

3. Balance

Label class share per split

lclass train dev calibration test train rows
<listed option> 4.5% 5.9% 5.6% 5.9% 1,136
none 15.0% 15.0% 15.0% 15.0% 3,807
p1 7.9% 8.9% 8.6% 8.1% 1,994
p10 3.2% 3.4% 2.5% 3.7% 804
p11 2.6% 3.1% 2.6% 2.7% 669
p12 2.4% 2.9% 2.1% 2.4% 603
p13 2.0% 1.8% 1.9% 1.7% 495
p14 1.9% 1.7% 2.2% 2.1% 492
p15 1.7% 1.2% 1.7% 1.2% 421
p16 1.5% 1.8% 1.8% 1.6% 377
p17 1.2% 1.0% 1.2% 1.6% 314
p18 1.3% 1.0% 1.7% 0.9% 320
p19 1.0% 0.7% 0.7% 1.2% 245
p2 7.8% 9.6% 8.2% 7.3% 1,991
p20 0.8% 0.4% 0.8% 1.1% 210
p21 0.6% 0.4% 0.3% 0.6% 151
p22 0.6% 0.4% 0.3% 0.7% 145
p3 7.7% 7.8% 9.2% 7.5% 1,960
p4 7.8% 7.1% 6.1% 6.8% 1,977
p5 7.6% 7.4% 8.6% 7.5% 1,931
p6 6.9% 5.7% 6.2% 6.3% 1,740
p7 5.6% 5.8% 5.3% 5.8% 1,421
p8 4.6% 3.9% 4.3% 4.0% 1,166
p9 4.0% 3.1% 3.1% 4.4% 1,011

Row kind share per split

kind train dev calibration test train rows
rerank-docs 11.5% 6.1% 8.2% 9.7% 2,923
rerank-docs-unanswerable 2.5% 1.1% 1.1% 1.5% 639
rerank-hotpot 13.9% 11.8% 15.6% 15.5% 3,518
rerank-hotpot-removed 6.8% 6.3% 7.5% 8.8% 1,714
rerank-squad 59.1% 66.0% 60.7% 59.1% 15,007
rerank-squad-nearmiss 0.5% 1.1% 0.6% 0.7% 125
rerank-squad-nearmiss-none 0.3% 0.6% 0.3% 0.2% 74
rerank-squad-removed 4.4% 5.9% 5.6% 3.3% 1,124
rerank-squad-unanswerable 1.0% 1.1% 0.6% 1.2% 256

4. Format

split row-level format problems
train none
dev none
calibration none
test none

Options per choice row

split min median p99 max
train 6 13 40 41
dev 6 13 40 41
calibration 6 14 40 41
test 6 14 41 41

Prompt length in tokens

split measure median p99 max > 8192
train source.input_tokens (25380/25380 rows) 2165 7222 8126 0
dev source.input_tokens (900/900 rows) 2228 7084 8065 0
calibration source.input_tokens (720/720 rows) 2184 7233 7574 0
test source.input_tokens (2080/2080 rows) 2290 7282 8025 0

State key sets (train)

keys rows
question 25,380 (100.0%)

Instructions (train)

  • Canonical (the most common text) 70.3%, reworded 26.7% (80 distinct rewordings), none 3.0%. Target about 70 / 27 / 3.
  • Canonical text: "Which passage answers the question? If none of them does, choose none."
  • source instruction tag: canonical 70.3%, variant 26.7%, none 3.0%
class canonical none
<listed option> 71.4% 2.5%
none 70.3% 2.8%
p1 69.0% 3.2%
p10 72.0% 2.9%
p11 72.2% 3.4%
p12 67.0% 3.8%
p13 70.7% 3.6%
p14 69.9% 2.4%
p15 71.0% 2.9%
p16 72.1% 3.2%
p17 63.7% 3.8%
p18 72.2% 1.9%
p19 66.9% 3.7%
p2 69.9% 2.8%
p20 66.7% 3.3%
p21 67.5% 0.7%
p22 66.2% 4.8%
p3 71.8% 3.4%
p4 69.0% 3.0%
p5 70.7% 3.3%
p6 70.6% 3.1%
p7 71.9% 3.3%
p8 69.3% 2.7%
p9 70.7% 3.2%

1. Shortcuts

Phrase statistics and models use a label-stratified sample of 25,380 train rows; models are scored on the full test file (2,080 rows).

Text length by label class (main text, characters)

split class rows p10 median p90 mean
train <listed option> 1136 39 61 114 72
train none 3807 42 71 139 84
train p1 1994 38 60 108 70
train p10 804 38 61 116 72
train p11 669 37 58 105 67
train p12 603 37 60 112 69
train p13 495 39 59 108 69
train p14 492 37 59 112 68
train p15 421 39 59 113 69
train p16 377 37 60 105 68
train p17 314 39 60 115 71
train p18 320 38 61 111 70
train p19 245 36 56 96 63
train p2 1991 37 60 111 70
train p20 210 40 58 100 68
train p21 151 35 55 97 64
train p22 145 34 62 108 67
train p3 1960 37 60 108 69
train p4 1977 37 60 111 69
train p5 1931 39 61 110 71
train p6 1740 38 59 107 69
train p7 1421 37 59 111 69
train p8 1166 37 59 110 69
train p9 1011 37 62 113 71
test <listed option> 123 40 61 128 74
test none 312 46 83 147 92
test p1 169 37 61 108 68
test p10 77 40 65 122 75
test p11 56 45 67 114 76
test p12 49 31 64 117 70
test p13 36 39 57 93 63
test p14 43 33 64 110 73
test p15 24 40 75 173 100
test p16 33 37 71 108 80
test p17 33 39 67 93 68
test p18 19 32 45 98 57
test p19 24 38 57 94 62
test p2 151 37 65 121 73
test p20 22 43 64 120 75
test p21 13 50 68 161 89
test p22 14 36 71 111 75
test p3 157 41 66 119 74
test p4 142 37 59 109 69
test p5 155 36 63 114 72
test p6 132 38 63 100 69
test p7 121 37 64 119 77
test p8 84 35 59 101 67
test p9 91 34 57 98 64

By row kind (train): main-text length, length of the rest of the state, options

kind rows median chars mean chars median other-state chars median options
rerank-docs 2923 54 55 0 14
rerank-docs-unanswerable 639 62 63 0 13
rerank-hotpot 3518 111 128 0 13
rerank-hotpot-removed 1714 99 115 0 14
rerank-squad 15007 56 59 0 13
rerank-squad-nearmiss 125 59 59 0 11
rerank-squad-nearmiss-none 74 58 60 0 13
rerank-squad-removed 1124 57 58 0 13
rerank-squad-unanswerable 256 48 50 0 13

Correct option: longest / shortest / position / key

For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):

split rows correct is longest correct is shortest chance (1/listed) mean relative position (0 first, 1 last; 0.5 expected) position fifths
train 1136 9.9% 10.3% 11.8% 0.493 23% / 17% / 19% / 18% / 23%
dev 53 7.5% 7.5% 11.8% 0.554 15% / 23% / 11% / 25% / 26%
calibration 40 7.5% 12.5% 12.4% 0.493 20% / 15% / 28% / 12% / 25%
test 123 17.1% 10.6% 11.9% 0.466 29% / 17% / 20% / 11% / 24%

Correct key and position by option count

split options rows mean options top correct keys most common position (0-based)
train 6-10 9239 8.1 none 15.3%, p2 12.8%, p4 12.8%, p3 12.5%, p1 12.2% 1 (13.1%)
train 11-30 13162 17.5 none 14.6%, p6 6.1%, p1 6.0%, p5 5.7%, p2 5.7% 0 (6.6%)
train 31-80 2979 35.7 none 15.9%, p18 3.0%, p15 2.9%, p12 2.8%, p16 2.7% 14 (3.4%)
test 6-10 698 8.1 none 14.2%, p3 13.3%, p1 13.0%, p2 12.6%, p5 11.6% 3 (14.0%)
test 11-30 1097 17.5 none 16.0%, p1 6.6%, p10 6.5%, p9 6.4%, p5 5.9% 10 (8.0%)
test 31-80 285 35.9 none 13.3%, p24 5.3%, p23 4.9%, p20 4.2%, p26 3.9% 26 (5.6%)
  • Train rows whose correct listed key is the first listed key (t1/o1): 12.1%, chance 11.8%.

Option count by label class (train)

class rows min median mean max
<listed option> 1136 24 35 34.4 41
none 3807 6 13 16.3 41
p1 1994 6 10 12.2 41
p10 804 11 15 17.4 41
p11 669 12 18 19.7 41
p12 603 13 18 20.5 41
p13 495 14 19 21.5 41
p14 492 15 20 22.2 41
p15 421 16 21 23.5 41
p16 377 17 21 24.3 41
p17 314 18 21 24.1 41
p18 320 19 23 25.9 41
p19 245 20 26 26.8 41
p2 1991 6 9 11.6 41
p20 210 21 27 28.3 40
p21 151 22 28 29.5 41
p22 145 23 31 31.0 41
p3 1960 6 9 11.8 41
p4 1977 6 9 11.8 41
p5 1931 6 10 12.0 41
p6 1740 7 11 13.0 41
p7 1421 8 11 13.9 41
p8 1166 9 12 14.8 41
p9 1011 10 14 16.1 41

Source fields by label class (train)

Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.

Overall majority: 15.0%.

source field values purity top values → classes
kind 9 23.0% rerank-squad: p1 9.4%, p2 9.4%; rerank-hotpot: p4 9.4%, p5 9.2%; rerank-docs: p6 9.2%, p7 8.6%; rerank-hotpot-removed: none 100.0%; rerank-squad-removed: none 100.0%; rerank-docs-unanswerable: none 100.0%
second_opinion 2 22.9% null: p1 9.2%, p2 9.2%; deepseek-v4-flash: none 100.0%
second_pool 2 15.7% null: none 18.5%, p2 7.6%; true: p4 9.4%, p3 8.9%
dataset 3 15.0% squad2: none 8.8%, p1 8.6%; hotpotqa: none 32.8%, p4 6.3%; teacher-docs: none 17.9%, p6 7.5%
license 2 15.0% CC BY-SA 4.0: none 14.5%, p4 8.0%; generated by the local teacher (qwen3.8-flash-next): none 17.9%, p6 7.5%
url 3 15.0% https://huggingface.co/datasets/rajpurkar/squad_v2: none 8.8%, p1 8.6%; https://huggingface.co/datasets/hotpotqa/hotpot_qa: none 32.8%, p4 6.3%; : none 17.9%, p6 7.5%
revision 3 15.0% 3ffb306f725f7d2ce8394bc1873b24868140c412: none 8.8%, p1 8.6%; 1908d6afbbead072334abe2965f91bd2709910ab: none 32.8%, p4 6.3%; : none 17.9%, p6 7.5%
titles_shown 2 15.0% true: none 14.9%, p2 8.1%; false: none 15.1%, p1 8.1%
text_teacher 3 15.0% none (public data and code): none 14.3%, p1 8.0%; local: qwen3.8-flash-next: none 18.1%, p6 7.5%; hosted: qwen3.8-max: none 21.0%, p5 8.1%

Formatting by label class (main text, share of rows)

feature lowest classes highest classes
ends with ? p14 98%, p16 98%, p2 98% p21 100%, p19 99%, p18 99%
ends with . p11 0%, p13 0%, p14 0% p15 1%, p22 1%, p7 1%
ends with ! <listed option> 0%, none 0%, p1 0% p9 0%, p8 0%, p7 0%
no end punctuation p21 0%, p22 0%, p19 0% p16 2%, p14 2%, p17 2%
starts lowercase p20 0%, p22 0%, p12 0% p17 2%, p6 1%, p21 1%
all lowercase p20 0%, p22 0%, p11 0% p17 1%, p10 1%, p19 1%
has a digit p11 13%, p8 13%, p20 13% none 21%, p17 19%, <listed option> 18%
has newline <listed option> 0%, none 0%, p1 0% p9 0%, p8 0%, p7 0%
has quotes p19 2%, p21 2%, p11 2% none 7%, p16 7%, p18 7%
has markup (HTML/markdown) <listed option> 0%, none 0%, p10 0% p14 0%, p4 0%, p1 0%
has URL <listed option> 0%, none 0%, p1 0% p9 0%, p8 0%, p7 0%
non-ASCII p12 0%, p17 1%, p22 1% p21 3%, none 2%, p18 2%
non-Latin script <listed option> 0%, p1 0%, p10 0% p22 1%, p9 0%, p5 0%
emoji <listed option> 0%, none 0%, p1 0% p6 0%, p9 0%, p8 0%
ALL-CAPS word (4+) p11 1%, p6 1%, p5 1% p13 3%, p1 3%, p10 2%
contains ' - ' or — p10 0%, p12 0%, p13 0% p14 0%, p1 0%, p11 0%

Same, by row kind

feature rerank-docs rerank-docs-unanswerable rerank-hotpot rerank-hotpot-removed rerank-squad rerank-squad-nearmiss rerank-squad-nearmiss-none rerank-squad-removed rerank-squad-unanswerable
ends with ? 99% 99% 97% 97% 99% 100% 100% 98% 98%
ends with . 0% 0% 1% 1% 0% 0% 0% 0% 0%
ends with ! 0% 0% 0% 0% 0% 0% 0% 0% 0%
no end punctuation 1% 1% 2% 1% 1% 0% 0% 1% 2%
starts lowercase 3% 3% 1% 1% 1% 0% 1% 0% 0%
all lowercase 3% 2% 0% 0% 0% 0% 0% 0% 0%
has a digit 4% 9% 40% 35% 13% 16% 11% 11% 12%
has newline 0% 0% 0% 0% 0% 0% 0% 0% 0%
has quotes 0% 0% 19% 15% 2% 0% 1% 1% 1%
has markup (HTML/markdown) 0% 0% 0% 0% 0% 0% 0% 0% 0%
has URL 0% 0% 0% 0% 0% 0% 0% 0% 0%
non-ASCII 0% 0% 4% 5% 1% 0% 0% 1% 0%
non-Latin script 0% 0% 0% 0% 0% 0% 0% 0% 0%
emoji 0% 0% 0% 0% 0% 0% 0% 0% 0%
ALL-CAPS word (4+) 0% 0% 3% 3% 2% 2% 3% 2% 3%
contains ' - ' or — 0% 0% 0% 0% 0% 0% 0% 0% 0%

Over-represented words and phrases per label class (main text)

Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.

Words, <listed option>: nc z=4 0.4% vs 0.0%; albert z=4 0.5% vs 0.1%; n z=4 0.4% vs 0.0%; meetings z=4 0.3% vs 0.0%; regulators z=4 0.3% vs 0.0%; adolescent z=4 0.3% vs 0.0%; communist z=4 0.4% vs 0.0%; earliest z=4 0.6% vs 0.1%; rings z=4 0.3% vs 0.0%; philippine z=4 0.3% vs 0.0%; tertiary z=4 0.3% vs 0.0%; difficult z=3 0.4% vs 0.0%; tram z=3 0.3% vs 0.0%; specifically z=3 0.3% vs 0.0%; surviving z=3 0.3% vs 0.0%; heir z=3 0.3% vs 0.0%; hunting z=3 0.5% vs 0.1%; princess z=3 0.4% vs 0.0%; helps z=3 0.3% vs 0.0%; quake z=3 0.3% vs 0.0%

Words, none: there z=13 4.6% vs 1.1%; born z=10 4.1% vs 1.2%; available z=10 1.7% vs 0.2%; offer z=9 1.4% vs 0.2%; specific z=7 1.1% vs 0.2%; parking z=7 0.8% vs 0.0%; a z=7 27.5% vs 18.5%; discount z=7 0.8% vs 0.0%; for z=7 18.2% vs 11.6%; that z=7 10.7% vs 6.2%; team z=7 2.7% vs 1.0%; which z=6 15.8% vs 10.2%; album z=6 2.2% vs 0.9%; film z=6 5.4% vs 2.8%; founded z=6 1.8% vs 0.7%; dress z=6 0.5% vs 0.0%; county z=6 1.2% vs 0.4%; won z=5 1.5% vs 0.5%; by z=5 12.0% vs 7.8%; mobile z=5 0.5% vs 0.1%

Words, p1: lies z=4 0.3% vs 0.0%; cultures z=4 0.4% vs 0.1%; philosophical z=4 0.2% vs 0.0%; coastal z=3 0.2% vs 0.0%; foundation z=3 0.4% vs 0.1%; universities z=3 0.5% vs 0.1%; 1948 z=3 0.3% vs 0.0%; cause z=3 0.7% vs 0.2%; hands z=3 0.2% vs 0.0%; industries z=3 0.2% vs 0.0%; society z=3 0.5% vs 0.2%; vegas z=3 0.2% vs 0.0%; basis z=3 0.2% vs 0.0%; officer z=3 0.5% vs 0.1%; absorb z=3 0.2% vs 0.0%; commanding z=3 0.2% vs 0.0%; descended z=3 0.2% vs 0.0%; adding z=3 0.2% vs 0.0%; acre z=3 0.3% vs 0.0%; destroy z=3 0.3% vs 0.0%

Words, p10: initially z=5 0.7% vs 0.1%; planning z=4 0.5% vs 0.0%; courses z=4 0.4% vs 0.0%; bermuda z=4 0.6% vs 0.1%; edwards z=4 0.4% vs 0.0%; vi z=4 0.6% vs 0.1%; impact z=4 0.5% vs 0.1%; recognized z=4 0.4% vs 0.0%; attendance z=4 0.4% vs 0.0%; paul z=4 1.0% vs 0.2%; fly z=3 0.4% vs 0.0%; upload z=3 0.4% vs 0.0%; sciences z=3 0.4% vs 0.0%; 1990s z=3 0.5% vs 0.1%; america's z=3 0.4% vs 0.0%; boys z=3 0.5% vs 0.1%; studying z=3 0.4% vs 0.0%; rated z=3 0.4% vs 0.0%; highest z=3 1.0% vs 0.3%; thompson z=3 0.2% vs 0.0%

Words, p11: 1975 z=4 0.6% vs 0.1%; religion z=4 1.2% vs 0.3%; floor z=4 0.6% vs 0.1%; cartoon z=4 0.4% vs 0.0%; multiple z=4 0.7% vs 0.1%; peninsula z=4 0.4% vs 0.0%; emotion z=4 0.4% vs 0.0%; quickly z=3 0.9% vs 0.2%; graduate z=3 0.4% vs 0.0%; immigrant z=3 0.3% vs 0.0%; adoption z=3 0.3% vs 0.0%; pigment z=3 0.3% vs 0.0%; samoan z=3 0.3% vs 0.0%; doctrine z=3 0.4% vs 0.1%; 17th z=3 0.3% vs 0.0%; traits z=3 0.3% vs 0.0%; agreed z=3 0.3% vs 0.0%; mice z=3 0.3% vs 0.0%; mammals z=3 0.3% vs 0.0%; long z=3 3.7% vs 2.0%

Words, p12: mans z=4 0.5% vs 0.0%; crew z=4 0.5% vs 0.0%; ge z=4 0.5% vs 0.0%; cultural z=4 0.8% vs 0.1%; opposed z=4 0.5% vs 0.0%; largest z=4 2.2% vs 0.7%; 1899 z=4 0.3% vs 0.0%; stays z=4 0.3% vs 0.0%; purposes z=4 0.3% vs 0.0%; container z=4 0.3% vs 0.0%; dangerous z=4 0.3% vs 0.0%; inch z=4 0.3% vs 0.0%; macarthur z=4 0.3% vs 0.0%; penalty z=4 0.8% vs 0.1%; laserdisc z=3 0.5% vs 0.0%; whitehead's z=3 0.5% vs 0.0%; explain z=3 0.5% vs 0.0%; whether z=3 0.3% vs 0.0%; romanian z=3 0.3% vs 0.0%; capita z=3 0.3% vs 0.0%

Words, p13: serving z=4 0.6% vs 0.0%; connection z=4 0.6% vs 0.0%; destinations z=4 0.4% vs 0.0%; rodney z=4 0.4% vs 0.0%; montgomery z=4 0.4% vs 0.0%; millennium z=4 0.4% vs 0.0%; sentences z=4 0.4% vs 0.0%; patriot z=4 0.4% vs 0.0%; tito z=4 0.8% vs 0.1%; proto z=4 0.4% vs 0.0%; significantly z=4 0.4% vs 0.0%; link's z=4 0.4% vs 0.0%; pavilion z=4 0.4% vs 0.0%; concern z=4 0.4% vs 0.0%; ghana z=4 0.4% vs 0.0%; szlachta z=3 0.4% vs 0.0%; jordan z=3 0.4% vs 0.0%; split z=3 0.6% vs 0.1%; tuberculosis z=3 0.4% vs 0.0%; distinguished z=3 0.4% vs 0.0%

Words, p14: prior z=5 1.4% vs 0.2%; listed z=4 1.0% vs 0.1%; schwarzenegger z=4 1.0% vs 0.2%; variation z=4 0.6% vs 0.0%; oldest z=4 1.0% vs 0.2%; unicef z=4 0.4% vs 0.0%; federation z=4 0.4% vs 0.0%; input z=4 0.4% vs 0.0%; skater z=4 0.4% vs 0.0%; poultry z=3 0.4% vs 0.0%; plateau z=3 0.4% vs 0.0%; biology z=3 0.4% vs 0.0%; bills z=3 0.4% vs 0.0%; dave z=3 0.4% vs 0.0%; eu z=3 0.4% vs 0.0%; metro z=3 0.6% vs 0.1%; bell z=3 1.0% vs 0.2%; night z=3 1.0% vs 0.2%; 1860 z=3 0.4% vs 0.0%; galicia z=3 0.4% vs 0.0%

Words, p15: classification z=5 1.0% vs 0.1%; midna z=4 0.5% vs 0.0%; veto z=4 0.5% vs 0.0%; hackers z=4 0.5% vs 0.0%; climate z=4 1.0% vs 0.1%; 1952 z=4 0.5% vs 0.0%; assention z=4 0.5% vs 0.0%; banking z=4 0.5% vs 0.0%; facet z=4 0.5% vs 0.0%; rulers z=3 0.5% vs 0.0%; financially z=3 0.5% vs 0.0%; taiwan z=3 0.5% vs 0.0%; boards z=3 0.5% vs 0.0%; northwest z=3 0.5% vs 0.0%; square z=3 0.7% vs 0.1%; technologies z=3 0.5% vs 0.0%; login z=3 0.5% vs 0.0%; ask z=3 0.5% vs 0.0%; logo z=3 0.5% vs 0.0%; option z=3 0.5% vs 0.0%

Words, p16: passengers z=5 0.8% vs 0.0%; followed z=4 0.8% vs 0.0%; notice z=4 1.1% vs 0.1%; rejection z=4 0.5% vs 0.0%; expanded z=4 0.5% vs 0.0%; sale z=4 0.5% vs 0.0%; ranks z=4 0.5% vs 0.0%; protest z=4 0.5% vs 0.0%; purple z=4 0.5% vs 0.0%; country z=4 3.4% vs 1.2%; city's z=4 0.8% vs 0.1%; middle z=3 1.3% vs 0.2%; header z=3 0.5% vs 0.0%; lab z=3 0.5% vs 0.0%; luxury z=3 0.5% vs 0.0%; tank z=3 0.5% vs 0.0%; principle z=3 0.5% vs 0.0%; automatic z=3 0.5% vs 0.0%; give z=3 1.3% vs 0.3%; decision z=3 0.8% vs 0.1%

Words, p17: mid z=4 1.3% vs 0.1%; stain z=4 0.6% vs 0.0%; 1987 z=4 1.0% vs 0.1%; barcelona's z=4 0.6% vs 0.0%; indies z=4 0.6% vs 0.0%; seattle's z=4 0.6% vs 0.0%; butler z=4 0.6% vs 0.0%; specialized z=4 0.6% vs 0.0%; declined z=4 0.6% vs 0.0%; societies z=4 0.6% vs 0.0%; thames z=3 0.6% vs 0.0%; performing z=3 0.6% vs 0.0%; same z=3 1.9% vs 0.4%; agencies z=3 0.6% vs 0.0%; heritage z=3 0.6% vs 0.0%; underground z=3 0.6% vs 0.0%; ago z=3 0.6% vs 0.0%; went z=3 1.0% vs 0.1%; horse z=3 0.6% vs 0.1%; tibetan z=3 0.6% vs 0.1%

Words, p18: 33 z=5 0.9% vs 0.0%; reportedly z=4 0.6% vs 0.0%; past z=4 0.9% vs 0.1%; cutoff z=4 0.6% vs 0.0%; enact z=4 0.6% vs 0.0%; healing z=4 0.6% vs 0.0%; wire z=4 0.6% vs 0.0%; 9 z=3 0.9% vs 0.1%; what's z=3 1.6% vs 0.3%; throw z=3 0.6% vs 0.0%; 1937 z=3 0.6% vs 0.0%; parental z=3 0.6% vs 0.0%; tablets z=3 0.6% vs 0.0%; empires z=3 0.6% vs 0.0%; range z=3 1.2% vs 0.2%; july z=3 0.9% vs 0.1%; shape z=3 0.6% vs 0.0%; nickname z=3 1.2% vs 0.2%; guam z=3 0.6% vs 0.0%; devices z=3 0.9% vs 0.1%

Words, p19: jay z=4 0.8% vs 0.0%; xbox z=4 0.8% vs 0.0%; quantum z=4 0.8% vs 0.0%; paperwork z=4 0.8% vs 0.0%; playstation z=4 0.8% vs 0.0%; platform z=4 0.8% vs 0.0%; ps3 z=3 0.8% vs 0.0%; slaves z=3 0.8% vs 0.1%; aspect z=3 0.8% vs 0.1%; account z=3 2.0% vs 0.4%; actors z=3 0.8% vs 0.1%; songs z=3 1.2% vs 0.2%; spectre z=3 0.8% vs 0.1%; joint z=3 0.8% vs 0.1%; jazz z=3 0.8% vs 0.1%; testing z=3 0.8% vs 0.1%; die z=3 1.2% vs 0.2%; city's z=3 0.8% vs 0.1%; 1985 z=3 0.8% vs 0.1%; 2011 z=3 1.6% vs 0.3%

Words, p2: argument z=4 0.3% vs 0.0%; congressional z=4 0.3% vs 0.0%; kitchen z=4 0.2% vs 0.0%; n z=4 0.3% vs 0.0%; obesity z=4 0.2% vs 0.0%; empty z=3 0.3% vs 0.0%; stated z=3 0.4% vs 0.1%; sanskrit z=3 0.3% vs 0.1%; help z=3 0.7% vs 0.2%; transportation z=3 0.3% vs 0.0%; books z=3 0.5% vs 0.2%; 1938 z=3 0.2% vs 0.0%; decides z=3 0.2% vs 0.0%; terminate z=3 0.2% vs 0.0%; refused z=3 0.2% vs 0.0%; kathmandu's z=3 0.2% vs 0.0%; motorcycle z=3 0.2% vs 0.0%; schumann z=3 0.2% vs 0.0%; recession z=3 0.2% vs 0.0%; java z=3 0.2% vs 0.0%

Words, p20: canada z=5 1.9% vs 0.1%; resident z=5 1.4% vs 0.1%; rangers z=4 1.0% vs 0.0%; john's z=4 1.0% vs 0.0%; reserve z=4 1.0% vs 0.0%; equipment z=4 1.0% vs 0.0%; transport z=4 1.0% vs 0.0%; resource z=3 1.0% vs 0.1%; low z=3 1.4% vs 0.2%; assistant z=3 1.0% vs 0.1%; issued z=3 1.0% vs 0.1%; projects z=3 1.0% vs 0.1%; fail z=3 1.0% vs 0.1%; metro z=3 1.0% vs 0.1%; mall z=3 1.0% vs 0.1%; chopin z=3 1.9% vs 0.3%; joined z=3 1.0% vs 0.1%; emergency z=3 1.0% vs 0.1%; defeat z=3 1.0% vs 0.1%; hub z=3 1.0% vs 0.1%

Words, p21: statistics z=5 1.3% vs 0.0%; failed z=4 2.0% vs 0.1%; lawyer z=4 1.3% vs 0.1%; retail z=4 1.3% vs 0.1%; colors z=4 1.3% vs 0.1%; cubs z=3 1.3% vs 0.1%; bowl z=3 1.3% vs 0.1%; heart z=3 1.3% vs 0.1%; francis z=3 1.3% vs 0.1%; defined z=3 1.3% vs 0.1%; faith z=3 1.3% vs 0.1%; records z=3 2.0% vs 0.2%; profession z=3 1.3% vs 0.1%; white z=3 2.0% vs 0.2%; files z=3 1.3% vs 0.1%; 500 z=3 1.3% vs 0.1%; anti z=3 1.3% vs 0.1%; address z=3 1.3% vs 0.1%; jaws z=3 0.7% vs 0.0%; investor z=3 0.7% vs 0.0%

Words, p22: mistake z=5 1.4% vs 0.0%; strategy z=4 1.4% vs 0.0%; regards z=4 1.4% vs 0.1%; coast z=3 1.4% vs 0.1%; 1988 z=3 1.4% vs 0.1%; zone z=3 1.4% vs 0.1%; 1947 z=3 0.7% vs 0.0%; rebel z=3 0.7% vs 0.0%; katrina z=3 0.7% vs 0.0%; sells z=3 0.7% vs 0.0%; everyday z=3 0.7% vs 0.0%; judaism z=3 0.7% vs 0.0%; feeling z=3 0.7% vs 0.0%; shelters z=3 0.7% vs 0.0%; showtime z=3 0.7% vs 0.0%; wit z=3 0.7% vs 0.0%; exhibit z=3 0.7% vs 0.0%; alcoholic z=3 0.7% vs 0.0%; arch z=3 0.7% vs 0.0%; strasbourg z=3 0.7% vs 0.0%

Words, p3: refer z=4 0.4% vs 0.1%; consumer z=4 0.3% vs 0.0%; spot z=4 0.3% vs 0.0%; frequency z=4 0.4% vs 0.1%; allies z=4 0.2% vs 0.0%; boston z=4 0.6% vs 0.2%; liam z=3 0.2% vs 0.0%; pre z=3 0.4% vs 0.1%; consonants z=3 0.2% vs 0.0%; infringement z=3 0.2% vs 0.0%; appear z=3 0.5% vs 0.1%; higher z=3 0.4% vs 0.1%; aggression z=3 0.2% vs 0.0%; kid z=3 0.2% vs 0.0%; aspirated z=3 0.2% vs 0.0%; allowing z=3 0.2% vs 0.0%; acting z=3 0.3% vs 0.0%; spoke z=3 0.2% vs 0.0%; details z=3 0.2% vs 0.0%; 1857 z=3 0.2% vs 0.0%

Words, p4: student z=4 0.6% vs 0.2%; excessive z=4 0.3% vs 0.0%; friars z=3 0.2% vs 0.0%; clubs z=3 0.2% vs 0.0%; wider z=3 0.2% vs 0.0%; emerge z=3 0.2% vs 0.0%; combine z=3 0.2% vs 0.0%; rhine z=3 0.2% vs 0.0%; innovator z=3 0.2% vs 0.0%; 1891 z=3 0.2% vs 0.0%; divisions z=3 0.3% vs 0.0%; isles z=3 0.3% vs 0.0%; sand z=3 0.3% vs 0.0%; torch z=3 0.3% vs 0.1%; bern z=3 0.2% vs 0.0%; attempted z=3 0.2% vs 0.0%; imperial's z=3 0.2% vs 0.0%; collaborate z=3 0.2% vs 0.0%; settlements z=3 0.2% vs 0.0%; coverage z=3 0.3% vs 0.1%

Words, p5: qing z=4 0.4% vs 0.1%; endangered z=4 0.3% vs 0.0%; relationship z=4 0.5% vs 0.1%; vote z=3 0.4% vs 0.1%; needed z=3 0.3% vs 0.1%; throughout z=3 0.3% vs 0.1%; graphics z=3 0.2% vs 0.0%; rule z=3 0.6% vs 0.2%; vertical z=3 0.2% vs 0.0%; relay z=3 0.3% vs 0.0%; torch z=3 0.3% vs 0.1%; pain z=3 0.4% vs 0.1%; increase z=3 0.4% vs 0.1%; sequel z=3 0.2% vs 0.0%; chuck z=3 0.2% vs 0.0%; households z=3 0.2% vs 0.0%; mtv z=3 0.2% vs 0.0%; displayed z=3 0.2% vs 0.0%; applications z=3 0.2% vs 0.0%; subjects z=3 0.3% vs 0.1%

Words, p6: hd z=4 0.3% vs 0.0%; control z=4 0.7% vs 0.2%; 000 z=4 0.5% vs 0.1%; server z=4 0.3% vs 0.0%; montini z=4 0.3% vs 0.0%; rule z=4 0.7% vs 0.2%; half z=4 0.5% vs 0.1%; foreign z=3 0.5% vs 0.1%; returned z=3 0.2% vs 0.0%; can't z=3 0.3% vs 0.1%; brown z=3 0.4% vs 0.1%; germany z=3 0.5% vs 0.1%; need z=3 2.4% vs 1.4%; 24 z=3 0.3% vs 0.1%; hours z=3 0.3% vs 0.1%; forget z=3 0.3% vs 0.1%; eliminate z=3 0.2% vs 0.0%; accredited z=3 0.2% vs 0.0%; kerry's z=3 0.2% vs 0.0%; suburban z=3 0.2% vs 0.0%

Words, p7: it z=4 4.4% vs 2.7%; posses z=4 0.2% vs 0.0%; couple z=3 0.3% vs 0.0%; frequencies z=3 0.2% vs 0.0%; nationalist z=3 0.2% vs 0.0%; handle z=3 0.4% vs 0.1%; breaks z=3 0.3% vs 0.0%; younger z=3 0.3% vs 0.0%; philosopher z=3 0.4% vs 0.1%; photographer z=3 0.2% vs 0.0%; warm z=3 0.2% vs 0.0%; read z=3 0.5% vs 0.1%; biggest z=3 0.5% vs 0.1%; religions z=3 0.3% vs 0.0%; okay z=3 0.2% vs 0.0%; asked z=3 0.2% vs 0.0%; invaded z=3 0.2% vs 0.0%; get z=3 2.4% vs 1.4%; constitutional z=3 0.3% vs 0.0%; concerns z=3 0.3% vs 0.0%

Words, p8: attempt z=5 0.7% vs 0.1%; makes z=4 0.8% vs 0.2%; becoming z=4 0.3% vs 0.0%; concentrated z=4 0.3% vs 0.0%; order z=4 1.0% vs 0.3%; currency z=4 0.4% vs 0.1%; mammals z=4 0.3% vs 0.0%; membership z=3 0.3% vs 0.0%; crimes z=3 0.3% vs 0.0%; silent z=3 0.3% vs 0.0%; neighborhood z=3 0.3% vs 0.0%; southern z=3 0.7% vs 0.2%; humanism z=3 0.3% vs 0.1%; winner z=3 0.4% vs 0.1%; fiction z=3 0.6% vs 0.2%; claim z=3 0.9% vs 0.3%; lands z=3 0.3% vs 0.0%; 1925 z=3 0.3% vs 0.0%; green z=3 0.5% vs 0.1%; islam z=3 0.3% vs 0.1%

Words, p9: intake z=4 0.3% vs 0.0%; northwestern z=4 0.5% vs 0.1%; cricketer z=4 0.3% vs 0.0%; fibers z=4 0.3% vs 0.0%; ad z=4 0.4% vs 0.0%; future z=3 0.5% vs 0.1%; orders z=3 0.3% vs 0.0%; nigeria z=3 0.5% vs 0.1%; firm z=3 0.3% vs 0.0%; organized z=3 0.3% vs 0.0%; let z=3 0.3% vs 0.0%; missing z=3 0.3% vs 0.0%; critical z=3 0.4% vs 0.1%; printing z=3 0.2% vs 0.0%; chapel z=3 0.2% vs 0.0%; unincorporated z=3 0.2% vs 0.0%; protocols z=3 0.2% vs 0.0%; customs z=3 0.2% vs 0.0%; interdisciplinary z=3 0.2% vs 0.0%; 1839 z=3 0.2% vs 0.0%

2–4-word phrases, <listed option>: company based z=4 0.4% vs 0.0%; of solar z=4 0.4% vs 0.0%; was previously z=4 0.4% vs 0.0%; how long has z=4 0.4% vs 0.0%; long has z=4 0.4% vs 0.0%; what person z=4 0.4% vs 0.0%; during a z=4 0.4% vs 0.0%; is the earliest z=4 0.4% vs 0.0%; the earliest z=4 0.6% vs 0.1%; company based in z=4 0.3% vs 0.0%; of prince z=4 0.3% vs 0.0%; the biggest z=4 0.5% vs 0.1%; what property z=4 0.3% vs 0.0%; a report z=4 0.3% vs 0.0%; is it called z=4 0.4% vs 0.0%; what is it called z=4 0.4% vs 0.0%; tertiary education z=4 0.3% vs 0.0%; a video game z=4 0.3% vs 0.0%; game released z=4 0.3% vs 0.0%; how many people in z=4 0.3% vs 0.0%

2–4-word phrases, none: is there z=14 3.3% vs 0.4%; is there a z=14 3.0% vs 0.3%; there a z=13 3.0% vs 0.3%; does the z=9 5.1% vs 2.0%; available for z=8 1.0% vs 0.0%; in which z=7 3.5% vs 1.5%; by an z=6 0.9% vs 0.2%; a discount z=6 0.7% vs 0.0%; was born z=6 1.3% vs 0.3%; policy cover z=6 0.6% vs 0.0%; by a z=6 1.5% vs 0.4%; are there any z=6 0.6% vs 0.0%; the api z=6 0.5% vs 0.0%; are there z=6 1.1% vs 0.3%; there any z=6 0.6% vs 0.1%; is a z=6 4.2% vs 2.0%; in what z=6 6.5% vs 3.6%; discount for z=5 0.5% vs 0.0%; born in z=5 1.3% vs 0.4%; the film z=5 1.4% vs 0.5%

2–4-word phrases, p1: the las vegas z=4 0.2% vs 0.0%; the las z=4 0.2% vs 0.0%; of the federal z=4 0.2% vs 0.0%; to drive z=3 0.2% vs 0.0%; who commanded the z=3 0.2% vs 0.0%; who commanded z=3 0.2% vs 0.0%; can cause z=3 0.2% vs 0.0%; to be z=3 2.1% vs 1.2%; las vegas z=3 0.2% vs 0.0%; become the first z=3 0.2% vs 0.0%; team to win z=3 0.2% vs 0.0%; my plot z=3 0.2% vs 0.0%; chief executive officer of z=3 0.2% vs 0.0%; on my plot z=3 0.2% vs 0.0%; executive officer of z=3 0.2% vs 0.0%; and led z=3 0.2% vs 0.0%; american businessman z=3 0.2% vs 0.0%; has held z=3 0.2% vs 0.0%; of the first single z=3 0.2% vs 0.0%; that isn't z=3 0.2% vs 0.0%

2–4-word phrases, p10: paul vi z=4 0.6% vs 0.1%; what club z=4 0.4% vs 0.0%; maximum file z=4 0.4% vs 0.0%; the self titled z=4 0.4% vs 0.0%; did paul vi z=4 0.4% vs 0.0%; in the 1990s z=4 0.4% vs 0.0%; an issue z=4 0.4% vs 0.0%; of representatives z=3 0.4% vs 0.0%; in the 19th z=3 0.4% vs 0.0%; the self z=3 0.4% vs 0.0%; house of representatives z=3 0.4% vs 0.0%; self titled z=3 0.4% vs 0.0%; i am z=3 0.5% vs 0.1%; what happens to z=3 0.6% vs 0.1%; did paul z=3 0.4% vs 0.0%; the 1990s z=3 0.4% vs 0.0%; of one z=3 0.4% vs 0.0%; and author z=3 0.4% vs 0.0%; happens to z=3 0.6% vs 0.1%; was marked z=3 0.2% vs 0.0%

2–4-word phrases, p11: long can z=5 1.0% vs 0.1%; how long can z=5 1.0% vs 0.1%; decade was z=4 0.4% vs 0.0%; is the number z=4 0.4% vs 0.0%; how quickly z=4 0.9% vs 0.1%; in 1975 z=4 0.4% vs 0.0%; when does the z=4 0.9% vs 0.2%; can a z=4 0.7% vs 0.1%; population in z=3 0.4% vs 0.0%; that was used z=3 0.3% vs 0.0%; in what period z=3 0.3% vs 0.0%; is the number of z=3 0.3% vs 0.0%; which doctrine z=3 0.3% vs 0.0%; to keep the z=3 0.3% vs 0.0%; which spanish z=3 0.3% vs 0.0%; file is z=3 0.3% vs 0.0%; of city z=3 0.3% vs 0.0%; my plot z=3 0.3% vs 0.0%; what is the number z=3 0.3% vs 0.0%; what field z=3 0.3% vs 0.0%

2–4-word phrases, p12: penalty if z=4 0.5% vs 0.0%; what song was z=4 0.5% vs 0.0%; need to pass z=4 0.5% vs 0.0%; on november z=4 0.5% vs 0.0%; released by z=4 0.7% vs 0.1%; to pass z=4 0.5% vs 0.0%; song was z=4 0.5% vs 0.0%; the wife z=4 0.5% vs 0.0%; the wife of z=4 0.5% vs 0.0%; in miami z=4 0.5% vs 0.0%; a limit on how z=4 0.5% vs 0.0%; there a limit on z=4 0.5% vs 0.0%; aircraft that z=3 0.3% vs 0.0%; the dead z=3 0.3% vs 0.0%; i need to pass z=3 0.3% vs 0.0%; in when z=3 0.3% vs 0.0%; in mathematics z=3 0.3% vs 0.0%; the book the z=3 0.3% vs 0.0%; its obligations z=3 0.3% vs 0.0%; a native z=3 0.3% vs 0.0%

2–4-word phrases, p13: of all the z=5 0.6% vs 0.0%; name of the book z=4 0.6% vs 0.0%; the major z=4 0.8% vs 0.1%; was the largest z=4 0.6% vs 0.0%; the pacific z=4 0.4% vs 0.0%; credit for z=4 0.4% vs 0.0%; like a z=4 0.4% vs 0.0%; what is the earliest z=4 0.4% vs 0.0%; i get if my z=4 0.4% vs 0.0%; proto indo z=4 0.4% vs 0.0%; hotel in z=4 0.4% vs 0.0%; about how much of z=4 0.4% vs 0.0%; the cincinnati z=4 0.4% vs 0.0%; at any z=4 0.4% vs 0.0%; get if my z=4 0.4% vs 0.0%; apply for the z=3 0.4% vs 0.0%; about how much z=3 0.4% vs 0.0%; of the major z=3 0.4% vs 0.0%; american drama film z=3 0.4% vs 0.0%; the people in z=3 0.4% vs 0.0%

2–4-word phrases, p14: prior to z=4 1.2% vs 0.2%; result of z=4 1.0% vs 0.1%; the result of z=4 0.6% vs 0.0%; band from z=4 0.6% vs 0.0%; record high z=4 0.4% vs 0.0%; update my z=4 0.4% vs 0.0%; the plant z=4 0.4% vs 0.0%; look at z=4 0.4% vs 0.0%; in 1860 z=4 0.4% vs 0.0%; which season z=4 0.4% vs 0.0%; stage name z=4 0.4% vs 0.0%; was the first president z=4 0.4% vs 0.0%; what colors z=4 0.4% vs 0.0%; will you z=4 0.4% vs 0.0%; into a z=3 0.6% vs 0.1%; on time z=3 0.4% vs 0.0%; the actor born z=3 0.4% vs 0.0%; was the highest z=3 0.4% vs 0.0%; the chicago z=3 0.4% vs 0.0%; the result z=3 0.6% vs 0.1%

2–4-word phrases, p15: a 2005 z=4 0.7% vs 0.0%; the airport z=4 0.7% vs 0.0%; of valencia z=4 0.5% vs 0.0%; type of climate z=4 0.5% vs 0.0%; year at z=4 0.5% vs 0.0%; year at the z=4 0.5% vs 0.0%; what type of climate z=4 0.5% vs 0.0%; also played for z=4 0.5% vs 0.0%; else in z=4 0.5% vs 0.0%; all india z=4 0.5% vs 0.0%; who worked with z=4 0.5% vs 0.0%; decrease in z=4 0.5% vs 0.0%; organization was z=4 0.5% vs 0.0%; be considered z=4 0.7% vs 0.1%; how much in z=4 0.5% vs 0.0%; american film director z=4 0.5% vs 0.0%; what facet z=4 0.5% vs 0.0%; where else z=4 0.5% vs 0.0%; much in z=4 0.5% vs 0.0%; facet of z=4 0.5% vs 0.0%

2–4-word phrases, p16: what country z=5 2.4% vs 0.4%; a singer z=4 0.8% vs 0.0%; the swiss z=4 0.8% vs 0.0%; what role did z=4 0.8% vs 0.0%; role did z=4 0.8% vs 0.1%; how many passengers z=4 0.5% vs 0.0%; high middle z=4 0.5% vs 0.0%; what level z=4 0.5% vs 0.0%; sale of z=4 0.5% vs 0.0%; to and z=4 0.5% vs 0.0%; give for z=4 0.5% vs 0.0%; the sale z=4 0.5% vs 0.0%; what level of z=4 0.5% vs 0.0%; rejection of z=4 0.5% vs 0.0%; many passengers z=4 0.5% vs 0.0%; high middle ages z=4 0.5% vs 0.0%; the sale of z=4 0.5% vs 0.0%; much notice do i z=4 0.8% vs 0.1%; notice do i need z=4 0.8% vs 0.1%; notice do i z=4 0.8% vs 0.1%

2–4-word phrases, p17: start in z=4 0.6% vs 0.0%; which singer z=4 0.6% vs 0.0%; that he z=4 0.6% vs 0.0%; performing arts z=4 0.6% vs 0.0%; does seattle z=4 0.6% vs 0.0%; in 1999 z=4 0.6% vs 0.0%; mid 20th z=4 0.6% vs 0.0%; footballer who plays z=4 0.6% vs 0.0%; mid 20th century z=4 0.6% vs 0.0%; what are these z=4 0.6% vs 0.0%; how many times did z=4 0.6% vs 0.0%; many times did z=4 0.6% vs 0.0%; developed in z=4 0.6% vs 0.0%; a performance z=4 0.6% vs 0.0%; roles as z=4 0.6% vs 0.0%; times did z=4 0.6% vs 0.0%; are more z=4 0.6% vs 0.0%; in the 1990s z=3 0.6% vs 0.0%; went to z=3 0.6% vs 0.0%; the same z=3 1.9% vs 0.4%

2–4-word phrases, p18: in the past z=5 0.9% vs 0.0%; the past z=5 0.9% vs 0.0%; should i throw z=4 0.6% vs 0.0%; it has been z=4 0.6% vs 0.0%; to enact z=4 0.6% vs 0.0%; of northern z=4 0.6% vs 0.0%; originally called z=4 0.6% vs 0.0%; should i throw away z=4 0.6% vs 0.0%; the nickname z=4 1.2% vs 0.1%; were killed z=4 0.9% vs 0.1%; above the z=4 0.6% vs 0.0%; of the 20th century z=4 0.6% vs 0.0%; of healing z=4 0.6% vs 0.0%; book of healing z=4 0.6% vs 0.0%; was george z=4 0.6% vs 0.0%; the nickname of z=4 0.9% vs 0.1%; for when z=4 0.6% vs 0.0%; why has z=4 0.6% vs 0.0%; i throw z=4 0.6% vs 0.0%; throw away z=4 0.6% vs 0.0%

2–4-word phrases, p19: the xbox z=4 0.8% vs 0.0%; population was z=4 0.8% vs 0.0%; into which z=4 0.8% vs 0.0%; featured in the z=4 1.2% vs 0.1%; the most recent z=4 0.8% vs 0.0%; of large z=4 0.8% vs 0.0%; france and z=4 0.8% vs 0.0%; do the z=4 2.4% vs 0.4%; of mary z=4 0.8% vs 0.0%; most recent z=4 0.8% vs 0.0%; the brother of z=4 0.8% vs 0.0%; times was z=4 0.8% vs 0.0%; prime minister of z=4 0.8% vs 0.0%; the brother z=4 0.8% vs 0.0%; at the time of z=3 0.8% vs 0.0%; the city's z=3 0.8% vs 0.0%; is the oldest z=3 0.8% vs 0.0%; the leader of the z=3 0.8% vs 0.0%; how long did the z=3 0.8% vs 0.1%; long did the z=3 0.8% vs 0.1%

2–4-word phrases, p2: to give z=4 0.4% vs 0.1%; take place z=4 0.7% vs 0.2%; the 19th century z=3 0.4% vs 0.1%; are found z=3 0.2% vs 0.0%; the border z=3 0.2% vs 0.0%; get my money z=3 0.4% vs 0.1%; get my money back z=3 0.4% vs 0.1%; came to z=3 0.3% vs 0.0%; who decides z=3 0.2% vs 0.0%; of income z=3 0.2% vs 0.0%; of bermuda's z=3 0.2% vs 0.0%; the main character of z=3 0.2% vs 0.0%; main character of z=3 0.2% vs 0.0%; km north z=3 0.2% vs 0.0%; almost all z=3 0.2% vs 0.0%; replaced the z=3 0.2% vs 0.0%; who stated that z=3 0.2% vs 0.0%; was dedicated to z=3 0.2% vs 0.0%; who stated z=3 0.2% vs 0.0%; who was the father z=3 0.2% vs 0.0%

2–4-word phrases, p20: the notre dame z=4 1.0% vs 0.0%; the assistant z=4 1.0% vs 0.0%; the defeat z=4 1.0% vs 0.0%; take part z=4 1.0% vs 0.0%; the notre z=4 1.0% vs 0.0%; take part in z=4 1.0% vs 0.0%; the defeat of z=4 1.0% vs 0.0%; are in the z=4 1.4% vs 0.1%; defeat of z=4 1.0% vs 0.0%; how large is z=4 1.0% vs 0.0%; of the new york z=4 1.0% vs 0.0%; in canada z=4 1.0% vs 0.0%; large is z=4 1.0% vs 0.0%; what will z=4 1.0% vs 0.0%; was known for z=4 1.0% vs 0.0%; up a z=3 1.0% vs 0.0%; are in z=3 1.9% vs 0.2%; how large z=3 1.0% vs 0.1%; co wrote z=3 1.0% vs 0.1%; the new york z=3 1.4% vs 0.1%

2–4-word phrases, p21: why would z=4 1.3% vs 0.0%; anti aircraft z=4 1.3% vs 0.0%; is written z=4 1.3% vs 0.0%; what in the z=4 1.3% vs 0.0%; who recorded the z=4 1.3% vs 0.0%; who recorded z=4 1.3% vs 0.0%; a failed z=4 1.3% vs 0.0%; recorded the z=4 1.3% vs 0.0%; home of the z=4 1.3% vs 0.0%; what channel z=4 1.3% vs 0.0%; my own z=4 1.3% vs 0.0%; the cubs z=4 1.3% vs 0.1%; the science z=4 1.3% vs 0.1%; home of z=4 1.3% vs 0.1%; how did z=3 2.6% vs 0.4%; why did the z=3 1.3% vs 0.1%; where are the z=3 1.3% vs 0.1%; the ruler of z=3 0.7% vs 0.0%; the restaurant z=3 0.7% vs 0.0%; which location z=3 0.7% vs 0.0%

2–4-word phrases, p22: known as a z=5 1.4% vs 0.0%; has there been z=4 1.4% vs 0.0%; has there z=4 1.4% vs 0.0%; there been z=4 1.4% vs 0.0%; in regards z=4 1.4% vs 0.0%; in regards to z=4 1.4% vs 0.0%; regards to z=4 1.4% vs 0.0%; was part of z=3 1.4% vs 0.1%; was part z=3 1.4% vs 0.1%; charge the battery z=3 0.7% vs 0.0%; served in the z=3 0.7% vs 0.0%; in french z=3 0.7% vs 0.0%; a 1985 z=3 0.7% vs 0.0%; university did z=3 0.7% vs 0.0%; a character named z=3 0.7% vs 0.0%; yongle emperor z=3 0.7% vs 0.0%; show hosted z=3 0.7% vs 0.0%; show hosted by z=3 0.7% vs 0.0%; in taiwan z=3 0.7% vs 0.0%; a response z=3 0.7% vs 0.0%

2–4-word phrases, p3: day was z=4 0.3% vs 0.0%; refer to z=4 0.4% vs 0.1%; start to z=4 0.2% vs 0.0%; child of z=4 0.2% vs 0.0%; what day z=3 0.4% vs 0.1%; what day was z=3 0.2% vs 0.0%; secretary of state z=3 0.2% vs 0.0%; took place on z=3 0.2% vs 0.0%; did the united z=3 0.2% vs 0.0%; a branch z=3 0.2% vs 0.0%; how much was the z=3 0.2% vs 0.0%; cap on z=3 0.2% vs 0.0%; to london z=3 0.2% vs 0.0%; much was the z=3 0.2% vs 0.0%; died from z=3 0.2% vs 0.0%; what tradition z=3 0.2% vs 0.0%; acting in z=3 0.2% vs 0.0%; did general z=3 0.2% vs 0.0%; is another name for z=3 0.4% vs 0.1%; on how much z=3 0.3% vs 0.0%

2–4-word phrases, p4: life of z=4 0.3% vs 0.0%; played at z=3 0.3% vs 0.0%; working on z=3 0.2% vs 0.0%; torch relay z=3 0.2% vs 0.0%; a university z=3 0.3% vs 0.1%; adult contemporary z=3 0.3% vs 0.0%; in western z=3 0.2% vs 0.0%; hit the z=3 0.2% vs 0.0%; was featured on z=3 0.2% vs 0.0%; attempted to z=3 0.2% vs 0.0%; did the u s z=3 0.2% vs 0.0%; did the u z=3 0.2% vs 0.0%; how is z=3 0.6% vs 0.2%; type of z=3 2.6% vs 1.6%; many people were z=3 0.4% vs 0.1%; how many people were z=3 0.4% vs 0.1%; and sand z=3 0.2% vs 0.0%; coverage of z=3 0.2% vs 0.0%; being what z=3 0.2% vs 0.0%; a period of z=3 0.2% vs 0.0%

2–4-word phrases, p5: if an z=4 0.2% vs 0.0%; produced by z=4 0.7% vs 0.2%; of the current z=4 0.2% vs 0.0%; from a z=4 0.7% vs 0.3%; behind the z=3 0.2% vs 0.0%; other artists z=3 0.2% vs 0.0%; activity in z=3 0.2% vs 0.0%; is given to z=3 0.2% vs 0.0%; of plymouth z=3 0.2% vs 0.0%; did the qing z=3 0.2% vs 0.0%; of the roman z=3 0.2% vs 0.0%; if i don't pick z=3 0.2% vs 0.0%; i don't pick z=3 0.2% vs 0.0%; don't pick z=3 0.2% vs 0.0%; carry the z=3 0.2% vs 0.0%; song for z=3 0.2% vs 0.0%; war of the z=3 0.2% vs 0.0%; the qing z=3 0.3% vs 0.0%; the torch z=3 0.3% vs 0.0%; to many z=3 0.2% vs 0.0%

2–4-word phrases, p6: control of z=4 0.4% vs 0.1%; happens if i forget z=4 0.3% vs 0.0%; i need z=4 2.1% vs 1.1%; do i z=4 3.7% vs 2.3%; for creating z=4 0.2% vs 0.0%; the us air z=4 0.2% vs 0.0%; the us air force z=4 0.2% vs 0.0%; forget to z=4 0.3% vs 0.1%; i forget to z=4 0.3% vs 0.1%; if i forget to z=4 0.3% vs 0.1%; if i forget z=3 0.3% vs 0.1%; i forget z=3 0.3% vs 0.1%; host the z=3 0.2% vs 0.0%; i need to z=3 1.7% vs 0.9%; first in z=3 0.2% vs 0.0%; september 2014 z=3 0.2% vs 0.0%; renew my z=3 0.2% vs 0.0%; cost if z=3 0.2% vs 0.0%; giving the z=3 0.2% vs 0.0%; returned for z=3 0.2% vs 0.0%

2–4-word phrases, p7: when does z=4 0.9% vs 0.3%; couple of z=4 0.3% vs 0.0%; a couple of z=4 0.3% vs 0.0%; was the founder z=4 0.3% vs 0.0%; was the founder of z=4 0.3% vs 0.0%; to get z=4 0.9% vs 0.3%; a couple z=4 0.3% vs 0.0%; does my z=4 0.5% vs 0.1%; write in z=3 0.2% vs 0.0%; can my z=3 0.4% vs 0.1%; was the leader z=3 0.2% vs 0.0%; who was the leader z=3 0.2% vs 0.0%; have to use z=3 0.2% vs 0.0%; was the leader of z=3 0.2% vs 0.0%; present day z=3 0.3% vs 0.0%; all the z=3 0.4% vs 0.1%; top level z=3 0.2% vs 0.0%; go into z=3 0.2% vs 0.0%; writer who z=3 0.2% vs 0.0%; who was the founder z=3 0.2% vs 0.0%

2–4-word phrases, p8: the time of the z=4 0.3% vs 0.0%; time of the z=4 0.3% vs 0.0%; at the time z=4 0.4% vs 0.0%; of humanism z=4 0.3% vs 0.0%; the winner of the z=4 0.3% vs 0.0%; at the time of z=4 0.3% vs 0.0%; in india in z=4 0.3% vs 0.0%; winner of z=4 0.4% vs 0.1%; what percent of z=4 0.6% vs 0.1%; what is one of z=4 0.4% vs 0.1%; the time of z=4 0.4% vs 0.1%; how long are z=4 0.3% vs 0.0%; is played by z=4 0.3% vs 0.0%; india in z=4 0.3% vs 0.0%; long are z=4 0.3% vs 0.0%; a show z=3 0.3% vs 0.0%; the winner of z=3 0.3% vs 0.0%; winner of the z=3 0.3% vs 0.0%; the winner z=3 0.3% vs 0.0%; focus of z=3 0.3% vs 0.0%

2–4-word phrases, p9: one country z=4 0.4% vs 0.0%; law review z=4 0.3% vs 0.0%; to cancel my z=4 0.3% vs 0.0%; was the most z=4 0.5% vs 0.1%; which city was z=3 0.3% vs 0.0%; had to z=3 0.3% vs 0.0%; to cancel z=3 0.4% vs 0.0%; in 2017 z=3 0.3% vs 0.0%; what does the term z=3 0.3% vs 0.0%; city was z=3 0.6% vs 0.1%; and two z=3 0.3% vs 0.0%; was william z=3 0.3% vs 0.0%; the book of z=3 0.3% vs 0.0%; does the term z=3 0.3% vs 0.0%; agree to z=3 0.3% vs 0.0%; how can z=3 0.5% vs 0.1%; city was the z=3 0.4% vs 0.1%; loosely based on z=3 0.3% vs 0.0%; the largest city z=3 0.3% vs 0.0%; loosely based z=3 0.3% vs 0.0%

Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):

  • none

Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):

  • all rows (25,380 rows): none

Shortcut models

Predicting the label class on test (2,080 rows). Chance 4.2%, majority class ('none') 15.0%; balanced chance 4.2%.

model (logistic regression, trained on the train sample) test accuracy balanced accuracy (mean recall)
bag of words, whole state (words and word pairs) 14.0% 5.1%
bag of words, main text only (question) 14.0% 5.1%
surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) 18.1% 7.8%
surface features of the main text only 15.1% 4.3%

Strongest single surface features (logistic regression on one feature, balanced accuracy on test):

feature accuracy balanced accuracy
n_options(log) 17.9% 7.6%
chars(log) 15.1% 4.3%
state_chars(log) 15.1% 4.3%
words(log) 15.1% 4.2%
count_! 15.0% 4.2%
upper_ratio 15.0% 4.2%
digit_ratio 15.0% 4.2%
nonascii_ratio 15.0% 4.2%
nonlatin 15.0% 4.2%
emoji 15.0% 4.2%

Other state fields alone (predicting the label class on test from one field, without the main text):

field treated as accuracy balanced accuracy

No-meaning option picker: a logistic ranker scores each option from its position, length, key type, fixed-option identity and shape (commas, brackets, capitals), never reading the state or the option's words, and picks the top option per row.

  • Test accuracy 15.0% against uniform chance 7.9% (this includes the fixed options, whose share is a class prior).
  • Among rows whose answer is a listed option (123), picking only among listed options: 10.6% against chance 11.9%.

2. Duplicates and split separation

Families shared between splits

splits shared families examples
train ∩ dev 0
train ∩ calibration 0
train ∩ test 0
dev ∩ calibration 0
dev ∩ test 0
calibration ∩ test 0
  • Train rows whose main text repeats an earlier row's (normalised): 5,123 (20.2%); groups: 5,107; largest group 6.
  • Train rows identical in the whole prompt (state, options, instructions): 0.
  • Identical whole prompt, different answer: 0 groups (0 rows).
  • Identical main text, different label class: 4580 groups (9176 rows). (Can be legitimate when the rest of the state or the options differ.)
    • "What are the most common side effects?" ×6: p2 2, p7 2, p1 1, p13 1
    • "How long does my API key last before it expires?" ×4: p6 1, p10 1, p7 1, p8 1
    • "Who published A Theory of Justice?" ×4: p9 1, p1 1, p11 1, p10 1
    • "What is the most common side effect?" ×4: p1 2, p4 1, p3 1
    • "How long can I keep the bottle after opening it?" ×4: p8 1, p11 1, p4 1, p15 1
    • "How much does it cost to print a color page?" ×4: p3 2, p1 1, p9 1
    • "How quickly will I get a response if I submit a support ticket?" ×3: p22 1, p1 1, p5 1
    • "How long can I keep the bottle once I've opened it?" ×3: p3 1, p12 1, p7 1

Most repeated main texts in train:

  • ×6: "what are the most common side effects?" (p2 2, p7 2, p1 1)
  • ×4: "who published a theory of justice?" (p9 1, p1 1, p11 1)
  • ×4: "what is the most common side effect?" (p1 2, p4 1, p3 1)
  • ×4: "how much does it cost to print a color page?" (p3 2, p1 1, p9 1)
  • ×4: "how long does my api key last before it expires?" (p6 1, p10 1, p7 1)
  • ×4: "how long can i keep the bottle after opening it?" (p8 1, p11 1, p4 1)
  • ×3: "how quickly will i get a response if i submit a support ticket?" (p22 1, p1 1, p5 1)
  • ×3: "how long can i keep the bottle once i've opened it?" (p3 1, p12 1, p7 1)

Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):

split rows examples
dev 0 (0.0%)
calibration 0 (0.0%)
test 0 (0.0%)

Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)

  • Train: 5,169 near-duplicate pairs; 10,245 rows (40.4%) sit in 5,111 clusters; largest cluster 6; excess rows (cluster size − 1) 5,134 (20.2%).
  • Clusters with more than one label class: 4,584 (9,190 rows).
    • ×6: "What are the most common side effects?" → p2 2, p7 2, p1 1, p13 1
    • ×4: "How long does my API key last before it expires?" → p6 1, p10 1, p7 1, p8 1
    • ×4: "Who published A Theory of Justice?" → p9 1, p1 1, p11 1, p10 1
    • ×4: "What is the most common side effect?" → p1 2, p4 1, p3 1
    • ×4: "How long can I keep the bottle after opening it?" → p8 1, p11 1, p4 1, p15 1
  • Held-out rows with a near duplicate in train: dev 0 (0.0%), calibration 1 (0.1%), test 1 (0.0%)
    • train "How often does the software check my license online?" ~ calibration "How often does the software check my license?" (J=0.86)
    • train "What should I do if I forget to take my daily dose?" ~ test "What should I do if I forget to take my daily capsule?" (J=0.82)
    • train "How often does the software check my license online?" ~ calibration "How often does the software check my license?" (J=0.86)

Largest train clusters:

  • ×6: "What are the most common side effects?"
  • ×4: "How long does my API key last before it expires?"
  • ×4: "Who published A Theory of Justice?"
  • ×4: "What is the most common side effect?"
  • ×4: "How long can I keep the bottle after opening it?"

5. Junk

split empty main text main text under 10 characters
train 0 0
dev 0 0
calibration 0 0
test 0 0

Pattern scan of train main texts (count, then the share of each class's rows):

pattern rows by class
placeholder [NAME]-style 0 (0.0%)
lorem ipsum 0 (0.0%)
TODO/TBD/FIXME 0 (0.0%)
'As an AI' / refusal 0 (0.0%)
chat preamble ('Here is/are...', 'Sure!') 0 (0.0%)
meta words (example/variation/message:) 0 (0.0%)
model thinking tags 0 (0.0%)
JSON/code-fence leftovers 0 (0.0%)
encoding garbage (mojibake/replacement char) 0 (0.0%)
HTML tag 0 (0.0%)
HTML entity 0 (0.0%)
base64-like run (40+ chars) 0 (0.0%)
URL 0 (0.0%)
  • Possibly cut off: 0 of 148 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: <listed option> 0.0%, none 0.0%, p1 0.0%, p10 0.0%, p11 0.0%, p12 0.0%, p13 0.0%, p14 0.0%, p15 0.0%, p16 0.0%, p17 0.0%, p18 0.0%, p2 0.0%, p20 0.0%, p3 0.0%, p4 0.0%, p5 0.0%, p6 0.0%, p7 0.0%, p8 0.0%, p9 0.0%

Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):

  • none

6. Samples

20 random train rows per kind: ground-rerank-samples.txt. Reading notes are in the findings above.