Skip to content
JeffHub

QA report: nav

The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.

The third independent review of this adapter's training, development, calibration and test files, written before training by a reviewer who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.

QA: nav (third independent review, rebuild of 2026-10-01 00:48)

Verdict: PASS WITH NOTES

Reviewed 2026-10-01 01:40 BST by an independent reviewer. The hand-over is READY-nav.staging. The train file is train.jsonl, 72,299 rows. I recomputed its SHA-256 as e1ef9066bd088796e1ebeff86164b65dec3941127cf19437fe2fe44297ae6d65, which matches the staging file. The dev, calibration and test checksums also match. The previous review (FIX FIRST, 23:45) is kept as nav-review2-fix-first.md, and the one before it as nav-before-length-fix.md.

Deciding reasons:

  1. The length signal is gone. Every kind except noise has the same word-count histogram, matching to 0.1 points at every word count from 2 to 16+. A length-only model scores 35.9% balanced on the 3 answers (chance 33.3%), 19.9% on the 5 kinds without noise (chance 20%), and 50.4% on navigation against "none of these" (chance 50%). The previous review's script gave 47.3% before this fix.
  2. The "none of these" share no longer depends on style or round. In train it is 14.7–15.0% in every generation round (original rows, round 2, round 4, round 5) and 14.7% at every length band. Round 4–5 requests were written alike and then made navigation or "none" at random. Words alone separate "none" from navigation within rounds 4–5 at 56.9% balanced, about the same as for the whole set (56.5%). That leftover is content, because the requested item is absent. No new shortcut from list shape or from per-cell sampling generalises to held-out apps (details below).
  3. Hand reading found 0 clear label errors in 78 rows. There were 3 borderline rows and 2 unnatural ones.

Scripts and outputs are all read-only and saved in qa/review3/:

  • a.out and b.out: the previous reviewer's nav_lenfix.py and nav_lenfix2.py, unchanged (they tag rounds 1–3 only).
  • c.out and d.out: the builder's nav_lenfix_r45.py and nav_lenfix2_r45.py. I diffed them against the originals. The only change is the round list (adding extra4 and extra5) and the file name passed to exec.
  • rev3.py, rev3b.py, rev3c.py, rev3d.py: my own independent checks, with outputs in the matching .out files. rev3c.out holds the rows I read by hand.

Builder's claims against my recomputation

claim builder mine agree?
length-only model, 3 answers 35.9% (34.2% without noise) reviewer's script, rerun: 35.9% / 34.2%. My own model (adds characters per word): 35.4% / 35.1% on test, 36.1% / 33.3% on a held-back half of train agree
length-only model, 5 kinds without noise 19.9% vs 20% 19.9% (my model 21.6%) agree
share of each kind at ≤5 words / ≥10 words (train) 32% / 21% 32.1% / 20.8% for every kind except noise (test 32% / 28–29%) agree
"none" share 14.7% at every length band and in every round yes 14.7% at every band; by round: original 14.7, round 2 15.0, round 4 14.8, round 5 14.8 (test 14.5 / 14.2 / 16.7 / 13.3, small counts) agree
questions 10.2% at every length band yes 10.2% in every band (test 10.0–10.4%) agree
vague pointers 0% 0% 0% in clean, garbled, none, question and unrelated rows (noise 1%) agree
articles: questions of 4+ words 99% new vs 64% old; navigation 53–85% yes determiner (articles and possessives): new questions 100% in every round, original questions 66%. Navigation 53–95% depending on round. Article (the/a/an) only: new questions 71–89%, original 57% agree; see note 1
fillers: unrelated 13%, questions 14%, navigation and none 21–24% yes clean 23%, garbled 21%, none 24%, question 14%, unrelated 13% agree
correct option longest / shortest 3.6% / 4.9% vs 3.4% 3.5% / 4.9% vs 3.4% chance; first 3.3%, last 3.5%, mean relative position 0.500; rounds 4 and 5 longest 3.8% / 4.0% vs 3.4–3.5% agree
meaning-free option picker at chance 4.9% vs 5.0% (qa.py) not rerun not checked
surface-feature model 18.7% vs 16.7% (6 kinds) 18.7% not comparable. With length, articles, fillers, punctuation, "please", i/me/you and a question-word flag, I get 51.1% on 6 kinds; without the question-word flag, 37.5% on 5 kinds. Navigation against none without that flag: 53.1%. The signal sits in question against request form (question words, "please", "could you", "me"), which is what the class means partly; see note 3
size 72,299 / 1,996 / 1,994 / 3,300 same (line counts) agree
split held out by app yes train 286 apps, test 34, overlap 0; dev and calibration share no app with train agree
train SHA-256 e1ef9066… e1ef9066bd088796e1ebeff86164b65dec3941127cf19437fe2fe44297ae6d65 agree
leak gate passed yes not rerun. Note: 144 of 3,300 test transcripts also occur in train as exact text (generic ones such as "open settings" and "help me out", on other apps' screens) not rerun

What I checked for new shortcuts

  • Round-4 "none" rows: does the shortened list give away the removed item?
    • Across apps, no. A model given only the list's shape (option count, how many items of each entity type, spread of those group sizes, from IDs or from the visible first word of each option) scores 52.1% and 51.0% on held-out apps against 50% chance. Group sizes are uneven in 81–85% of rows of every kind. Option keys have no numbering gap in any kind: the keys run o1…oN (or item_1…, or other key names) without holes.
    • Within one screen, yes, but only by memorising that screen. A screen's "none" rows have one option fewer than its other rows: 98% of rows that are exactly one option short of their screen's usual count are "none", and these cover about 75% of "none" rows. This holds in the original rows just as in rounds 4–5, so it is a property of the whole design, not new. A model could only use it by memorising each screen's item count. Test apps are never seen in training, so a model that relied on it would fail on test rather than look good. See note 2.
    • Small screens use a stand-in item (1,132 train rows), so the option count is unchanged there.
  • Per-cell sampling and app concentration. Predicting the kind from the app, from the app plus exact word count, from the screen, or from the screen plus word count scores 20.0–22.3% balanced on 5 kinds (chance 20%). Within any kind-by-word-count cell of 100 or more rows, no app holds more than 5%. Screen type, format (canonical against varied, 69–70% / 30–31%) and instruction wording are spread evenly across kinds.
  • Articles over-correction (new questions at 100% determiners). Among non-noise rows of 4–5 words, the question share is 22.8% with an article against 4.0% without (test: 27% against 2.6%). Almost all of that comes from question words: among rows that start with a question word, an article raises the question share from 67.7% to 93.4% (483 and 982 rows). Among rows that do not start with one, the question share is 0.1–0.5% either way. Over all rows of 4+ words that start with a question word, the effect almost disappears (54.8% with an article against 51.7% without). This is a small, local cue, not a general shortcut. At 2–3 words, questions almost never have an article (0–6%), and neither do commands.
  • Round style markers. Rounds 4–5 requests often end in "now" (11–17% against about 1% in original requests), and contain "for me" (7–12%) or "come and now"/"command" (5–9%). These rates are equal for navigation and "none" rows in the same round, so they carry no navigation-against-none label. They never occur in questions, but neither do they in original questions, and they are request phrasing.
  • Label mix by round (train): original 54,793 rows, round 1 478, round 2 3,518, round 3 488, round 4 7,412, round 5 5,610. Round 4's question batch (Lk) also gave 109 short navigation rows, and those are correctly labelled.
  • Repeats. Transcripts with more than one kind: 1,421, almost all the same generic command appearing as clean on one screen and "none" on another (for example "help me out"). That is correct by design. The most repeated question is "is the dashboard loading" (9 times).

Hand reading (78 rows; rounds 4–5 plus 4 original questions; review3/rev3c.out)

I read 18 round-4/5 "none" rows, 16 round-4/5 navigation rows (clean and garbled), 16 round-4/5 questions, 7 short navigation rows from the question batches, 8 unrelated remarks and 4 original questions. For "none" rows I checked every listed option that shares a word with the request.

  • Clear label errors: 0.
  • Borderline: 3.
    • "show helen and dale an a ver sary" is labelled "none". The anniversary photo was removed, but "photo Wedding of Helen and Dale 1948" is still listed. It is a different photo, but close.
    • "what is comedy" is labelled "question", and "Comedy" is listed. It reads as a question, so the label is defensible.
    • "uh clothes the gate" (an airline app) and "tape that" are labelled unrelated. Either could be a command, but unrelated and "none of these" give the same answer.
  • Unnatural: 2. "is the my hearth updated" (the screen is "My Hearth"; 17 questions contain "the my …") and "is welcome loading". Questions in rounds 4–5 lean on the template "is the … updated/loading/ready" (13–14% of their questions, against 0.7% of original questions). This is repetitive but stays within the question class.
  • The removed targets in "none" rows were absent in every case I checked, including close same-type distractors (PF-1004 absent with PF-1001/1003/1006/1007 listed; a cloud-sync project absent with other projects listed).

Notes (not blocking)

  1. Determiners in new questions are 100%, against 66% in original questions and 53–95% in requests. At 4–5 words this adds a small cue on top of question words (see above). If there is another round, let some generated questions of 4+ words go without an article ("who signed off on budget", "any word on refund") instead of dropping all of them.
  2. "None of these" rows are always a screen with one item taken out. Only per-screen memory can exploit this, so it does not transfer to held-out apps, and it is not new. A cheap future hardening: in some navigation rows, also remove one random item that is not the target, so that "one item missing" stops meaning "none".
  3. Request politeness never appears in questions. "please", "could you", "for me", "go ahead" and "so i can" appear in 0.0% of questions, against 1–14% of requests. Real users do say "could you tell me …". Adding a few polite questions in a future round would make the question class less tied to form.
  4. Rounds 4–5 questions over-use "is the … updated/loading". Train shrank to 72,299 rows because some cells were short (long clean navigation at 11–15 words, questions at 2 and 4 words, none at 9–10 words). That is acceptable for an adapter.
  5. Test was built the same way as train. The checks above that use held-out apps are the meaningful ones; test scores will not reveal per-screen memorisation, though that cannot help on unseen apps either.