QA report: tools
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.
The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.
QA: tools
Checked 2026-09-30 13:48 by adapters/qa/qa.py (READY file READY-tools, 2026-09-30T12:42:50Z).
Verdict: PASS WITH NOTES
Tools v3 (READY-tools written 2026-09-30 12:42Z / 13:42 BST; tools/data/final-v2, train sha256 4a7935b3…3ba38, which matches READY; train 39,203 / dev 2,093 / calibration 2,195 / test 5,157). The earlier notes are in notes-tools-v1.md and notes-tools-v2.md.
Earlier findings, now fixed (checked)
- Openers are balanced across kinds: "can you" opens 3.8–4.8% of messages in every kind (it was 10.8% of tool rows against about 2% elsewhere), and "I need" opens 1.9% of tool, ask-the-user and out-of-scope messages and 0.3% of answer-directly ones (it was 15.9% of ask-the-user).
- Conversation presence is balanced: earlier turns appear in 41–45% of rows in every kind (it was 58% against 36%). The conversation's length and emptiness alone are at chance (25.1% balanced).
- Styles stay balanced: casual markers 8.7–11.3%, "and then" 5.0–6.8%, polite style rare everywhere, the thank-you sentence gone.
- No option shortcut: the correct tool is the longest 5.1% of the time against 5.0% chance, positions are uniform, and the option picker scores 5.5% against 4.8%.
- Mix and format: kinds are exactly 70/12/10/8; instructions are 70.0/26.9/3.1; the format is valid; prompts reach at most 4,527 tokens.
- Separation: families are disjoint; 0.8–1.0% of held-out rows have a near duplicate in train (generic follow-ups).
Notes (meaning; for the model card)
- A surface-only model scores 74.5% against a 70% majority baseline (42.0% balanced against 25% chance), and 35.9% on the user message alone. Its drivers are meaning:
- digits and ID tokens (IDs are in 26.5% of tool messages against 8.7% of ask-the-user ones, because ask-the-user rows lack a value by design);
- question marks, which end 49% of answer-directly messages (general and "what did you say" questions).
- The conversation's words alone give 38.7% balanced accuracy. That is meaning: in answer-directly rows, the conversation holds the answer.
- The standard phrase flags are answer-directly vocabulary: "difference between", "explain", "remind me", "did you", "thanks".
- Train shrank to 39,203 rows (from 80,330 in v1) because of the balancing.
- The noisiest kinds are still ask-the-user (some missing values could be looked up) and out-of-scope (many requests are extreme). The label audit estimated 0.4% label error on v1.
- User messages were written by the hosted qwen3.8-max model.
Automatic flags (for the reviewer to judge; not all are problems)
- formatting 'ends with ?' differs by class: answer_directly 49%, out_of_scope 27%, tool 19%, ask_user 13%
- formatting 'ends with .' differs by class: ask_user 68%, out_of_scope 57%, tool 51%, answer_directly 32%
- formatting 'no end punctuation' differs by class: tool 28%, ask_user 17%, out_of_scope 16%, answer_directly 13%
- formatting 'has a digit' differs by class: tool 62%, ask_user 30%, out_of_scope 26%, answer_directly 9%
- 22 strong phrase flags (see list): review whether they are meaning or leakage
- 42 standard phrase flags (≥2% of a class, mostly that class): review
Data checked
| split | rows | families | file |
|---|---|---|---|
| train | 39,203 | 365 | train.jsonl |
| dev | 2,093 | 20 | dev.jsonl |
| calibration | 2,195 | 16 | calibration.jsonl |
| test | 5,157 | 45 | test.jsonl |
- Train sha256:
4a7935b3f46418354b86a8912a02229a237aa804e94ad795329d4c8b61f3ba38(READY says4a7935b3f46418354b86a8912a02229a237aa804e94ad795329d4c8b61f3ba38: match) - Main text field (the text the phrase and length checks use):
state.user_message. - Label classes: answer_directly, ask_user, out_of_scope, tool (
<listed option>= one of the per-row listed options such as t3 or o12). Row kinds (source.kind): answer_directly, ask_user, out_of_scope, tool.
3. Balance
Label class share per split
| lclass | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| answer_directly | 12.0% | 12.0% | 12.0% | 12.0% | 4,704 |
| ask_user | 10.0% | 10.0% | 10.0% | 10.0% | 3,920 |
| out_of_scope | 8.0% | 8.0% | 8.0% | 8.0% | 3,136 |
| tool | 70.0% | 70.0% | 70.1% | 70.0% | 27,443 |
Row kind share per split
| kind | train | dev | calibration | test | train rows |
|---|---|---|---|---|---|
| answer_directly | 12.0% | 12.0% | 12.0% | 12.0% | 4,704 |
| ask_user | 10.0% | 10.0% | 10.0% | 10.0% | 3,920 |
| out_of_scope | 8.0% | 8.0% | 8.0% | 8.0% | 3,136 |
| tool | 70.0% | 70.0% | 70.1% | 70.0% | 27,443 |
Against the spec (train)
| value | spec | train | note |
|---|---|---|---|
| tool | 70% | 70.0% | |
| answer_directly | 12% | 12.0% | |
| ask_user | 10% | 10.0% | |
| out_of_scope | 8% | 8.0% |
4. Format
| split | row-level format problems |
|---|---|
| train | none |
| dev | none |
| calibration | none |
| test | none |
Options per choice row
| split | min | median | p99 | max |
|---|---|---|---|---|
| train | 4 | 34 | 146 | 152 |
| dev | 5 | 34 | 97 | 97 |
| calibration | 6 | 40 | 135 | 135 |
| test | 4 | 37 | 136 | 136 |
Prompt length in tokens
| split | measure | median | p99 | max | > 8192 |
|---|---|---|---|---|---|
| train | source.input_tokens (39203/39203 rows) | 1266 | 4125 | 4527 | 0 |
| dev | source.input_tokens (2093/2093 rows) | 1243 | 2356 | 2517 | 0 |
| calibration | source.input_tokens (2195/2195 rows) | 1410 | 3917 | 4020 | 0 |
| test | source.input_tokens (5157/5157 rows) | 1279 | 3546 | 3663 | 0 |
State key sets (train)
| keys | rows |
|---|---|
| agent, conversation, user_message | 39,203 (100.0%) |
Instructions (train)
- Canonical (the most common text) 70.0%, reworded 26.9% (80 distinct rewordings), none 3.1%. Target about 70 / 27 / 3.
- Canonical text: "Which tool should the agent call next to handle the user's latest message? Use the conversation for context. If no tool is needed, choose answer directly. If a tool is needed but information it requires is missing, choose ask the user."
sourceinstruction tag: canonical 70.0%, variant 26.9%, none 3.1%
| class | canonical | none |
|---|---|---|
| answer_directly | 69.9% | 3.4% |
| ask_user | 69.6% | 3.1% |
| out_of_scope | 70.5% | 3.2% |
| tool | 70.0% | 3.0% |
1. Shortcuts
Phrase statistics and models use a label-stratified sample of 39,203 train rows; models are scored on the full test file (5,157 rows).
Text length by label class (main text, characters)
| split | class | rows | p10 | median | p90 | mean |
|---|---|---|---|---|---|---|
| train | answer_directly | 4704 | 49 | 97 | 177 | 106 |
| train | ask_user | 3920 | 47 | 92 | 156 | 98 |
| train | out_of_scope | 3136 | 74 | 118 | 186 | 125 |
| train | tool | 27443 | 27 | 91 | 240 | 117 |
| test | answer_directly | 619 | 52 | 100 | 185 | 110 |
| test | ask_user | 515 | 46 | 88 | 149 | 93 |
| test | out_of_scope | 412 | 71 | 119 | 188 | 125 |
| test | tool | 3611 | 28 | 93 | 244 | 119 |
By row kind (train): main-text length, length of the rest of the state, options
| kind | rows | median chars | mean chars | median other-state chars | median options |
|---|---|---|---|---|---|
| answer_directly | 4704 | 97 | 106 | 142 | 32 |
| ask_user | 3920 | 92 | 98 | 144 | 22 |
| out_of_scope | 3136 | 118 | 125 | 143 | 28 |
| tool | 27443 | 91 | 117 | 152 | 40 |
Correct option: longest / shortest / position / key
For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):
| split | rows | correct is longest | correct is shortest | chance (1/listed) | mean relative position (0 first, 1 last; 0.5 expected) | position fifths |
|---|---|---|---|---|---|---|
| train | 27443 | 5.1% | 5.4% | 5.0% | 0.501 | 21% / 19% / 19% / 19% / 23% |
| dev | 1466 | 4.2% | 5.1% | 5.0% | 0.515 | 19% / 18% / 20% / 18% / 25% |
| calibration | 1538 | 3.8% | 3.6% | 3.6% | 0.490 | 22% / 20% / 19% / 19% / 21% |
| test | 3611 | 5.0% | 5.8% | 4.8% | 0.491 | 22% / 20% / 19% / 19% / 21% |
Correct key and position by option count
| split | options | rows | mean options | top correct keys | most common position (0-based) |
|---|---|---|---|---|---|
| train | 2-5 | 1312 | 4.6 | answer_directly 31.4%, ask_user 27.9%, t2 16.8%, t1 15.2%, t3 8.6% | 0 (31.4%) |
| train | 6-10 | 3682 | 8.3 | answer_directly 26.7%, ask_user 20.1%, t3 9.2%, t2 8.8%, t1 8.5% | 0 (26.7%) |
| train | 11-30 | 12615 | 20.7 | answer_directly 20.3%, ask_user 10.7%, t8 4.3%, t3 4.2%, t5 4.1% | 0 (20.3%) |
| train | 31-80 | 17184 | 48.5 | answer_directly 17.2%, ask_user 6.7%, t9 1.9%, t12 1.9%, t19 1.8% | 0 (17.2%) |
| train | 81+ | 4410 | 122.6 | answer_directly 21.0%, ask_user 7.2%, t48 0.9%, t75 0.8%, t12 0.8% | 0 (21.0%) |
| test | 2-5 | 162 | 4.6 | ask_user 32.7%, answer_directly 30.2%, t1 13.6%, t2 12.3%, t3 11.1% | 1 (32.7%) |
| test | 6-10 | 393 | 8.0 | answer_directly 30.0%, ask_user 18.1%, t3 12.5%, t4 9.2%, t1 8.9% | 0 (30.0%) |
| test | 11-30 | 1407 | 19.4 | answer_directly 18.7%, ask_user 12.2%, t4 5.7%, t9 5.4%, t1 5.3% | 0 (18.7%) |
| test | 31-80 | 2496 | 40.8 | answer_directly 17.9%, ask_user 6.9%, t15 2.5%, t4 2.5%, t20 2.4% | 0 (17.9%) |
| test | 81+ | 699 | 119.0 | answer_directly 22.2%, ask_user 6.7%, t36 1.4%, t23 1.4%, t60 1.4% | 0 (22.2%) |
- Train rows whose correct listed key is the first listed key (t1/o1): 4.9%, chance 5.0%.
Option count by label class (train)
| class | rows | min | median | mean | max |
|---|---|---|---|---|---|
| answer_directly | 4704 | 4 | 32 | 42.2 | 152 |
| ask_user | 3920 | 4 | 22 | 32.3 | 152 |
| out_of_scope | 3136 | 4 | 28 | 38.2 | 152 |
| tool | 27443 | 4 | 40 | 44.7 | 152 |
Source fields by label class (train)
Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.
Overall majority: 70.0%.
| source field | values | purity | top values → classes |
|---|---|---|---|
| kind | 4 | 100.0% | tool: tool 100.0%; answer_directly: answer_directly 100.0%; ask_user: ask_user 100.0%; out_of_scope: out_of_scope 100.0% |
| message_kind | 20 | 92.4% | follow_up: tool 82.3%, out_of_scope 9.7%; terse: tool 100.0%; casual: tool 78.7%, ask_user 11.1%; first_step: tool 100.0%; indirect: tool 100.0%; other: tool 100.0% |
| message_version | 2 | 75.6% | 1: tool 78.9%, answer_directly 11.2%; 2: ask_user 49.6%, out_of_scope 32.2% |
| area | 57 | 70.0% | government and city services: tool 72.3%, answer_directly 11.2%; translation and localisation: tool 72.7%, answer_directly 10.3%; job seekers looking for work: tool 73.9%, answer_directly 11.8%; database administration: tool 68.9%, answer_directly 11.1%; recruiting and hiring: tool 72.1%, answer_directly 12.4%; scientific laboratories: tool 69.3%, answer_directly 13.4% |
| size_band | 3 | 70.0% | medium: tool 71.5%, answer_directly 11.6%; large: tool 70.9%, answer_directly 13.4%; small: tool 44.3%, ask_user 25.8% |
| style | 5 | 70.0% | terse: tool 71.0%, answer_directly 11.7%; python: tool 70.5%, answer_directly 11.9%; typescript: tool 71.1%, answer_directly 12.3%; long: tool 65.8%, ask_user 13.3%; prose: tool 69.1%, answer_directly 11.8% |
| naming | 5 | 70.0% | snake_case: tool 70.3%, answer_directly 11.9%; camelCase: tool 67.6%, ask_user 12.1%; PascalCase: tool 71.8%, answer_directly 12.1%; dotted namespace: tool 69.9%, answer_directly 12.6%; kebab-case: tool 71.2%, answer_directly 12.1% |
| near_duplicate_pairs_on_list | 19 | 70.0% | 1: tool 59.6%, ask_user 15.9%; 2: tool 69.4%, answer_directly 12.2%; 5: tool 75.5%, answer_directly 11.5%; 4: tool 75.0%, answer_directly 11.1%; 3: tool 73.3%, answer_directly 10.8%; 7: tool 78.5%, answer_directly 9.8% |
| style_prompt | 5 | 70.0% | plain: tool 65.1%, answer_directly 15.6%; follow_up: tool 82.3%, out_of_scope 9.7%; multi_step: tool 77.3%, ask_user 9.7%; casual: tool 78.7%, ask_user 11.1%; polite: tool 79.4%, answer_directly 9.5% |
Formatting by label class (main text, share of rows)
| feature | answer_directly | ask_user | out_of_scope | tool | |
|---|---|---|---|---|---|
| ends with ? | 48.9% | 12.6% | 26.6% | 18.8% | gap |
| ends with . | 32.4% | 68.0% | 57.4% | 51.0% | gap |
| ends with ! | 6.1% | 2.5% | 0.1% | 0.5% | |
| no end punctuation | 12.6% | 16.5% | 15.9% | 27.9% | gap |
| starts lowercase | 27.4% | 18.0% | 19.6% | 31.7% | |
| all lowercase | 24.8% | 15.6% | 17.7% | 16.0% | |
| has a digit | 9.1% | 30.4% | 26.2% | 62.2% | gap |
| has newline | 0.0% | 0.0% | 0.0% | 0.2% | |
| has quotes | 0.0% | 0.7% | 0.0% | 2.2% | |
| has markup (HTML/markdown) | 0.0% | 0.0% | 0.0% | 0.0% | |
| has URL | 0.0% | 0.2% | 0.0% | 1.2% | |
| non-ASCII | 1.8% | 0.1% | 0.3% | 0.5% | |
| non-Latin script | 0.0% | 0.0% | 0.0% | 0.0% | |
| emoji | 0.4% | 0.0% | 0.0% | 0.0% | |
| ALL-CAPS word (4+) | 1.4% | 4.1% | 3.3% | 7.7% | |
| contains ' - ' or — | 1.6% | 0.2% | 0.1% | 0.6% |
Same, by row kind
| feature | answer_directly | ask_user | out_of_scope | tool |
|---|---|---|---|---|
| ends with ? | 49% | 13% | 27% | 19% |
| ends with . | 32% | 68% | 57% | 51% |
| ends with ! | 6% | 2% | 0% | 0% |
| no end punctuation | 13% | 17% | 16% | 28% |
| starts lowercase | 27% | 18% | 20% | 32% |
| all lowercase | 25% | 16% | 18% | 16% |
| has a digit | 9% | 30% | 26% | 62% |
| has newline | 0% | 0% | 0% | 0% |
| has quotes | 0% | 1% | 0% | 2% |
| has markup (HTML/markdown) | 0% | 0% | 0% | 0% |
| has URL | 0% | 0% | 0% | 1% |
| non-ASCII | 2% | 0% | 0% | 0% |
| non-Latin script | 0% | 0% | 0% | 0% |
| emoji | 0% | 0% | 0% | 0% |
| ALL-CAPS word (4+) | 1% | 4% | 3% | 8% |
| contains ' - ' or — | 2% | 0% | 0% | 1% |
Over-represented words and phrases per label class (main text)
Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.
Words, answer_directly: what z=54 35.6% vs 6.8%; again z=34 11.0% vs 1.2%; you z=34 41.6% vs 18.2%; thanks z=32 9.0% vs 0.3%; did z=31 9.1% vs 1.0%; explain z=30 8.2% vs 0.3%; help z=29 7.5% vs 0.6%; does z=29 7.6% vs 0.8%; or z=27 12.9% vs 3.5%; how z=27 13.7% vs 4.0%; between z=27 6.8% vs 0.8%; wait z=26 8.1% vs 1.4%; good z=26 7.3% vs 1.2%; say z=25 5.5% vs 0.5%; was z=23 10.1% vs 2.9%; remind z=23 4.9% vs 0.1%; exactly z=21 9.1% vs 2.8%; me z=21 24.1% vs 12.8%; things z=20 3.7% vs 0.4%; makes z=19 3.3% vs 0.1%
Words, ask_user: the z=16 70.3% vs 57.7%; need z=15 23.0% vs 15.1%; oh z=14 2.3% vs 0.4%; yet z=13 4.4% vs 1.5%; i z=13 42.4% vs 33.7%; haven't z=13 1.9% vs 0.3%; for z=13 49.8% vs 41.4%; forgot z=12 1.9% vs 0.4%; check z=12 10.7% vs 6.4%; decided z=11 1.1% vs 0.1%; right z=11 8.1% vs 4.6%; up z=11 14.4% vs 9.8%; asap z=11 2.8% vs 1.0%; to z=10 47.1% vs 40.9%; i'm z=10 7.3% vs 4.3%; me z=10 18.3% vs 13.7%; want z=10 7.2% vs 4.2%; idk z=10 0.9% vs 0.1%; run z=10 5.4% vs 2.9%; but z=9 7.0% vs 4.2%
Words, out_of_scope: directly z=23 6.2% vs 0.6%; our z=20 18.8% vs 6.8%; into z=18 8.8% vs 2.3%; automatically z=17 3.4% vs 0.3%; from z=16 14.1% vs 5.6%; account z=15 5.8% vs 1.4%; portal z=15 3.1% vs 0.4%; change z=13 4.7% vs 1.2%; transfer z=13 2.3% vs 0.3%; now z=12 16.7% vs 8.5%; to z=12 59.1% vs 40.0%; great z=12 4.0% vs 1.1%; submit z=12 3.4% vs 0.8%; last z=12 7.5% vs 2.9%; database z=11 2.7% vs 0.6%; password z=11 1.4% vs 0.1%; delete z=10 1.8% vs 0.3%; month z=10 3.6% vs 1.1%; their z=10 8.0% vs 3.6%; payroll z=10 1.3% vs 0.1%
Words, tool: 2024 z=21 7.7% vs 0.9%; id z=17 5.8% vs 1.3%; yes z=15 3.8% vs 0.6%; yeah z=14 5.4% vs 1.9%; let's z=14 3.7% vs 0.8%; any z=13 5.1% vs 1.9%; com z=12 2.7% vs 0.4%; 00 z=12 2.6% vs 0.5%; 01 z=11 2.2% vs 0.4%; ahead z=11 6.4% vs 3.5%; 10 z=10 2.3% vs 0.7%; show z=10 3.2% vs 1.3%; 11 z=10 1.7% vs 0.3%; currently z=10 1.8% vs 0.4%; 2023 z=10 1.6% vs 0.3%; plz z=9 1.6% vs 0.1%; 12 z=9 2.0% vs 0.6%; b z=9 1.8% vs 0.5%; all z=9 6.2% vs 3.7%; been z=9 3.0% vs 1.4%
2–4-word phrases, answer_directly: me what z=24 5.9% vs 0.6%; help me z=23 4.8% vs 0.2%; what was z=21 4.5% vs 0.1%; do i z=21 5.8% vs 1.0%; remind me z=20 4.9% vs 0.1%; got it z=20 4.2% vs 0.1%; me with z=20 3.6% vs 0.1%; if i z=19 5.2% vs 0.9%; what exactly z=19 5.0% vs 0.1%; was the z=19 4.0% vs 0.1%; tell me what z=18 3.4% vs 0.4%; you can z=18 3.1% vs 0.2%; good morning z=18 3.2% vs 0.1%; kind of z=18 2.9% vs 0.1%; and a z=17 2.8% vs 0.1%; what kind of z=17 2.8% vs 0.1%; what kind z=17 2.8% vs 0.1%; able to z=17 2.7% vs 0.2%; you explain z=17 4.1% vs 0.0%; just to z=17 2.7% vs 0.1%
2–4-word phrases, ask_user: i haven't z=13 1.6% vs 0.1%; but i z=13 3.5% vs 1.0%; for me z=12 5.6% vs 2.3%; i want z=12 6.0% vs 2.7%; i forgot z=11 1.4% vs 0.2%; hey i need to z=11 1.1% vs 0.1%; hey i z=11 1.9% vs 0.4%; i want to z=11 5.3% vs 2.4%; up the z=11 5.5% vs 2.6%; hey i need z=10 1.5% vs 0.3%; forgot to z=10 1.0% vs 0.1%; run the z=10 2.4% vs 0.8%; to check z=10 1.9% vs 0.5%; right away z=10 1.2% vs 0.2%; but i haven't z=10 1.1% vs 0.0%; and i need z=10 1.7% vs 0.4%; need to z=10 13.9% vs 9.3%; the other z=9 1.3% vs 0.2%; oh wait z=9 0.8% vs 0.0%; yo need z=9 0.8% vs 0.1%
2–4-word phrases, out_of_scope: how do i z=22 5.0% vs 0.4%; how do z=22 5.1% vs 0.4%; do i z=19 6.2% vs 1.2%; directly to z=17 2.8% vs 0.2%; to the z=16 10.9% vs 4.2%; you to z=16 5.9% vs 1.6%; great now z=16 3.1% vs 0.4%; i need you z=14 3.8% vs 0.8%; i need you to z=14 3.8% vs 0.8%; need you to z=14 4.2% vs 1.0%; need you z=14 4.2% vs 1.0%; is there a z=14 2.5% vs 0.3%; there a z=14 2.5% vs 0.3%; is there a way z=14 2.0% vs 0.2%; there a way z=14 2.0% vs 0.2%; to our z=13 2.6% vs 0.4%; a way z=13 2.1% vs 0.2%; there a way to z=13 1.8% vs 0.1%; way to z=13 2.1% vs 0.3%; a way to z=13 1.8% vs 0.2%
2–4-word phrases, tool: so i z=13 10.0% vs 5.8%; so i can z=12 7.3% vs 4.0%; go ahead z=11 6.3% vs 3.3%; i can z=11 8.0% vs 4.7%; go ahead and z=11 6.1% vs 3.3%; ahead and z=11 6.1% vs 3.3%; id is z=10 1.8% vs 0.3%; do the z=10 1.8% vs 0.3%; show me z=9 2.5% vs 1.0%; hey can u z=8 1.2% vs 0.2%; need to see z=8 1.6% vs 0.5%; 2024 11 z=8 1.1% vs 0.1%; second one z=8 1.1% vs 0.2%; you pull z=8 1.5% vs 0.4%; do the same z=8 1.1% vs 0.2%; hey can z=8 1.3% vs 0.3%; just need z=8 1.2% vs 0.2%; you pull up z=8 1.2% vs 0.3%; make sure z=8 2.1% vs 1.0%; pull up z=8 3.9% vs 2.3%
Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):
- answer_directly:
what35.6% vs 6.8% - answer_directly:
again11.0% vs 1.2% - answer_directly:
did9.1% vs 1.0% - answer_directly:
thanks9.0% vs 0.3% - answer_directly:
explain8.2% vs 0.3% - answer_directly:
wait8.1% vs 1.4% - tool:
20247.7% vs 0.9% - answer_directly:
does7.6% vs 0.8% - answer_directly:
help7.5% vs 0.6% - answer_directly:
good7.3% vs 1.2% - answer_directly:
between6.8% vs 0.8% - out_of_scope:
directly6.2% vs 0.6% - out_of_scope:
do i6.2% vs 1.2% - answer_directly:
me what5.9% vs 0.6% - tool:
id5.8% vs 1.3% - answer_directly:
do i5.8% vs 1.0% - out_of_scope:
account5.8% vs 1.4% - answer_directly:
say5.5% vs 0.5% - answer_directly:
if i5.2% vs 0.9% - out_of_scope:
how do5.1% vs 0.4% - answer_directly:
what exactly5.0% vs 0.1% - out_of_scope:
how do i5.0% vs 0.4%
Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):
- all rows (39,203 rows): answer_directly:
thanks9.0% of class, 79% of its 536 rows; answer_directly:explain8.2% of class, 80% of its 484 rows; answer_directly:difference5.7% of class, 95% of its 281 rows; answer_directly:did you5.4% of class, 98% of its 260 rows; answer_directly:difference between5.1% of class, 100% of its 243 rows; answer_directly:what exactly5.0% of class, 92% of its 257 rows; answer_directly:remind me4.9% of class, 89% of its 258 rows; answer_directly:help me4.8% of class, 76% of its 298 rows; answer_directly:what was4.5% of class, 83% of its 256 rows; answer_directly:got it4.2% of class, 85% of its 232 rows; answer_directly:the difference4.1% of class, 95% of its 203 rows; answer_directly:thanks for4.0% of class, 95% of its 198 rows; answer_directly:was the4.0% of class, 88% of its 213 rows; answer_directly:can you explain3.8% of class, 94% of its 191 rows; answer_directly:the difference between3.7% of class, 99% of its 174 rows; answer_directly:me with3.6% of class, 78% of its 218 rows; answer_directly:are you3.5% of class, 91% of its 182 rows; answer_directly:did you say3.5% of class, 100% of its 163 rows; answer_directly:help me with3.3% of class, 98% of its 158 rows; answer_directly:makes3.3% of class, 77% of its 200 rows; answer_directly:sense3.2% of class, 78% of its 194 rows; answer_directly:good morning3.2% of class, 80% of its 188 rows; answer_directly:between a3.1% of class, 99% of its 149 rows; answer_directly:mean3.1% of class, 78% of its 190 rows; answer_directly:you can3.1% of class, 71% of its 203 rows; answer_directly:what was the3.1% of class, 95% of its 151 rows; answer_directly:question2.9% of class, 79% of its 175 rows; answer_directly:kind of2.9% of class, 74% of its 183 rows; answer_directly:what kind of2.8% of class, 81% of its 163 rows; answer_directly:and a2.8% of class, 76% of its 173 rows; answer_directly:just to2.7% of class, 80% of its 161 rows; answer_directly:or do2.7% of class, 93% of its 139 rows; answer_directly:you're2.7% of class, 73% of its 173 rows; answer_directly:difference between a2.5% of class, 100% of its 119 rows; answer_directly:that makes2.5% of class, 92% of its 127 rows; answer_directly:explain how2.4% of class, 94% of its 121 rows; answer_directly:what does2.4% of class, 75% of its 150 rows; answer_directly:and then explain2.3% of class, 97% of its 110 rows; answer_directly:quick question2.2% of class, 98% of its 106 rows; answer_directly:explain the2.2% of class, 78% of its 131 rows
Shortcut models
Predicting the label class on test (5,157 rows). Chance 25.0%, majority class ('tool') 70.0%; balanced chance 25.0%.
| model (logistic regression, trained on the train sample) | test accuracy | balanced accuracy (mean recall) |
|---|---|---|
| bag of words, whole state (words and word pairs) | 84.2% | 60.5% |
bag of words, main text only (user_message) |
85.4% | 63.6% |
| surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) | 74.5% | 42.0% |
| surface features of the main text only | 72.7% | 35.9% |
Strongest single surface features (logistic regression on one feature, balanced accuracy on test):
| feature | accuracy | balanced accuracy |
|---|---|---|
| count_! | 70.7% | 27.4% |
| ends_! | 70.4% | 26.4% |
| chars(log) | 70.0% | 25.0% |
| words(log) | 70.0% | 25.0% |
| upper_ratio | 70.0% | 25.0% |
| digit_ratio | 70.0% | 25.0% |
| nonascii_ratio | 70.0% | 25.0% |
| nonlatin | 70.0% | 25.0% |
| emoji | 70.0% | 25.0% |
| newlines | 70.0% | 25.0% |
Other state fields alone (predicting the label class on test from one field, without the main text):
| field | treated as | accuracy | balanced accuracy |
|---|---|---|---|
| agent | text: bag of words / length+empty | 70.0% / 70.0% | 25.0% / 25.0% |
| conversation | text: bag of words / length+empty | 74.6% / 70.1% | 38.7% / 25.1% |
No-meaning option picker: a logistic ranker scores each option from its position, length, key type, fixed-option identity and shape (commas, brackets, capitals), never reading the state or the option's words, and picks the top option per row.
- Test accuracy 20.0% against uniform chance 4.5% (this includes the fixed options, whose share is a class prior).
- Among rows whose answer is a listed option (3,611), picking only among listed options: 5.5% against chance 4.8%.
2. Duplicates and split separation
Families shared between splits
| splits | shared families | examples |
|---|---|---|
| train ∩ dev | 0 | |
| train ∩ calibration | 0 | |
| train ∩ test | 0 | |
| dev ∩ calibration | 0 | |
| dev ∩ test | 0 | |
| calibration ∩ test | 0 |
- Train rows whose main text repeats an earlier row's (normalised): 153 (0.4%); groups: 79; largest group 10.
- Train rows identical in the whole prompt (state, options, instructions): 0.
- Identical whole prompt, different answer: 0 groups (0 rows).
- Identical main text, different label class: 4 groups (13 rows). (Can be legitimate when the rest of the state or the options differ.)
- "What about the other one?" ×6: ask_user 5, tool 1
- "Now do the same for the second one." ×3: tool 2, ask_user 1
- "Go ahead and book it." ×2: tool 1, ask_user 1
- "Go ahead and run that search now." ×2: ask_user 1, tool 1
Most repeated main texts in train:
- ×10: "yes, do it." (tool 10)
- ×8: "yes, go ahead and book it." (tool 8)
- ×8: "yeah do it" (tool 8)
- ×7: "yes do it" (tool 7)
- ×6: "yes, go ahead and send it." (tool 6)
- ×6: "yes, go ahead and run it." (tool 6)
- ×6: "what about the other one?" (ask_user 5, tool 1)
- ×6: "got it thanks bye" (answer_directly 6)
Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):
| split | rows | examples |
|---|---|---|
| dev | 15 (0.7%) | "what about the other one?"; "Yeah go ahead with it."; "Yeah, do it."; "parking rules" |
| calibration | 17 (0.8%) | "Yes, do it for that one."; "Yes, do it."; "got it, thx"; "Yes, go ahead and run it." |
| test | 22 (0.4%) | "Yes, go ahead and generate it."; "Yes, go ahead and pull that up."; "Yes do it."; "Yes, go ahead and run it." |
Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)
- Train: 695 near-duplicate pairs; 322 rows (0.8%) sit in 96 clusters; largest cluster 23; excess rows (cluster size − 1) 226 (0.6%).
- Clusters with more than one label class: 4 (36 rows).
- ×23: "Yes, do the same for the second one." → tool 22, ask_user 1
- ×6: "What about the other one?" → ask_user 5, tool 1
- ×5: "Okay, go ahead and run that search now." → tool 4, ask_user 1
- ×2: "Go ahead and book it." → tool 1, ask_user 1
- Held-out rows with a near duplicate in train: dev 16 (0.8%), calibration 23 (1.0%), test 42 (0.8%)
- train "yes do it" ~ calibration "Yes do it" (J=1.00)
- train "hey, you there?" ~ calibration "hey, you there?" (J=1.00)
- train "Yep, do it." ~ calibration "Yep do it" (J=1.00)
- train "Yes, go ahead and send it." ~ test "Yes, go ahead and send it." (J=1.00)
- train "yeah do it" ~ dev "Yeah, do it." (J=1.00)
Largest train clusters:
- ×23: "Yes, do the same for the second one."
- ×18: "yes do it"
- ×13: "Yeah, do it."
- ×10: "Yes, go ahead and run it now."
- ×10: "Yes, go ahead and send it."
5. Junk
| split | empty main text | main text under 10 characters |
|---|---|---|
| train | 0 | 84 |
| dev | 0 | 6 |
| calibration | 0 | 9 |
| test | 0 | 6 |
Very short train examples: "my holds" (tool); "yes do it" (tool); "yes do it" (tool); "hi there" (answer_directly); "API_Call" (tool); "Use 2021." (tool); "ok do it" (tool); "active rx" (tool); "FAQ docs" (tool); "yes do it" (tool); "code 504" (tool); "fireworks" (tool)
Pattern scan of train main texts (count, then the share of each class's rows):
| pattern | rows | by class |
|---|---|---|
| placeholder [NAME]-style | 0 (0.0%) | |
| lorem ipsum | 0 (0.0%) | |
| TODO/TBD/FIXME | 0 (0.0%) | |
| 'As an AI' / refusal | 0 (0.0%) | |
| chat preamble ('Here is/are...', 'Sure!') | 127 (0.3%) | ask_user 0.2%, out_of_scope 0.0%, tool 0.4% |
| meta words (example/variation/message:) | 0 (0.0%) | |
| model thinking tags | 0 (0.0%) | |
| JSON/code-fence leftovers | 2 (0.0%) | tool 0.0% |
| encoding garbage (mojibake/replacement char) | 0 (0.0%) | |
| HTML tag | 6 (0.0%) | tool 0.0% |
| HTML entity | 0 (0.0%) | |
| base64-like run (40+ chars) | 23 (0.1%) | tool 0.1% |
| URL | 343 (0.9%) | ask_user 0.2%, tool 1.2% |
chat preamble ('Here is/are...', 'Sure!'):
tools-equipment_maintenance_bot-119424(tool): Here is the first one for the report: QTR-990. |tools-permit_faster-078588(tool): Sure, I live in 33139. I just want to make sure I don't get fined by the city for putting it too close to the property line or making it to… |tools-b2b_procurement_portal-093260(tool): Here is the information you asked for: PO ID is PO-60199 and the URL is https://sharepoint.corp.com/fastenal/corrected_v2.pdf. Let's get th…JSON/code-fence leftovers:
tools-crypto_quant_analyst-124091(tool): Here is the JSON payload from the committee's decision engine to apply to our trading book: ⏎ { ⏎ "action": "rebalance", ⏎ "weights": {… |tools-clinic_frontdesk_bot-033088(tool): Here is the payload he gave me: ⏎ [ ⏎ {"first": "Liam", "last": "Neeson", "email": "liam.n@test.com"}, ⏎ {"first": "Morgan", "last": "F…HTML tag:
tools-edu_tutor_assistant-001882(tool): i got this error Traceback (most recent call last): ⏎ File "hw3.py", line 8, in <module> ⏎ result = divide(a, b) ⏎ File "hw3.py", l… |tools-edu_tutor_assistant-001881(tool): im working on the lab 2 assignment where we have to read a csv file and calculate averages. i keep running into this crash every time i try… |tools-edu_tutor_assistant-001885(tool): Hello, I hope you are doing well. I am terribly sorry to bother you, but I was wondering if you would be so kind as to help me understand w…base64-like run (40+ chars):
tools-soc_tier1_triage-120728(tool): Can you check if the SHA256 hash 8d7b3f9a1c2e4f6a8b0c2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a is known malware or benign? |tools-legal_contract_translator-138445(tool): Could you upload the executed master services agreement for matter MAT-2024-0891 to the secure vault? The base64 encoded file content is JV… |tools-tournament_ticker-047386(tool): First, upload the logo file winter_brawl_logo.svg as an image/svg+xml with the binary data PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9z…URL:
tools-global_expatriate_tax-128022(tool): Here it is: https://japan-tax.jp/rent_march.pdf |tools-ecommerce_product_sync-140894(tool): Great. Now do the same validation for the third image in that set, which is https://media.leathercraft.co/bags/model-X10/detail-stitching.j… |tools-small_business_press_kit-091194(tool): Could you log the new article mention that just went live at https://www.dailyherald.com/news/2024/local-bakery-expansion and calculate the…Possibly cut off: 13 of 1,709 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: answer_directly 0.0%, ask_user 0.0%, out_of_scope 0.0%, tool 0.8%
tools-legal_case_manager-068017: …act filename. could u search the files for that case using the keyword timeline so i can figure out which one it is? thxtools-executive_chief_of_staff-029287: … to archive everything coming from marketing-spam.net so he can actually see what matters before his 2pm call, pls hurrytools-ip_patent_prior_art-041295: … from our application, can u run an image match on it? https://storage.myfirm.internal/cases/4421/hinge-drawing-fig3.png
Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):
- none
6. Samples
20 random train rows per kind: tools-samples.txt. Reading notes are in the findings above.
