Skip to content
JeffHub

QA report: tools

The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged. Sample rows are shown as plain text.

The data-quality report on this adapter's training, development, calibration and test files, written by the maintainers' QA script before training and reviewed by someone who did not build the data. For publication, internal file paths were cut to file names and machine, service and account names were removed; every number, verdict and sample row is unchanged.

QA: tools

Checked 2026-09-30 13:48 by adapters/qa/qa.py (READY file READY-tools, 2026-09-30T12:42:50Z).

Verdict: PASS WITH NOTES

Tools v3 (READY-tools written 2026-09-30 12:42Z / 13:42 BST; tools/data/final-v2, train sha256 4a7935b3…3ba38, which matches READY; train 39,203 / dev 2,093 / calibration 2,195 / test 5,157). The earlier notes are in notes-tools-v1.md and notes-tools-v2.md.

Earlier findings, now fixed (checked)

  • Openers are balanced across kinds: "can you" opens 3.8–4.8% of messages in every kind (it was 10.8% of tool rows against about 2% elsewhere), and "I need" opens 1.9% of tool, ask-the-user and out-of-scope messages and 0.3% of answer-directly ones (it was 15.9% of ask-the-user).
  • Conversation presence is balanced: earlier turns appear in 41–45% of rows in every kind (it was 58% against 36%). The conversation's length and emptiness alone are at chance (25.1% balanced).
  • Styles stay balanced: casual markers 8.7–11.3%, "and then" 5.0–6.8%, polite style rare everywhere, the thank-you sentence gone.
  • No option shortcut: the correct tool is the longest 5.1% of the time against 5.0% chance, positions are uniform, and the option picker scores 5.5% against 4.8%.
  • Mix and format: kinds are exactly 70/12/10/8; instructions are 70.0/26.9/3.1; the format is valid; prompts reach at most 4,527 tokens.
  • Separation: families are disjoint; 0.8–1.0% of held-out rows have a near duplicate in train (generic follow-ups).

Notes (meaning; for the model card)

  1. A surface-only model scores 74.5% against a 70% majority baseline (42.0% balanced against 25% chance), and 35.9% on the user message alone. Its drivers are meaning:
    • digits and ID tokens (IDs are in 26.5% of tool messages against 8.7% of ask-the-user ones, because ask-the-user rows lack a value by design);
    • question marks, which end 49% of answer-directly messages (general and "what did you say" questions).
  2. The conversation's words alone give 38.7% balanced accuracy. That is meaning: in answer-directly rows, the conversation holds the answer.
  3. The standard phrase flags are answer-directly vocabulary: "difference between", "explain", "remind me", "did you", "thanks".
  4. Train shrank to 39,203 rows (from 80,330 in v1) because of the balancing.
  5. The noisiest kinds are still ask-the-user (some missing values could be looked up) and out-of-scope (many requests are extreme). The label audit estimated 0.4% label error on v1.
  6. User messages were written by the hosted qwen3.8-max model.

Automatic flags (for the reviewer to judge; not all are problems)

  • formatting 'ends with ?' differs by class: answer_directly 49%, out_of_scope 27%, tool 19%, ask_user 13%
  • formatting 'ends with .' differs by class: ask_user 68%, out_of_scope 57%, tool 51%, answer_directly 32%
  • formatting 'no end punctuation' differs by class: tool 28%, ask_user 17%, out_of_scope 16%, answer_directly 13%
  • formatting 'has a digit' differs by class: tool 62%, ask_user 30%, out_of_scope 26%, answer_directly 9%
  • 22 strong phrase flags (see list): review whether they are meaning or leakage
  • 42 standard phrase flags (≥2% of a class, mostly that class): review

Data checked

split rows families file
train 39,203 365 train.jsonl
dev 2,093 20 dev.jsonl
calibration 2,195 16 calibration.jsonl
test 5,157 45 test.jsonl
  • Train sha256: 4a7935b3f46418354b86a8912a02229a237aa804e94ad795329d4c8b61f3ba38 (READY says 4a7935b3f46418354b86a8912a02229a237aa804e94ad795329d4c8b61f3ba38: match)
  • Main text field (the text the phrase and length checks use): state.user_message.
  • Label classes: answer_directly, ask_user, out_of_scope, tool (<listed option> = one of the per-row listed options such as t3 or o12). Row kinds (source.kind): answer_directly, ask_user, out_of_scope, tool.

3. Balance

Label class share per split

lclass train dev calibration test train rows
answer_directly 12.0% 12.0% 12.0% 12.0% 4,704
ask_user 10.0% 10.0% 10.0% 10.0% 3,920
out_of_scope 8.0% 8.0% 8.0% 8.0% 3,136
tool 70.0% 70.0% 70.1% 70.0% 27,443

Row kind share per split

kind train dev calibration test train rows
answer_directly 12.0% 12.0% 12.0% 12.0% 4,704
ask_user 10.0% 10.0% 10.0% 10.0% 3,920
out_of_scope 8.0% 8.0% 8.0% 8.0% 3,136
tool 70.0% 70.0% 70.1% 70.0% 27,443

Against the spec (train)

value spec train note
tool 70% 70.0%
answer_directly 12% 12.0%
ask_user 10% 10.0%
out_of_scope 8% 8.0%

4. Format

split row-level format problems
train none
dev none
calibration none
test none

Options per choice row

split min median p99 max
train 4 34 146 152
dev 5 34 97 97
calibration 6 40 135 135
test 4 37 136 136

Prompt length in tokens

split measure median p99 max > 8192
train source.input_tokens (39203/39203 rows) 1266 4125 4527 0
dev source.input_tokens (2093/2093 rows) 1243 2356 2517 0
calibration source.input_tokens (2195/2195 rows) 1410 3917 4020 0
test source.input_tokens (5157/5157 rows) 1279 3546 3663 0

State key sets (train)

keys rows
agent, conversation, user_message 39,203 (100.0%)

Instructions (train)

  • Canonical (the most common text) 70.0%, reworded 26.9% (80 distinct rewordings), none 3.1%. Target about 70 / 27 / 3.
  • Canonical text: "Which tool should the agent call next to handle the user's latest message? Use the conversation for context. If no tool is needed, choose answer directly. If a tool is needed but information it requires is missing, choose ask the user."
  • source instruction tag: canonical 70.0%, variant 26.9%, none 3.1%
class canonical none
answer_directly 69.9% 3.4%
ask_user 69.6% 3.1%
out_of_scope 70.5% 3.2%
tool 70.0% 3.0%

1. Shortcuts

Phrase statistics and models use a label-stratified sample of 39,203 train rows; models are scored on the full test file (5,157 rows).

Text length by label class (main text, characters)

split class rows p10 median p90 mean
train answer_directly 4704 49 97 177 106
train ask_user 3920 47 92 156 98
train out_of_scope 3136 74 118 186 125
train tool 27443 27 91 240 117
test answer_directly 619 52 100 185 110
test ask_user 515 46 88 149 93
test out_of_scope 412 71 119 188 125
test tool 3611 28 93 244 119

By row kind (train): main-text length, length of the rest of the state, options

kind rows median chars mean chars median other-state chars median options
answer_directly 4704 97 106 142 32
ask_user 3920 92 98 144 22
out_of_scope 3136 118 125 143 28
tool 27443 91 117 152 40

Correct option: longest / shortest / position / key

For rows whose answer is one of the listed options (fixed options such as 'none of these' excluded):

split rows correct is longest correct is shortest chance (1/listed) mean relative position (0 first, 1 last; 0.5 expected) position fifths
train 27443 5.1% 5.4% 5.0% 0.501 21% / 19% / 19% / 19% / 23%
dev 1466 4.2% 5.1% 5.0% 0.515 19% / 18% / 20% / 18% / 25%
calibration 1538 3.8% 3.6% 3.6% 0.490 22% / 20% / 19% / 19% / 21%
test 3611 5.0% 5.8% 4.8% 0.491 22% / 20% / 19% / 19% / 21%

Correct key and position by option count

split options rows mean options top correct keys most common position (0-based)
train 2-5 1312 4.6 answer_directly 31.4%, ask_user 27.9%, t2 16.8%, t1 15.2%, t3 8.6% 0 (31.4%)
train 6-10 3682 8.3 answer_directly 26.7%, ask_user 20.1%, t3 9.2%, t2 8.8%, t1 8.5% 0 (26.7%)
train 11-30 12615 20.7 answer_directly 20.3%, ask_user 10.7%, t8 4.3%, t3 4.2%, t5 4.1% 0 (20.3%)
train 31-80 17184 48.5 answer_directly 17.2%, ask_user 6.7%, t9 1.9%, t12 1.9%, t19 1.8% 0 (17.2%)
train 81+ 4410 122.6 answer_directly 21.0%, ask_user 7.2%, t48 0.9%, t75 0.8%, t12 0.8% 0 (21.0%)
test 2-5 162 4.6 ask_user 32.7%, answer_directly 30.2%, t1 13.6%, t2 12.3%, t3 11.1% 1 (32.7%)
test 6-10 393 8.0 answer_directly 30.0%, ask_user 18.1%, t3 12.5%, t4 9.2%, t1 8.9% 0 (30.0%)
test 11-30 1407 19.4 answer_directly 18.7%, ask_user 12.2%, t4 5.7%, t9 5.4%, t1 5.3% 0 (18.7%)
test 31-80 2496 40.8 answer_directly 17.9%, ask_user 6.9%, t15 2.5%, t4 2.5%, t20 2.4% 0 (17.9%)
test 81+ 699 119.0 answer_directly 22.2%, ask_user 6.7%, t36 1.4%, t23 1.4%, t60 1.4% 0 (22.2%)
  • Train rows whose correct listed key is the first listed key (t1/o1): 4.9%, chance 5.0%.

Option count by label class (train)

class rows min median mean max
answer_directly 4704 4 32 42.2 152
ask_user 3920 4 22 32.3 152
out_of_scope 3136 4 28 38.2 152
tool 27443 4 40 44.7 152

Source fields by label class (train)

Scalar source fields with 2–60 values. 'Purity' = accuracy of predicting the label class from this field alone (per-value majority), against the overall majority. The model does not see source, but a field that predicts the label marks a confound: rows of one origin carry one label, so any style difference of that origin becomes a shortcut.

Overall majority: 70.0%.

source field values purity top values → classes
kind 4 100.0% tool: tool 100.0%; answer_directly: answer_directly 100.0%; ask_user: ask_user 100.0%; out_of_scope: out_of_scope 100.0%
message_kind 20 92.4% follow_up: tool 82.3%, out_of_scope 9.7%; terse: tool 100.0%; casual: tool 78.7%, ask_user 11.1%; first_step: tool 100.0%; indirect: tool 100.0%; other: tool 100.0%
message_version 2 75.6% 1: tool 78.9%, answer_directly 11.2%; 2: ask_user 49.6%, out_of_scope 32.2%
area 57 70.0% government and city services: tool 72.3%, answer_directly 11.2%; translation and localisation: tool 72.7%, answer_directly 10.3%; job seekers looking for work: tool 73.9%, answer_directly 11.8%; database administration: tool 68.9%, answer_directly 11.1%; recruiting and hiring: tool 72.1%, answer_directly 12.4%; scientific laboratories: tool 69.3%, answer_directly 13.4%
size_band 3 70.0% medium: tool 71.5%, answer_directly 11.6%; large: tool 70.9%, answer_directly 13.4%; small: tool 44.3%, ask_user 25.8%
style 5 70.0% terse: tool 71.0%, answer_directly 11.7%; python: tool 70.5%, answer_directly 11.9%; typescript: tool 71.1%, answer_directly 12.3%; long: tool 65.8%, ask_user 13.3%; prose: tool 69.1%, answer_directly 11.8%
naming 5 70.0% snake_case: tool 70.3%, answer_directly 11.9%; camelCase: tool 67.6%, ask_user 12.1%; PascalCase: tool 71.8%, answer_directly 12.1%; dotted namespace: tool 69.9%, answer_directly 12.6%; kebab-case: tool 71.2%, answer_directly 12.1%
near_duplicate_pairs_on_list 19 70.0% 1: tool 59.6%, ask_user 15.9%; 2: tool 69.4%, answer_directly 12.2%; 5: tool 75.5%, answer_directly 11.5%; 4: tool 75.0%, answer_directly 11.1%; 3: tool 73.3%, answer_directly 10.8%; 7: tool 78.5%, answer_directly 9.8%
style_prompt 5 70.0% plain: tool 65.1%, answer_directly 15.6%; follow_up: tool 82.3%, out_of_scope 9.7%; multi_step: tool 77.3%, ask_user 9.7%; casual: tool 78.7%, ask_user 11.1%; polite: tool 79.4%, answer_directly 9.5%

Formatting by label class (main text, share of rows)

feature answer_directly ask_user out_of_scope tool
ends with ? 48.9% 12.6% 26.6% 18.8% gap
ends with . 32.4% 68.0% 57.4% 51.0% gap
ends with ! 6.1% 2.5% 0.1% 0.5%
no end punctuation 12.6% 16.5% 15.9% 27.9% gap
starts lowercase 27.4% 18.0% 19.6% 31.7%
all lowercase 24.8% 15.6% 17.7% 16.0%
has a digit 9.1% 30.4% 26.2% 62.2% gap
has newline 0.0% 0.0% 0.0% 0.2%
has quotes 0.0% 0.7% 0.0% 2.2%
has markup (HTML/markdown) 0.0% 0.0% 0.0% 0.0%
has URL 0.0% 0.2% 0.0% 1.2%
non-ASCII 1.8% 0.1% 0.3% 0.5%
non-Latin script 0.0% 0.0% 0.0% 0.0%
emoji 0.4% 0.0% 0.0% 0.0%
ALL-CAPS word (4+) 1.4% 4.1% 3.3% 7.7%
contains ' - ' or — 1.6% 0.2% 0.1% 0.6%

Same, by row kind

feature answer_directly ask_user out_of_scope tool
ends with ? 49% 13% 27% 19%
ends with . 32% 68% 57% 51%
ends with ! 6% 2% 0% 0%
no end punctuation 13% 17% 16% 28%
starts lowercase 27% 18% 20% 32%
all lowercase 25% 16% 18% 16%
has a digit 9% 30% 26% 62%
has newline 0% 0% 0% 0%
has quotes 0% 1% 0% 2%
has markup (HTML/markdown) 0% 0% 0% 0%
has URL 0% 0% 0% 1%
non-ASCII 2% 0% 0% 0%
non-Latin script 0% 0% 0% 0%
emoji 0% 0% 0% 0%
ALL-CAPS word (4+) 1% 4% 3% 8%
contains ' - ' or — 2% 0% 0% 1%

Over-represented words and phrases per label class (main text)

Log-odds ratio with an informative Dirichlet prior (Monroe et al. 2008), each class against all the others; z-score, then the share of rows in the class and in the other classes that contain the phrase. Counted once per row.

Words, answer_directly: what z=54 35.6% vs 6.8%; again z=34 11.0% vs 1.2%; you z=34 41.6% vs 18.2%; thanks z=32 9.0% vs 0.3%; did z=31 9.1% vs 1.0%; explain z=30 8.2% vs 0.3%; help z=29 7.5% vs 0.6%; does z=29 7.6% vs 0.8%; or z=27 12.9% vs 3.5%; how z=27 13.7% vs 4.0%; between z=27 6.8% vs 0.8%; wait z=26 8.1% vs 1.4%; good z=26 7.3% vs 1.2%; say z=25 5.5% vs 0.5%; was z=23 10.1% vs 2.9%; remind z=23 4.9% vs 0.1%; exactly z=21 9.1% vs 2.8%; me z=21 24.1% vs 12.8%; things z=20 3.7% vs 0.4%; makes z=19 3.3% vs 0.1%

Words, ask_user: the z=16 70.3% vs 57.7%; need z=15 23.0% vs 15.1%; oh z=14 2.3% vs 0.4%; yet z=13 4.4% vs 1.5%; i z=13 42.4% vs 33.7%; haven't z=13 1.9% vs 0.3%; for z=13 49.8% vs 41.4%; forgot z=12 1.9% vs 0.4%; check z=12 10.7% vs 6.4%; decided z=11 1.1% vs 0.1%; right z=11 8.1% vs 4.6%; up z=11 14.4% vs 9.8%; asap z=11 2.8% vs 1.0%; to z=10 47.1% vs 40.9%; i'm z=10 7.3% vs 4.3%; me z=10 18.3% vs 13.7%; want z=10 7.2% vs 4.2%; idk z=10 0.9% vs 0.1%; run z=10 5.4% vs 2.9%; but z=9 7.0% vs 4.2%

Words, out_of_scope: directly z=23 6.2% vs 0.6%; our z=20 18.8% vs 6.8%; into z=18 8.8% vs 2.3%; automatically z=17 3.4% vs 0.3%; from z=16 14.1% vs 5.6%; account z=15 5.8% vs 1.4%; portal z=15 3.1% vs 0.4%; change z=13 4.7% vs 1.2%; transfer z=13 2.3% vs 0.3%; now z=12 16.7% vs 8.5%; to z=12 59.1% vs 40.0%; great z=12 4.0% vs 1.1%; submit z=12 3.4% vs 0.8%; last z=12 7.5% vs 2.9%; database z=11 2.7% vs 0.6%; password z=11 1.4% vs 0.1%; delete z=10 1.8% vs 0.3%; month z=10 3.6% vs 1.1%; their z=10 8.0% vs 3.6%; payroll z=10 1.3% vs 0.1%

Words, tool: 2024 z=21 7.7% vs 0.9%; id z=17 5.8% vs 1.3%; yes z=15 3.8% vs 0.6%; yeah z=14 5.4% vs 1.9%; let's z=14 3.7% vs 0.8%; any z=13 5.1% vs 1.9%; com z=12 2.7% vs 0.4%; 00 z=12 2.6% vs 0.5%; 01 z=11 2.2% vs 0.4%; ahead z=11 6.4% vs 3.5%; 10 z=10 2.3% vs 0.7%; show z=10 3.2% vs 1.3%; 11 z=10 1.7% vs 0.3%; currently z=10 1.8% vs 0.4%; 2023 z=10 1.6% vs 0.3%; plz z=9 1.6% vs 0.1%; 12 z=9 2.0% vs 0.6%; b z=9 1.8% vs 0.5%; all z=9 6.2% vs 3.7%; been z=9 3.0% vs 1.4%

2–4-word phrases, answer_directly: me what z=24 5.9% vs 0.6%; help me z=23 4.8% vs 0.2%; what was z=21 4.5% vs 0.1%; do i z=21 5.8% vs 1.0%; remind me z=20 4.9% vs 0.1%; got it z=20 4.2% vs 0.1%; me with z=20 3.6% vs 0.1%; if i z=19 5.2% vs 0.9%; what exactly z=19 5.0% vs 0.1%; was the z=19 4.0% vs 0.1%; tell me what z=18 3.4% vs 0.4%; you can z=18 3.1% vs 0.2%; good morning z=18 3.2% vs 0.1%; kind of z=18 2.9% vs 0.1%; and a z=17 2.8% vs 0.1%; what kind of z=17 2.8% vs 0.1%; what kind z=17 2.8% vs 0.1%; able to z=17 2.7% vs 0.2%; you explain z=17 4.1% vs 0.0%; just to z=17 2.7% vs 0.1%

2–4-word phrases, ask_user: i haven't z=13 1.6% vs 0.1%; but i z=13 3.5% vs 1.0%; for me z=12 5.6% vs 2.3%; i want z=12 6.0% vs 2.7%; i forgot z=11 1.4% vs 0.2%; hey i need to z=11 1.1% vs 0.1%; hey i z=11 1.9% vs 0.4%; i want to z=11 5.3% vs 2.4%; up the z=11 5.5% vs 2.6%; hey i need z=10 1.5% vs 0.3%; forgot to z=10 1.0% vs 0.1%; run the z=10 2.4% vs 0.8%; to check z=10 1.9% vs 0.5%; right away z=10 1.2% vs 0.2%; but i haven't z=10 1.1% vs 0.0%; and i need z=10 1.7% vs 0.4%; need to z=10 13.9% vs 9.3%; the other z=9 1.3% vs 0.2%; oh wait z=9 0.8% vs 0.0%; yo need z=9 0.8% vs 0.1%

2–4-word phrases, out_of_scope: how do i z=22 5.0% vs 0.4%; how do z=22 5.1% vs 0.4%; do i z=19 6.2% vs 1.2%; directly to z=17 2.8% vs 0.2%; to the z=16 10.9% vs 4.2%; you to z=16 5.9% vs 1.6%; great now z=16 3.1% vs 0.4%; i need you z=14 3.8% vs 0.8%; i need you to z=14 3.8% vs 0.8%; need you to z=14 4.2% vs 1.0%; need you z=14 4.2% vs 1.0%; is there a z=14 2.5% vs 0.3%; there a z=14 2.5% vs 0.3%; is there a way z=14 2.0% vs 0.2%; there a way z=14 2.0% vs 0.2%; to our z=13 2.6% vs 0.4%; a way z=13 2.1% vs 0.2%; there a way to z=13 1.8% vs 0.1%; way to z=13 2.1% vs 0.3%; a way to z=13 1.8% vs 0.2%

2–4-word phrases, tool: so i z=13 10.0% vs 5.8%; so i can z=12 7.3% vs 4.0%; go ahead z=11 6.3% vs 3.3%; i can z=11 8.0% vs 4.7%; go ahead and z=11 6.1% vs 3.3%; ahead and z=11 6.1% vs 3.3%; id is z=10 1.8% vs 0.3%; do the z=10 1.8% vs 0.3%; show me z=9 2.5% vs 1.0%; hey can u z=8 1.2% vs 0.2%; need to see z=8 1.6% vs 0.5%; 2024 11 z=8 1.1% vs 0.1%; second one z=8 1.1% vs 0.2%; you pull z=8 1.5% vs 0.4%; do the same z=8 1.1% vs 0.2%; hey can z=8 1.3% vs 0.3%; just need z=8 1.2% vs 0.2%; you pull up z=8 1.2% vs 0.3%; make sure z=8 2.1% vs 1.0%; pull up z=8 3.9% vs 2.3%

Strong phrase flags (in ≥5% of one class's rows and at ≥4× the rate in the others):

  • answer_directly: what 35.6% vs 6.8%
  • answer_directly: again 11.0% vs 1.2%
  • answer_directly: did 9.1% vs 1.0%
  • answer_directly: thanks 9.0% vs 0.3%
  • answer_directly: explain 8.2% vs 0.3%
  • answer_directly: wait 8.1% vs 1.4%
  • tool: 2024 7.7% vs 0.9%
  • answer_directly: does 7.6% vs 0.8%
  • answer_directly: help 7.5% vs 0.6%
  • answer_directly: good 7.3% vs 1.2%
  • answer_directly: between 6.8% vs 0.8%
  • out_of_scope: directly 6.2% vs 0.6%
  • out_of_scope: do i 6.2% vs 1.2%
  • answer_directly: me what 5.9% vs 0.6%
  • tool: id 5.8% vs 1.3%
  • answer_directly: do i 5.8% vs 1.0%
  • out_of_scope: account 5.8% vs 1.4%
  • answer_directly: say 5.5% vs 0.5%
  • answer_directly: if i 5.2% vs 0.9%
  • out_of_scope: how do 5.1% vs 0.4%
  • answer_directly: what exactly 5.0% vs 0.1%
  • out_of_scope: how do i 5.0% vs 0.4%

Standard flags (owner's rule: a word or phrase in more than 2% of one class's rows, of whose rows at least 70% (and at least twice the base rate) belong to that class; the reviewer decides whether each is meaning or a shortcut):

  • all rows (39,203 rows): answer_directly: thanks 9.0% of class, 79% of its 536 rows; answer_directly: explain 8.2% of class, 80% of its 484 rows; answer_directly: difference 5.7% of class, 95% of its 281 rows; answer_directly: did you 5.4% of class, 98% of its 260 rows; answer_directly: difference between 5.1% of class, 100% of its 243 rows; answer_directly: what exactly 5.0% of class, 92% of its 257 rows; answer_directly: remind me 4.9% of class, 89% of its 258 rows; answer_directly: help me 4.8% of class, 76% of its 298 rows; answer_directly: what was 4.5% of class, 83% of its 256 rows; answer_directly: got it 4.2% of class, 85% of its 232 rows; answer_directly: the difference 4.1% of class, 95% of its 203 rows; answer_directly: thanks for 4.0% of class, 95% of its 198 rows; answer_directly: was the 4.0% of class, 88% of its 213 rows; answer_directly: can you explain 3.8% of class, 94% of its 191 rows; answer_directly: the difference between 3.7% of class, 99% of its 174 rows; answer_directly: me with 3.6% of class, 78% of its 218 rows; answer_directly: are you 3.5% of class, 91% of its 182 rows; answer_directly: did you say 3.5% of class, 100% of its 163 rows; answer_directly: help me with 3.3% of class, 98% of its 158 rows; answer_directly: makes 3.3% of class, 77% of its 200 rows; answer_directly: sense 3.2% of class, 78% of its 194 rows; answer_directly: good morning 3.2% of class, 80% of its 188 rows; answer_directly: between a 3.1% of class, 99% of its 149 rows; answer_directly: mean 3.1% of class, 78% of its 190 rows; answer_directly: you can 3.1% of class, 71% of its 203 rows; answer_directly: what was the 3.1% of class, 95% of its 151 rows; answer_directly: question 2.9% of class, 79% of its 175 rows; answer_directly: kind of 2.9% of class, 74% of its 183 rows; answer_directly: what kind of 2.8% of class, 81% of its 163 rows; answer_directly: and a 2.8% of class, 76% of its 173 rows; answer_directly: just to 2.7% of class, 80% of its 161 rows; answer_directly: or do 2.7% of class, 93% of its 139 rows; answer_directly: you're 2.7% of class, 73% of its 173 rows; answer_directly: difference between a 2.5% of class, 100% of its 119 rows; answer_directly: that makes 2.5% of class, 92% of its 127 rows; answer_directly: explain how 2.4% of class, 94% of its 121 rows; answer_directly: what does 2.4% of class, 75% of its 150 rows; answer_directly: and then explain 2.3% of class, 97% of its 110 rows; answer_directly: quick question 2.2% of class, 98% of its 106 rows; answer_directly: explain the 2.2% of class, 78% of its 131 rows

Shortcut models

Predicting the label class on test (5,157 rows). Chance 25.0%, majority class ('tool') 70.0%; balanced chance 25.0%.

model (logistic regression, trained on the train sample) test accuracy balanced accuracy (mean recall)
bag of words, whole state (words and word pairs) 84.2% 60.5%
bag of words, main text only (user_message) 85.4% 63.6%
surface features only (no words: length, punctuation, case, markup, digits, script, state sizes, option count, instruction kind) 74.5% 42.0%
surface features of the main text only 72.7% 35.9%

Strongest single surface features (logistic regression on one feature, balanced accuracy on test):

feature accuracy balanced accuracy
count_! 70.7% 27.4%
ends_! 70.4% 26.4%
chars(log) 70.0% 25.0%
words(log) 70.0% 25.0%
upper_ratio 70.0% 25.0%
digit_ratio 70.0% 25.0%
nonascii_ratio 70.0% 25.0%
nonlatin 70.0% 25.0%
emoji 70.0% 25.0%
newlines 70.0% 25.0%

Other state fields alone (predicting the label class on test from one field, without the main text):

field treated as accuracy balanced accuracy
agent text: bag of words / length+empty 70.0% / 70.0% 25.0% / 25.0%
conversation text: bag of words / length+empty 74.6% / 70.1% 38.7% / 25.1%

No-meaning option picker: a logistic ranker scores each option from its position, length, key type, fixed-option identity and shape (commas, brackets, capitals), never reading the state or the option's words, and picks the top option per row.

  • Test accuracy 20.0% against uniform chance 4.5% (this includes the fixed options, whose share is a class prior).
  • Among rows whose answer is a listed option (3,611), picking only among listed options: 5.5% against chance 4.8%.

2. Duplicates and split separation

Families shared between splits

splits shared families examples
train ∩ dev 0
train ∩ calibration 0
train ∩ test 0
dev ∩ calibration 0
dev ∩ test 0
calibration ∩ test 0
  • Train rows whose main text repeats an earlier row's (normalised): 153 (0.4%); groups: 79; largest group 10.
  • Train rows identical in the whole prompt (state, options, instructions): 0.
  • Identical whole prompt, different answer: 0 groups (0 rows).
  • Identical main text, different label class: 4 groups (13 rows). (Can be legitimate when the rest of the state or the options differ.)
    • "What about the other one?" ×6: ask_user 5, tool 1
    • "Now do the same for the second one." ×3: tool 2, ask_user 1
    • "Go ahead and book it." ×2: tool 1, ask_user 1
    • "Go ahead and run that search now." ×2: ask_user 1, tool 1

Most repeated main texts in train:

  • ×10: "yes, do it." (tool 10)
  • ×8: "yes, go ahead and book it." (tool 8)
  • ×8: "yeah do it" (tool 8)
  • ×7: "yes do it" (tool 7)
  • ×6: "yes, go ahead and send it." (tool 6)
  • ×6: "yes, go ahead and run it." (tool 6)
  • ×6: "what about the other one?" (ask_user 5, tool 1)
  • ×6: "got it thanks bye" (answer_directly 6)

Main text of held-out rows found verbatim in train (normalised; the leak gate ignores short texts shared by many items):

split rows examples
dev 15 (0.7%) "what about the other one?"; "Yeah go ahead with it."; "Yeah, do it."; "parking rules"
calibration 17 (0.8%) "Yes, do it for that one."; "Yes, do it."; "got it, thx"; "Yes, go ahead and run it."
test 22 (0.4%) "Yes, go ahead and generate it."; "Yes, go ahead and pull that up."; "Yes do it."; "Yes, go ahead and run it."

Near duplicates (MinHash, word 3-gram Jaccard ≥ 0.8 on the main text)

  • Train: 695 near-duplicate pairs; 322 rows (0.8%) sit in 96 clusters; largest cluster 23; excess rows (cluster size − 1) 226 (0.6%).
  • Clusters with more than one label class: 4 (36 rows).
    • ×23: "Yes, do the same for the second one." → tool 22, ask_user 1
    • ×6: "What about the other one?" → ask_user 5, tool 1
    • ×5: "Okay, go ahead and run that search now." → tool 4, ask_user 1
    • ×2: "Go ahead and book it." → tool 1, ask_user 1
  • Held-out rows with a near duplicate in train: dev 16 (0.8%), calibration 23 (1.0%), test 42 (0.8%)
    • train "yes do it" ~ calibration "Yes do it" (J=1.00)
    • train "hey, you there?" ~ calibration "hey, you there?" (J=1.00)
    • train "Yep, do it." ~ calibration "Yep do it" (J=1.00)
    • train "Yes, go ahead and send it." ~ test "Yes, go ahead and send it." (J=1.00)
    • train "yeah do it" ~ dev "Yeah, do it." (J=1.00)

Largest train clusters:

  • ×23: "Yes, do the same for the second one."
  • ×18: "yes do it"
  • ×13: "Yeah, do it."
  • ×10: "Yes, go ahead and run it now."
  • ×10: "Yes, go ahead and send it."

5. Junk

split empty main text main text under 10 characters
train 0 84
dev 0 6
calibration 0 9
test 0 6

Very short train examples: "my holds" (tool); "yes do it" (tool); "yes do it" (tool); "hi there" (answer_directly); "API_Call" (tool); "Use 2021." (tool); "ok do it" (tool); "active rx" (tool); "FAQ docs" (tool); "yes do it" (tool); "code 504" (tool); "fireworks" (tool)

Pattern scan of train main texts (count, then the share of each class's rows):

pattern rows by class
placeholder [NAME]-style 0 (0.0%)
lorem ipsum 0 (0.0%)
TODO/TBD/FIXME 0 (0.0%)
'As an AI' / refusal 0 (0.0%)
chat preamble ('Here is/are...', 'Sure!') 127 (0.3%) ask_user 0.2%, out_of_scope 0.0%, tool 0.4%
meta words (example/variation/message:) 0 (0.0%)
model thinking tags 0 (0.0%)
JSON/code-fence leftovers 2 (0.0%) tool 0.0%
encoding garbage (mojibake/replacement char) 0 (0.0%)
HTML tag 6 (0.0%) tool 0.0%
HTML entity 0 (0.0%)
base64-like run (40+ chars) 23 (0.1%) tool 0.1%
URL 343 (0.9%) ask_user 0.2%, tool 1.2%
  • chat preamble ('Here is/are...', 'Sure!'): tools-equipment_maintenance_bot-119424 (tool): Here is the first one for the report: QTR-990. | tools-permit_faster-078588 (tool): Sure, I live in 33139. I just want to make sure I don't get fined by the city for putting it too close to the property line or making it to… | tools-b2b_procurement_portal-093260 (tool): Here is the information you asked for: PO ID is PO-60199 and the URL is https://sharepoint.corp.com/fastenal/corrected_v2.pdf. Let's get th…

  • JSON/code-fence leftovers: tools-crypto_quant_analyst-124091 (tool): Here is the JSON payload from the committee's decision engine to apply to our trading book: ⏎ { ⏎ "action": "rebalance", ⏎ "weights": {… | tools-clinic_frontdesk_bot-033088 (tool): Here is the payload he gave me: ⏎ [ ⏎ {"first": "Liam", "last": "Neeson", "email": "liam.n@test.com"}, ⏎ {"first": "Morgan", "last": "F…

  • HTML tag: tools-edu_tutor_assistant-001882 (tool): i got this error Traceback (most recent call last): ⏎ File "hw3.py", line 8, in <module> ⏎ result = divide(a, b) ⏎ File "hw3.py", l… | tools-edu_tutor_assistant-001881 (tool): im working on the lab 2 assignment where we have to read a csv file and calculate averages. i keep running into this crash every time i try… | tools-edu_tutor_assistant-001885 (tool): Hello, I hope you are doing well. I am terribly sorry to bother you, but I was wondering if you would be so kind as to help me understand w…

  • base64-like run (40+ chars): tools-soc_tier1_triage-120728 (tool): Can you check if the SHA256 hash 8d7b3f9a1c2e4f6a8b0c2d4e6f8a0b2c4d6e8f0a2b4c6d8e0f2a4b6c8d0e2f4a is known malware or benign? | tools-legal_contract_translator-138445 (tool): Could you upload the executed master services agreement for matter MAT-2024-0891 to the secure vault? The base64 encoded file content is JV… | tools-tournament_ticker-047386 (tool): First, upload the logo file winter_brawl_logo.svg as an image/svg+xml with the binary data PHN2ZyB4bWxucz0iaHR0cDovL3d3dy53My5vcmcvMjAwMC9z…

  • URL: tools-global_expatriate_tax-128022 (tool): Here it is: https://japan-tax.jp/rent_march.pdf | tools-ecommerce_product_sync-140894 (tool): Great. Now do the same validation for the third image in that set, which is https://media.leathercraft.co/bags/model-X10/detail-stitching.j… | tools-small_business_press_kit-091194 (tool): Could you log the new article mention that just went live at https://www.dailyherald.com/news/2024/local-bakery-expansion and calculate the…

  • Possibly cut off: 13 of 1,709 train main texts over 300 characters end mid-sentence (letter, digit or comma). By class: answer_directly 0.0%, ask_user 0.0%, out_of_scope 0.0%, tool 0.8%

    • tools-legal_case_manager-068017: …act filename. could u search the files for that case using the keyword timeline so i can figure out which one it is? thx
    • tools-executive_chief_of_staff-029287: … to archive everything coming from marketing-spam.net so he can actually see what matters before his 2pm call, pls hurry
    • tools-ip_patent_prior_art-041295: … from our application, can u run an image match on it? https://storage.myfirm.internal/cases/4421/hinge-drawing-fig3.png

Repeated sentences across rows (≥25 characters, in at least 0.2% of the sample):

  • none

6. Samples

20 random train rows per kind: tools-samples.txt. Reading notes are in the findings above.