Data guidelines
A ten-point checklist for an adapter's data set, so it trains well on v1.2 today and carries over to v1.3 unchanged.
Adapters are tied to the exact base model they were trained on, so an adapter trained on v1.2 has to be retrained for v1.3. Your data set is what carries over, unchanged, as long as it follows these rules. Invest in the data, not in v1.2 weights. The lessons behind the rules are in Training an adapter; the adapter kit will check most of them for you.
The checklist
- Row format. JSON Lines in Jeff's row format, with
id,family,state,question,labelandtarget. The label must be one of the row's options. Why: the trainer and the evaluator read exactly these fields. Check it with check-rows. - The changing state field last. Put what changes per request (the user's message, a transcript) last, and everything that stays the same for a screen, app or company before it. Why: v1.3 prepares the unchanging part of the prompt once and reuses it, which is what lets your data carry over. check-rows checks the order.
- Option keys never bare numbers. Use
o1,k2or a short word, and shuffle options per row unless their order means something (a score scale). Why: the clients refuse number-like keys, and a fixed order teaches the position. - Fixed instructions. Mostly one canonical wording per question, with some rewordings. Why: the adapter learns the wording it will see in production, without becoming brittle.
- Split by family, not by row. Hold out about 10% of whole families (apps, companies, sources) for the test. Why: the test then measures new situations, not memory. split does this.
- No leaks. No test row, and no near-copy of one, in training. Keep external benchmarks out of training entirely. Why: a leaked row inflates the score. leak-check fails on any.
- No shortcuts. Check that length, option position, option length, punctuation and filler words do not give the answer away. Re-check after every fix, because fixing one shortcut can create another. Why: shortcuts score well on look-alike test data and fail real users. See the lessons; shortcut-report measures them.
- Provenance. Record every source's licence, and for generated data, which model wrote and which checked each
row, and where it ran, in the row's
sourceobject (for example"source": {"written_by": "…", "checked_by": "…"}). Why: you need both for the data card, and readers need them to judge the data. - Measure three ways on the same test set: the untrained base, Jeff alone, and Jeff with the adapter. Why: only the comparison shows what the adapter adds. evaluate-three-ways runs all three.
- Someone who did not build the data reviews it before training. Why: our independent reviews caught problems the builders had missed.
Next: Training an adapter
