navVoice navigation
Turns a spoken command, even a misheard one, into an item on the current app screen, a question, or none of these.
Data: Generated mainly with Qwen3.8-Max (hosted, Alibaba Cloud DashScope), used to finish in time; apps and some stages by Qwen3.8-Flash-Next (local); checked by Qwen3.8-Flash and DeepSeek-V4-Flash (hosted).
Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap
Use it when
- Your app takes voice commands and you want the item on the current screen the user meant, matched by meaning and by sound.
- You can list every item on the screen as an option. Training screens mostly had 40 to 120 items, some up to 200, and dialogues 4 to 8.
- You also need to tell a question, or a request the screen does not offer, apart from a navigation command.
Not a good fit when
- The user's words exactly match an item's name. Match those in code first; training left such cases out.
- You need the command carried out or an answer to the user's question. Jeff only picks an option.
- You want one app's best accuracy and have data from that app. A separate full fine-tune on one app's own screens (a different model) scored 95.2% on navigation screens with a 10-item shortlist; this general adapter is not yet measured.
- Your users mostly speak languages other than English. All training data is English.
Request format
The state is an object with these fields, in this order. Only voice_transcript_may_contain_errors changes from request to request, so it comes last and the rest can be prepared in advance.
| State field | Changes per request | What goes in it |
|---|---|---|
current_screen | No | Where the user is, in plain words, for example "Tallow Books › Clients". |
previous_screen | No | The screen the user came from, or null. |
dialogue_step | No | Null, or the pending question when the app has asked the user to choose, such as "Which did you mean?". |
voice_transcript_may_contain_errors | Yes | What speech recognition heard. It may contain misheard words, split words and fillers. |
| Question | Type | What it decides |
|---|---|---|
target | Choice | Which item on the screen the user wants, or whether they are asking a question, or want something not listed. Options: |
- Use the instructions below word for word; the adapter was trained mostly on them.
- Keep the two fixed options first, with these exact texts.
ask_questionis "Not a navigation request: the user is asking a question";none_of_theseis "None of these: the user wants something not listed". - List every item on the screen; do not shortlist. Use at most 254 options in all. Item order does not matter; it was shuffled in training.
- Keep the state keys in this order, with the transcript last, so the unchanging part of the request can be prepared in advance.
General rules for every request are in the request format guide.
Example
The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).
from jeff import Client
from jeff.client import choice_question
jeff = Client("http://localhost:8765", model="nav")
state = {
"current_screen": "Tallow Books › Clients",
"previous_screen": "Tallow Books › Home",
"dialogue_step": None,
"voice_transcript_may_contain_errors": "um open bren and roofing supplies",
}
answers = jeff.ask(state, {
"target": choice_question(
{
"ask_question": "Not a navigation request: the user is asking a question",
"none_of_these": "None of these: the user wants something not listed",
"o1": "Home",
"o2": "Search",
"o3": "client Brennan Roofing",
"o4": "client Brennon Roofing Supplies",
"o5": "invoice 2231 → payments",
"o6": "Settings",
},
"Which of these does the user want? The voice transcript comes from speech recognition and may contain misheard words, so match by meaning and by sound. If the user is asking a question rather than asking to go somewhere or do something, choose the question option. If none of the listed options is what they want, choose none of these.",
),
})
print("target", answers.choice("target").key)import { Client, choiceQuestion } from '@jeff/client';
const jeff = new Client({ url: 'http://localhost:8765', model: 'nav' });
const state = {
current_screen: 'Tallow Books › Clients',
previous_screen: 'Tallow Books › Home',
dialogue_step: null,
voice_transcript_may_contain_errors: 'um open bren and roofing supplies',
};
const answers = await jeff.ask(state, {
target: choiceQuestion(
{
ask_question: 'Not a navigation request: the user is asking a question',
none_of_these: 'None of these: the user wants something not listed',
o1: 'Home',
o2: 'Search',
o3: 'client Brennan Roofing',
o4: 'client Brennon Roofing Supplies',
o5: 'invoice 2231 → payments',
o6: 'Settings',
},
'Which of these does the user want? The voice transcript comes from speech recognition and may contain misheard words, so match by meaning and by sound. If the user is asking a question rather than asking to go somewhere or do something, choose the question option. If none of the listed options is what they want, choose none of these.',
),
});
console.log('target', answers.target.key);curl -s http://localhost:8765/v1/systemone \
-H 'content-type: application/json' \
-d '{
"model": "nav",
"state": {
"current_screen": "Tallow Books › Clients",
"previous_screen": "Tallow Books › Home",
"dialogue_step": null,
"voice_transcript_may_contain_errors": "um open bren and roofing supplies"
},
"questions": {
"target": {
"type": "choice",
"instructions": "Which of these does the user want? The voice transcript comes from speech recognition and may contain misheard words, so match by meaning and by sound. If the user is asking a question rather than asking to go somewhere or do something, choose the question option. If none of the listed options is what they want, choose none of these.",
"criteria": {
"ask_question": "Not a navigation request: the user is asking a question",
"none_of_these": "None of these: the user wants something not listed",
"o1": "Home",
"o2": "Search",
"o3": "client Brennan Roofing",
"o4": "client Brennon Roofing Supplies",
"o5": "invoice 2231 → payments",
"o6": "Settings"
}
}
}
}'Response
{
"model": "nav",
"answers": {
"target": {
"type": "choice",
"probabilities": {
"ask_question": 0.000025296381382346897,
"none_of_these": 0.00012356457356308372,
"o1": 0.000019156673805856204,
"o2": 0.0000398589428784838,
"o3": 0.0010593527493508617,
"o4": 0.9986896920831212,
"o5": 0.00003204815887259614,
"o6": 0.000011030437025631684
},
"choice": "o4",
"confidence": 0.9985025052378528
}
},
"usage": {
"input_tokens": 275,
"output_tokens": 0,
"orders": 1
}
}Results
On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters
| Test set | Test rows | Qwen3.5-0.8B untrained | Jeff v1.2 0.8B alone | Jeff v1.2 0.8B + adapter |
|---|---|---|---|---|
nav | 3,300 | 12.6% · 0.038 | 23.8% · 0.195 | 97.0% · 0.005 |
nav3,300 test rows- Qwen3.5-0.8B untrained
- 12.6% · 0.038
- Jeff v1.2 0.8B alone
- 23.8% · 0.195
- Jeff v1.2 0.8B + adapter
- 97.0% · 0.005
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
How sure is it, and is it right?
Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.
When this adapter says it is about 99.5% sure, it is right about 99.5% of the time (3,024 test rows).
Point at a dot for its numbers. Bigger dots hold more rows.
| Model | Calibration error (ECE) | Brier score | Log loss |
|---|---|---|---|
| Qwen3.5-0.8B untrained | 0.038 | 0.947 | 3.691 |
| Jeff v1.2 0.8B alone | 0.195 | 0.931 | 2.977 |
| Jeff v1.2 0.8B + adapter | 0.005 | 0.044 | 0.099 |
All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.
Where it gets things wrong
none_of_theseread as a listed option: 50 rows- a listed option read as another listed option: 21 rows
- a listed option read as
none_of_these: 9 rows nothing_matchesread asq_question: 1 rowsnone_of_theseread asopt_62: 1 rows
The five commonest mistakes with the adapter, out of 3,300 test rows.
| Right answer | Test rows | Accuracy with the adapter |
|---|---|---|
| a listed option | 2,154 | 98.6% |
none_of_these | 557 | 89.6% |
ask_question | 318 | 100.0% |
none | 6 | 83.3% |
no_match | 5 | 80.0% |
not_listed | 5 | 100.0% |
question | 5 | 100.0% |
opt_5 | 5 | 100.0% |
item_25 | 4 | 100.0% |
nothing_matches | 4 | 75.0% |
Show all 194 answers
| Right answer | Test rows | Accuracy with the adapter |
|---|---|---|
| a listed option | 2,154 | 98.6% |
none_of_these | 557 | 89.6% |
ask_question | 318 | 100.0% |
none | 6 | 83.3% |
no_match | 5 | 80.0% |
not_listed | 5 | 100.0% |
question | 5 | 100.0% |
opt_5 | 5 | 100.0% |
item_25 | 4 | 100.0% |
nothing_matches | 4 | 75.0% |
q_question | 4 | 100.0% |
opt_4 | 3 | 100.0% |
opt_31 | 3 | 100.0% |
item_36 | 3 | 100.0% |
item_61 | 3 | 100.0% |
item_19 | 3 | 100.0% |
opt_2 | 3 | 100.0% |
item_21 | 3 | 100.0% |
opt_3 | 3 | 100.0% |
opt_41 | 3 | 66.7% |
opt_27 | 3 | 100.0% |
is_question | 3 | 100.0% |
opt_1 | 3 | 100.0% |
item_3 | 2 | 100.0% |
item_39 | 2 | 100.0% |
schedule_board | 2 | 100.0% |
opt_52 | 2 | 100.0% |
opt_9 | 2 | 100.0% |
opt_29 | 2 | 100.0% |
opt_83 | 2 | 100.0% |
item_74 | 2 | 100.0% |
item_26 | 2 | 100.0% |
opt_19 | 2 | 50.0% |
item_33 | 2 | 100.0% |
opt_42 | 2 | 100.0% |
item_30 | 2 | 100.0% |
item_46 | 2 | 100.0% |
opt_6 | 2 | 100.0% |
item_12 | 2 | 100.0% |
cancel | 2 | 100.0% |
item_38 | 2 | 100.0% |
log_reading | 2 | 50.0% |
item_7 | 2 | 100.0% |
item_27 | 2 | 100.0% |
item_4 | 2 | 100.0% |
item_64 | 2 | 100.0% |
item_22 | 2 | 100.0% |
back | 2 | 100.0% |
opt_23 | 2 | 100.0% |
machine_latpulldown_v3 | 1 | 100.0% |
document_board_consent_resol | 1 | 100.0% |
item_115 | 1 | 100.0% |
item_58 | 1 | 100.0% |
opt_34 | 1 | 100.0% |
item_18 | 1 | 100.0% |
opt_167 | 1 | 100.0% |
opt_97 | 1 | 100.0% |
schedule_december_holiday_co | 1 | 100.0% |
item_17 | 1 | 100.0% |
member_lily_age_14 | 1 | 100.0% |
item_48 | 1 | 100.0% |
researcher_clara_whitfield | 1 | 100.0% |
user_account_dr_silas_wrenn | 1 | 100.0% |
opt_46 | 1 | 100.0% |
add_guest | 1 | 100.0% |
opt_81 | 1 | 100.0% |
opt_15 | 1 | 100.0% |
paycheck_sept_overtime_slip | 1 | 100.0% |
item_14 | 1 | 100.0% |
item_45 | 1 | 100.0% |
opt_13 | 1 | 100.0% |
opt_12 | 1 | 100.0% |
opt_54 | 1 | 100.0% |
item_15 | 1 | 100.0% |
download_pdf | 1 | 100.0% |
invoice_inv_2024_010_alder_d | 1 | 100.0% |
item_32 | 1 | 100.0% |
dashboard | 1 | 100.0% |
user_zara_lin_activity | 1 | 100.0% |
sign_out | 1 | 100.0% |
session_session_4021b | 1 | 100.0% |
opt_118 | 1 | 100.0% |
item_91 | 1 | 100.0% |
vip_lounge | 1 | 100.0% |
opt_37 | 1 | 100.0% |
book_paper_boats_and_rain_no | 1 | 100.0% |
hotel_island_oasis_resort_ph | 1 | 100.0% |
opt_39 | 1 | 100.0% |
opt_35 | 1 | 100.0% |
opt_33 | 1 | 100.0% |
webhook_datadog_metrics_push | 1 | 100.0% |
archive_item_homestead_deed_ | 1 | 100.0% |
opt_78 | 1 | 100.0% |
opt_74 | 1 | 100.0% |
item_78 | 1 | 100.0% |
item_35 | 1 | 100.0% |
message_template_check_in_in_2 | 1 | 100.0% |
opt_73 | 1 | 100.0% |
opt_67 | 1 | 100.0% |
property_lantern_house | 1 | 100.0% |
export_pdf | 1 | 100.0% |
opt_180 | 1 | 100.0% |
mood_board_oceanic_blues_col | 1 | 100.0% |
opt_112 | 1 | 100.0% |
remove_from_shelf | 1 | 100.0% |
member_elias_thorne | 1 | 100.0% |
book_paper_boats_and_rain_re | 1 | 0.0% |
throttle_stream | 1 | 100.0% |
bookmark_letter_from_aunt_ma | 1 | 0.0% |
account_seed_fund_checking | 1 | 100.0% |
paycheck_tips_transfer_aug_8_3 | 1 | 100.0% |
search | 1 | 100.0% |
account_restful_nights_co | 1 | 100.0% |
notifications | 1 | 100.0% |
delete_template | 1 | 100.0% |
opt_36 | 1 | 100.0% |
opt_11 | 1 | 100.0% |
cleaning_task_inspection_dri | 1 | 100.0% |
item_56 | 1 | 100.0% |
automation_rule_vip_customer | 1 | 100.0% |
user_account_prof_elena_voss | 1 | 100.0% |
my_points | 1 | 100.0% |
booking_com | 1 | 100.0% |
payment_tarjeta_visa_oro | 1 | 100.0% |
document_opinion_of_counsel | 1 | 100.0% |
employee_kevin_hart | 1 | 100.0% |
purge_stream | 1 | 100.0% |
portfolio_future_me_money_ov | 1 | 100.0% |
pronunciation_metric_intonat | 1 | 100.0% |
opt_38 | 1 | 100.0% |
asset_brand_guidelines_pdf | 1 | 100.0% |
new_observation | 1 | 100.0% |
quick_add | 1 | 100.0% |
item_49 | 1 | 100.0% |
project_fairview_ancestral_r_2 | 1 | 100.0% |
opt_26 | 1 | 100.0% |
opt_14 | 1 | 100.0% |
zone_sector_bravo_7_nodes | 1 | 100.0% |
zone_d_visitor | 1 | 100.0% |
automation_rule_tracking_del | 1 | 100.0% |
filter_specialty | 1 | 100.0% |
new_clause | 1 | 100.0% |
pick_one | 1 | 100.0% |
item_67 | 1 | 100.0% |
opt_69 | 1 | 100.0% |
play_digest | 1 | 100.0% |
everyone | 1 | 100.0% |
clause_governing_law | 1 | 100.0% |
opt_66 | 1 | 100.0% |
item_82 | 1 | 100.0% |
item_41 | 1 | 100.0% |
item_24 | 1 | 100.0% |
item_108 | 1 | 100.0% |
clients | 1 | 100.0% |
opt_70 | 1 | 100.0% |
application_visa_app_mexico_ | 1 | 100.0% |
opt_51 | 1 | 100.0% |
team_member_kevin_doyle | 1 | 100.0% |
reservation_stay_july_12_pay | 1 | 100.0% |
item_112 | 1 | 100.0% |
thank_you_note_mrs_hendricks | 1 | 100.0% |
item_5 | 1 | 100.0% |
item_37 | 1 | 0.0% |
opt_71 | 1 | 100.0% |
contract_lisbon_lease | 1 | 100.0% |
opt_63 | 1 | 100.0% |
portfolio_steady_eddie_plan__2 | 1 | 100.0% |
candidate_michael_o_connor_s | 1 | 100.0% |
property_birchwood_apartment_2 | 1 | 100.0% |
policy_vidaplena_familiar_me | 1 | 100.0% |
pdf_document | 1 | 0.0% |
property_birchwould_flats | 1 | 100.0% |
item_122 | 1 | 100.0% |
budget_hair_cut_budget_break | 1 | 100.0% |
opt_50 | 1 | 100.0% |
item_1 | 1 | 100.0% |
export_all | 1 | 100.0% |
cancel_order | 1 | 100.0% |
item_89 | 1 | 100.0% |
guest_nathan_pryce | 1 | 100.0% |
project_kestrelos_contributo | 1 | 100.0% |
message_family_introduction__2 | 1 | 100.0% |
opt_104 | 1 | 100.0% |
opt_90 | 1 | 100.0% |
item_103 | 1 | 100.0% |
item_9 | 1 | 100.0% |
claim_emergencia_ambulancia | 1 | 100.0% |
project_oakridge_migration_s | 1 | 100.0% |
candidate_michael_o_connor_p | 1 | 100.0% |
delete_receipt | 1 | 100.0% |
station_gatehouse_terminal | 1 | 100.0% |
user_account_prof_elena_voss_2 | 1 | 100.0% |
session_weekend_idle_drain | 1 | 100.0% |
item_40 | 1 | 100.0% |
“A listed option” pools the rows whose right answer is one of the options listed in that request (keys such as o3 or t1, whose meaning changes from row to row); “another listed option” is a different one of them.
Against Qwen3.8-27B
More accurate, 35× faster than Qwen3.8-27B alone.
Gain over Qwen3.8-27B alone: +5.7 points [+2.7, +9.0] (95% interval).
On 300 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 91.3% of the time at 7.21 s per query on average; with Jeff and this adapter answering first, it was right 97.0% at 208 ms.
This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.
| Route | Mean | Median | 95th percentile | Typical prompt |
|---|---|---|---|---|
| Qwen3.8-27B alone | 7.21 s | 6.60 s | 13.81 s | 1,077 tokens |
| Jeff + adapter, its own answer | 208 ms | 202 ms | 406 ms | 1,044 tokens |
Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.
This task's threshold 0.00: accuracy 97.0%, 34.6× faster, 0.0% sent on to Qwen3.8-27B.
Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 300 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.
How it was trained
- Training rows
- 79,529
- Steps
- 1,243
- Training time
- 144 min
- Size as saved
- 41.5 MB
One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-nav-20261001-0352.
Source: jeff-finetunes/adapters/BASELINE.md
Data card
Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean
- Test set, so anyone can check the numbers
- Calibration rows, the rows its threshold is chosen on
- QA report, sanitised: the data-quality checks run before training
- How the test set was held out
- Requests from about 10% of the roughly 340 generated apps, never trained on.
- Training data
- Training data not published.
- Which models made the data, counted on the 72,299 training rows
- Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. nav's rows record only the checking model; the hand-over file (READY-nav) records which models wrote each stage, listed below.
What it did Model Where it ran Training rows Checked the row checkerQwen3.8-Flash hosted (Alibaba Cloud DashScope) 72,299 Which models did each stage, from the hand-over fileStage Models Apps (stage 1) and brand check Qwen3.8-Flash-Next (local) Navigation systems (stage 2) Qwen3.8-Max (hosted): 233 apps; Qwen3.8-Flash-Next (local): 107 apps Format pools (instructions, fixed-option wordings) Qwen3.8-Flash-Next (local) Utterances (stage 4) Qwen3.8-Max (hosted): 340 apps Supplement utterances Qwen3.8-Max (hosted): 340 apps Blind check (stage 6) Qwen3.8-Flash (hosted): 340 apps Second opinion on training none-of-these and question rows DeepSeek-V4-Flash (hosted): 340 apps Length-balance sentences round 1 (extra.py: 3-6-word questions and remarks, long requests) Qwen3.8-Max (hosted): 340 apps Length-balance sentences round 2 (extra2.py: 2-3-word and 6-8-word questions, 2-3-word remarks, 9-13-word requests for navigation and none-of-these rows) Qwen3.8-Max (hosted): 340 apps Length-balance sentences round 3 (extra3.py: 3-8-word questions with a filler word) Qwen3.8-Max (hosted): 340 apps Blind check (stage 6) of length-balance rows round 1: Qwen3.8-Flash (hosted): 340 apps; round 2: Qwen3.8-Flash (hosted): 340 apps; round 3: Qwen3.8-Flash (hosted): 340 apps Strict question check (every question candidate) round 1 (old and round-1 questions): DeepSeek-V4-Flash (hosted): 340 apps; round 2: DeepSeek-V4-Flash (hosted): 340 apps; round 3: DeepSeek-V4-Flash (hosted): 340 apps Length-balance sentences rounds 4 and 5 (extra4.py, extra5.py: requests of a stated word count, written alike and made navigation or none-of-these rows at random afterwards; 2-6-word questions with articles; remarks with fillers) round 4: Qwen3.8-Max (hosted): 340 apps; round 5: Qwen3.8-Max (hosted): 340 apps Blind check (stage 6) of rounds 4 and 5 round 4: Qwen3.8-Flash (hosted): 340 apps; round 5: Qwen3.8-Flash (hosted): 340 apps Strict question check rounds 4 and 5 round 4: DeepSeek-V4-Flash (hosted): 340 apps; round 5: DeepSeek-V4-Flash (hosted): 340 apps Second opinion on training none-of-these rows of rounds 2, 4 and 5 extra2: DeepSeek-V4-Flash (hosted): 286 apps; extra4: DeepSeek-V4-Flash (hosted): 271 apps; extra5: DeepSeek-V4-Flash (hosted): 270 apps Screens, garbling, code checks code, no model - The terms of the hosted model providers are being checked for training and publication use.
Data and licence
The adapter is released under Apache-2.0. It was trained on:
- Generated apps, screens and spoken commandsLicence: Released with the adapter under Apache-2.0 · Generated by Qwen3.8-Max (hosted) and Qwen3.8-Flash-Next (local), per stage; see the data card
About 340 made-up apps and websites, their screens and items, and how people say each command aloud, written by a language model and code, and checked by a second blind pass (see the data card). No real user data. Sound-alike speech-recognition errors were made in code with the CMU Pronouncing Dictionary (BSD-2-Clause, commit 74790861f652).
Changelog
- 0.1.0 · 2026-10-01Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.
Comments
Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.
