Skip to content
JeffHub
VoiceOfficialReproducedVersion 0.1.0

navVoice navigation

Turns a spoken command, even a misheard one, into an item on the current app screen, a question, or none of these.

Data: Generated mainly with Qwen3.8-Max (hosted, Alibaba Cloud DashScope), used to finish in time; apps and some stages by Qwen3.8-Flash-Next (local); checked by Qwen3.8-Flash and DeepSeek-V4-Flash (hosted).

Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap

Use it when

  • Your app takes voice commands and you want the item on the current screen the user meant, matched by meaning and by sound.
  • You can list every item on the screen as an option. Training screens mostly had 40 to 120 items, some up to 200, and dialogues 4 to 8.
  • You also need to tell a question, or a request the screen does not offer, apart from a navigation command.

Not a good fit when

  • The user's words exactly match an item's name. Match those in code first; training left such cases out.
  • You need the command carried out or an answer to the user's question. Jeff only picks an option.
  • You want one app's best accuracy and have data from that app. A separate full fine-tune on one app's own screens (a different model) scored 95.2% on navigation screens with a 10-item shortlist; this general adapter is not yet measured.
  • Your users mostly speak languages other than English. All training data is English.

Request format

The state is an object with these fields, in this order. Only voice_transcript_may_contain_errors changes from request to request, so it comes last and the rest can be prepared in advance.

State fieldChanges per requestWhat goes in it
current_screenNoWhere the user is, in plain words, for example "Tallow Books › Clients".
previous_screenNoThe screen the user came from, or null.
dialogue_stepNoNull, or the pending question when the app has asked the user to choose, such as "Which did you mean?".
voice_transcript_may_contain_errorsYesWhat speech recognition heard. It may contain misheard words, split words and fillers.
QuestionTypeWhat it decides
targetChoiceWhich item on the screen the user wants, or whether they are asking a question, or want something not listed.

Options: ask_question and none_of_these first, in that order, then every item on the screen as o1, o2, … with a short text in the app's own style, such as "Settings", "client Brennan Roofing" or "invoice 2231 → payments".

  • Use the instructions below word for word; the adapter was trained mostly on them.
  • Keep the two fixed options first, with these exact texts. ask_question is "Not a navigation request: the user is asking a question"; none_of_these is "None of these: the user wants something not listed".
  • List every item on the screen; do not shortlist. Use at most 254 options in all. Item order does not matter; it was shuffled in training.
  • Keep the state keys in this order, with the transcript last, so the unchanging part of the request can be prepared in advance.

General rules for every request are in the request format guide.

Example

The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).

from jeff import Client
from jeff.client import choice_question

jeff = Client("http://localhost:8765", model="nav")

state = {
    "current_screen": "Tallow Books › Clients",
    "previous_screen": "Tallow Books › Home",
    "dialogue_step": None,
    "voice_transcript_may_contain_errors": "um open bren and roofing supplies",
}

answers = jeff.ask(state, {
    "target": choice_question(
        {
            "ask_question": "Not a navigation request: the user is asking a question",
            "none_of_these": "None of these: the user wants something not listed",
            "o1": "Home",
            "o2": "Search",
            "o3": "client Brennan Roofing",
            "o4": "client Brennon Roofing Supplies",
            "o5": "invoice 2231 → payments",
            "o6": "Settings",
        },
        "Which of these does the user want? The voice transcript comes from speech recognition and may contain misheard words, so match by meaning and by sound. If the user is asking a question rather than asking to go somewhere or do something, choose the question option. If none of the listed options is what they want, choose none of these.",
    ),
})
print("target", answers.choice("target").key)

Response

{
  "model": "nav",
  "answers": {
    "target": {
      "type": "choice",
      "probabilities": {
        "ask_question": 0.000025296381382346897,
        "none_of_these": 0.00012356457356308372,
        "o1": 0.000019156673805856204,
        "o2": 0.0000398589428784838,
        "o3": 0.0010593527493508617,
        "o4": 0.9986896920831212,
        "o5": 0.00003204815887259614,
        "o6": 0.000011030437025631684
      },
      "choice": "o4",
      "confidence": 0.9985025052378528
    }
  },
  "usage": {
    "input_tokens": 275,
    "output_tokens": 0,
    "orders": 1
  }
}

Results

On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters

  • nav3,300 test rows
    Qwen3.5-0.8B untrained
    12.6% · 0.038
    Jeff v1.2 0.8B alone
    23.8% · 0.195
    97.0% · 0.005

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

How sure is it, and is it right?

Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.

When this adapter says it is about 99.5% sure, it is right about 99.5% of the time (3,024 test rows).

Jeff v1.2 0.8B aloneJeff v1.2 0.8B + adapterperfectly calibrated
0%0%25%25%50%50%75%75%100%100%Stated confidenceRight answersJeff v1.2 0.8B alone: 54 rows stated 12% on average and were right 20.4% of the timeJeff v1.2 0.8B alone: 270 rows stated 17% on average and were right 12.2% of the timeJeff v1.2 0.8B alone: 429 rows stated 24% on average and were right 21.7% of the timeJeff v1.2 0.8B alone: 507 rows stated 30% on average and were right 19.9% of the timeJeff v1.2 0.8B alone: 435 rows stated 37% on average and were right 20.2% of the timeJeff v1.2 0.8B alone: 363 rows stated 43% on average and were right 20.7% of the timeJeff v1.2 0.8B alone: 280 rows stated 50% on average and were right 23.6% of the timeJeff v1.2 0.8B alone: 256 rows stated 57% on average and were right 22.7% of the timeJeff v1.2 0.8B alone: 235 rows stated 63% on average and were right 31.5% of the timeJeff v1.2 0.8B alone: 178 rows stated 70% on average and were right 35.4% of the timeJeff v1.2 0.8B alone: 129 rows stated 77% on average and were right 33.3% of the timeJeff v1.2 0.8B alone: 91 rows stated 83% on average and were right 40.7% of the timeJeff v1.2 0.8B alone: 49 rows stated 90% on average and were right 61.2% of the timeJeff v1.2 0.8B alone: 24 rows stated 96% on average and were right 62.5% of the timeJeff v1.2 0.8B + adapter: 2 rows stated 25% on average and were right 0.0% of the timeJeff v1.2 0.8B + adapter: 3 rows stated 31% on average and were right 33.3% of the timeJeff v1.2 0.8B + adapter: 5 rows stated 37% on average and were right 40.0% of the timeJeff v1.2 0.8B + adapter: 6 rows stated 44% on average and were right 33.3% of the timeJeff v1.2 0.8B + adapter: 27 rows stated 51% on average and were right 37.0% of the timeJeff v1.2 0.8B + adapter: 20 rows stated 56% on average and were right 45.0% of the timeJeff v1.2 0.8B + adapter: 32 rows stated 63% on average and were right 62.5% of the timeJeff v1.2 0.8B + adapter: 24 rows stated 70% on average and were right 75.0% of the timeJeff v1.2 0.8B + adapter: 28 rows stated 77% on average and were right 71.4% of the timeJeff v1.2 0.8B + adapter: 44 rows stated 83% on average and were right 84.1% of the timeJeff v1.2 0.8B + adapter: 85 rows stated 91% on average and were right 85.9% of the timeJeff v1.2 0.8B + adapter: 3024 rows stated 100% on average and were right 99.5% of the time

Point at a dot for its numbers. Bigger dots hold more rows.

Calibration scores
ModelCalibration error (ECE)Brier scoreLog loss
Qwen3.5-0.8B untrained0.0380.9473.691
Jeff v1.2 0.8B alone0.1950.9312.977
Jeff v1.2 0.8B + adapter0.0050.0440.099

All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.

Where it gets things wrong

  • none_of_these read as a listed option: 50 rows
  • a listed option read as another listed option: 21 rows
  • a listed option read as none_of_these: 9 rows
  • nothing_matches read as q_question: 1 rows
  • none_of_these read as opt_62: 1 rows

The five commonest mistakes with the adapter, out of 3,300 test rows.

Accuracy per right answer
Right answerTest rowsAccuracy with the adapter
a listed option2,15498.6%
none_of_these55789.6%
ask_question318100.0%
none683.3%
no_match580.0%
not_listed5100.0%
question5100.0%
opt_55100.0%
item_254100.0%
nothing_matches475.0%
Show all 194 answers
Accuracy per right answer, all answers
Right answerTest rowsAccuracy with the adapter
a listed option2,15498.6%
none_of_these55789.6%
ask_question318100.0%
none683.3%
no_match580.0%
not_listed5100.0%
question5100.0%
opt_55100.0%
item_254100.0%
nothing_matches475.0%
q_question4100.0%
opt_43100.0%
opt_313100.0%
item_363100.0%
item_613100.0%
item_193100.0%
opt_23100.0%
item_213100.0%
opt_33100.0%
opt_41366.7%
opt_273100.0%
is_question3100.0%
opt_13100.0%
item_32100.0%
item_392100.0%
schedule_board2100.0%
opt_522100.0%
opt_92100.0%
opt_292100.0%
opt_832100.0%
item_742100.0%
item_262100.0%
opt_19250.0%
item_332100.0%
opt_422100.0%
item_302100.0%
item_462100.0%
opt_62100.0%
item_122100.0%
cancel2100.0%
item_382100.0%
log_reading250.0%
item_72100.0%
item_272100.0%
item_42100.0%
item_642100.0%
item_222100.0%
back2100.0%
opt_232100.0%
machine_latpulldown_v31100.0%
document_board_consent_resol1100.0%
item_1151100.0%
item_581100.0%
opt_341100.0%
item_181100.0%
opt_1671100.0%
opt_971100.0%
schedule_december_holiday_co1100.0%
item_171100.0%
member_lily_age_141100.0%
item_481100.0%
researcher_clara_whitfield1100.0%
user_account_dr_silas_wrenn1100.0%
opt_461100.0%
add_guest1100.0%
opt_811100.0%
opt_151100.0%
paycheck_sept_overtime_slip1100.0%
item_141100.0%
item_451100.0%
opt_131100.0%
opt_121100.0%
opt_541100.0%
item_151100.0%
download_pdf1100.0%
invoice_inv_2024_010_alder_d1100.0%
item_321100.0%
dashboard1100.0%
user_zara_lin_activity1100.0%
sign_out1100.0%
session_session_4021b1100.0%
opt_1181100.0%
item_911100.0%
vip_lounge1100.0%
opt_371100.0%
book_paper_boats_and_rain_no1100.0%
hotel_island_oasis_resort_ph1100.0%
opt_391100.0%
opt_351100.0%
opt_331100.0%
webhook_datadog_metrics_push1100.0%
archive_item_homestead_deed_1100.0%
opt_781100.0%
opt_741100.0%
item_781100.0%
item_351100.0%
message_template_check_in_in_21100.0%
opt_731100.0%
opt_671100.0%
property_lantern_house1100.0%
export_pdf1100.0%
opt_1801100.0%
mood_board_oceanic_blues_col1100.0%
opt_1121100.0%
remove_from_shelf1100.0%
member_elias_thorne1100.0%
book_paper_boats_and_rain_re10.0%
throttle_stream1100.0%
bookmark_letter_from_aunt_ma10.0%
account_seed_fund_checking1100.0%
paycheck_tips_transfer_aug_8_31100.0%
search1100.0%
account_restful_nights_co1100.0%
notifications1100.0%
delete_template1100.0%
opt_361100.0%
opt_111100.0%
cleaning_task_inspection_dri1100.0%
item_561100.0%
automation_rule_vip_customer1100.0%
user_account_prof_elena_voss1100.0%
my_points1100.0%
booking_com1100.0%
payment_tarjeta_visa_oro1100.0%
document_opinion_of_counsel1100.0%
employee_kevin_hart1100.0%
purge_stream1100.0%
portfolio_future_me_money_ov1100.0%
pronunciation_metric_intonat1100.0%
opt_381100.0%
asset_brand_guidelines_pdf1100.0%
new_observation1100.0%
quick_add1100.0%
item_491100.0%
project_fairview_ancestral_r_21100.0%
opt_261100.0%
opt_141100.0%
zone_sector_bravo_7_nodes1100.0%
zone_d_visitor1100.0%
automation_rule_tracking_del1100.0%
filter_specialty1100.0%
new_clause1100.0%
pick_one1100.0%
item_671100.0%
opt_691100.0%
play_digest1100.0%
everyone1100.0%
clause_governing_law1100.0%
opt_661100.0%
item_821100.0%
item_411100.0%
item_241100.0%
item_1081100.0%
clients1100.0%
opt_701100.0%
application_visa_app_mexico_1100.0%
opt_511100.0%
team_member_kevin_doyle1100.0%
reservation_stay_july_12_pay1100.0%
item_1121100.0%
thank_you_note_mrs_hendricks1100.0%
item_51100.0%
item_3710.0%
opt_711100.0%
contract_lisbon_lease1100.0%
opt_631100.0%
portfolio_steady_eddie_plan__21100.0%
candidate_michael_o_connor_s1100.0%
property_birchwood_apartment_21100.0%
policy_vidaplena_familiar_me1100.0%
pdf_document10.0%
property_birchwould_flats1100.0%
item_1221100.0%
budget_hair_cut_budget_break1100.0%
opt_501100.0%
item_11100.0%
export_all1100.0%
cancel_order1100.0%
item_891100.0%
guest_nathan_pryce1100.0%
project_kestrelos_contributo1100.0%
message_family_introduction__21100.0%
opt_1041100.0%
opt_901100.0%
item_1031100.0%
item_91100.0%
claim_emergencia_ambulancia1100.0%
project_oakridge_migration_s1100.0%
candidate_michael_o_connor_p1100.0%
delete_receipt1100.0%
station_gatehouse_terminal1100.0%
user_account_prof_elena_voss_21100.0%
session_weekend_idle_drain1100.0%
item_401100.0%

“A listed option” pools the rows whose right answer is one of the options listed in that request (keys such as o3 or t1, whose meaning changes from row to row); “another listed option” is a different one of them.

Against Qwen3.8-27B

More accurate, 35× faster than Qwen3.8-27B alone.

Gain over Qwen3.8-27B alone: +5.7 points [+2.7, +9.0] (95% interval).

On 300 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 91.3% of the time at 7.21 s per query on average; with Jeff and this adapter answering first, it was right 97.0% at 208 ms.

This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.

Time per query and prompt length for both routes
RouteMeanMedian95th percentileTypical prompt
Qwen3.8-27B alone7.21 s6.60 s13.81 s1,077 tokens
Jeff + adapter, its own answer208 ms202 ms406 ms1,044 tokens

Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.

This task's threshold 0.00: accuracy 97.0%, 34.6× faster, 0.0% sent on to Qwen3.8-27B.

Accuracy97.0%
90.0%95.0%100.0%0.000.250.500.751.00
Speed-up34.6×
0.0×20.0×40.0×0.000.250.500.751.00
Sent on to Qwen3.8-27B0.0%
0.0%50.0%100.0%0.000.250.500.751.00

Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 300 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.

All tasks, and how this was measured

How it was trained

Training rows
79,529
Steps
1,243
Training time
144 min
Size as saved
41.5 MB

One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-nav-20261001-0352.

Source: jeff-finetunes/adapters/BASELINE.md

Data card

Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean

How the test set was held out
Requests from about 10% of the roughly 340 generated apps, never trained on.
Training data
Training data not published.
Which models made the data, counted on the 72,299 training rows
What it didModelWhere it ranTraining rows
Checked the row checkerQwen3.8-Flashhosted (Alibaba Cloud DashScope)72,299
Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. nav's rows record only the checking model; the hand-over file (READY-nav) records which models wrote each stage, listed below.
Which models did each stage, from the hand-over file
StageModels
Apps (stage 1) and brand checkQwen3.8-Flash-Next (local)
Navigation systems (stage 2)Qwen3.8-Max (hosted): 233 apps; Qwen3.8-Flash-Next (local): 107 apps
Format pools (instructions, fixed-option wordings)Qwen3.8-Flash-Next (local)
Utterances (stage 4)Qwen3.8-Max (hosted): 340 apps
Supplement utterancesQwen3.8-Max (hosted): 340 apps
Blind check (stage 6)Qwen3.8-Flash (hosted): 340 apps
Second opinion on training none-of-these and question rowsDeepSeek-V4-Flash (hosted): 340 apps
Length-balance sentences round 1 (extra.py: 3-6-word questions and remarks, long requests)Qwen3.8-Max (hosted): 340 apps
Length-balance sentences round 2 (extra2.py: 2-3-word and 6-8-word questions, 2-3-word remarks, 9-13-word requests for navigation and none-of-these rows)Qwen3.8-Max (hosted): 340 apps
Length-balance sentences round 3 (extra3.py: 3-8-word questions with a filler word)Qwen3.8-Max (hosted): 340 apps
Blind check (stage 6) of length-balance rowsround 1: Qwen3.8-Flash (hosted): 340 apps; round 2: Qwen3.8-Flash (hosted): 340 apps; round 3: Qwen3.8-Flash (hosted): 340 apps
Strict question check (every question candidate)round 1 (old and round-1 questions): DeepSeek-V4-Flash (hosted): 340 apps; round 2: DeepSeek-V4-Flash (hosted): 340 apps; round 3: DeepSeek-V4-Flash (hosted): 340 apps
Length-balance sentences rounds 4 and 5 (extra4.py, extra5.py: requests of a stated word count, written alike and made navigation or none-of-these rows at random afterwards; 2-6-word questions with articles; remarks with fillers)round 4: Qwen3.8-Max (hosted): 340 apps; round 5: Qwen3.8-Max (hosted): 340 apps
Blind check (stage 6) of rounds 4 and 5round 4: Qwen3.8-Flash (hosted): 340 apps; round 5: Qwen3.8-Flash (hosted): 340 apps
Strict question check rounds 4 and 5round 4: DeepSeek-V4-Flash (hosted): 340 apps; round 5: DeepSeek-V4-Flash (hosted): 340 apps
Second opinion on training none-of-these rows of rounds 2, 4 and 5extra2: DeepSeek-V4-Flash (hosted): 286 apps; extra4: DeepSeek-V4-Flash (hosted): 271 apps; extra5: DeepSeek-V4-Flash (hosted): 270 apps
Screens, garbling, code checkscode, no model
The terms of the hosted model providers are being checked for training and publication use.

Data and licence

The adapter is released under Apache-2.0. It was trained on:

  • Generated apps, screens and spoken commandsLicence: Released with the adapter under Apache-2.0 · Generated by Qwen3.8-Max (hosted) and Qwen3.8-Flash-Next (local), per stage; see the data card

    About 340 made-up apps and websites, their screens and items, and how people say each command aloud, written by a language model and code, and checked by a second blind pass (see the data card). No real user data. Sound-alike speech-recognition errors were made in code with the CMU Pronouncing Dictionary (BSD-2-Clause, commit 74790861f652).

Changelog

  1. 0.1.0 · 2026-10-01Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.

Comments

Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.