Skip to content
JeffHub
RetrievalOfficialReproducedVersion 0.1.0

groundPassage re-ranking and answer grounding

Picks the passage that answers a question, and checks whether an answer is supported by its sources.

Data: Public data (SQuAD 2.0, HotpotQA, FEVER) plus text generated with Qwen3.8-Max (hosted, Alibaba Cloud DashScope) and Qwen3.8-Flash-Next (local); checked by Qwen3.8-Flash and DeepSeek-V4-Flash (hosted).

Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap

Use it when

  • You run retrieval-augmented generation and want to pick the best of 5 to 40 retrieved passages, or learn that none answers the question.
  • You want to check a generated answer against its sources before showing it, and tell apart supported, partly supported, contradicted and unsupported answers.
  • You want both checks from one small model, fast enough to sit on every request.

Not a good fit when

  • You need the answer written or the unsupported claim pointed out. Jeff only chooses between options; it does not generate text.
  • Your passages are very long or very many. Training passages were 30 to 300 words, and every request stayed under 8,192 tokens.
  • You need a judgement from outside knowledge. The grounding check is trained to judge only from the sources given.
  • Your texts are mostly not in English. The training data is English.

Request format

The state is an object with these fields, in this order. Only answer changes from request to request, so it comes last and the rest can be prepared in advance.

State fieldChanges per requestWhat goes in it
sourcesNoThe source passages the answer was written from, one to six in training, each with a short label such as "[1] Title" or "Source 1".
answerYesThe answer to check.
QuestionTypeWhat it decides
groundingChoiceWhether the answer is supported by the sources, judged only from the sources.

Options: Four fixed options: supported, partly_supported, contradicted and unsupported, each with the one-line meaning shown in the example.

  • Use the instructions below word for word; the adapter was trained mostly on them. Grounding instructions, "Is the answer supported by the sources? Judge only from the sources, not from outside knowledge."
  • Keep the four grounding option keys and their texts exactly as in the example. Their order does not matter; it was shuffled in training.
  • The adapter also re-ranks passages, with a second request shape. State: one key, question, holding the user's question. Options: none (text "None of these passages answers the question") plus the passages as p1, p2, … with the passage text as the option text, optionally starting with its title in square brackets. Instructions, word for word: "Which passage answers the question? If none of them does, choose none."
  • Re-ranking was trained with 5 to 40 passages of 30 to 300 words each. The passages change on every request, so nothing can be prepared in advance for them.
  • Send one question per request; the two shapes have different state keys and are separate requests.

General rules for every request are in the request format guide.

Example

The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).

from jeff import Client
from jeff.client import choice_question

jeff = Client("http://localhost:8765", model="ground")

state = {
    "sources": "[1] Harbour Bridge\nThe Harbour Bridge opened to traffic in March 1932 after eight years of work. It carries eight lanes of road traffic and two railway lines.\n\n[2] Harbour ferries\nFerries have crossed the harbour since the 1840s. The busiest route runs from the Quay to Manly.",
    "answer": "The Harbour Bridge opened in 1934 and carries two railway lines.",
}

answers = jeff.ask(state, {
    "grounding": choice_question(
        {
            "supported": "Everything the answer claims is stated in or follows directly from the sources",
            "partly_supported": "Some claims are supported, but at least one is not in the sources",
            "contradicted": "The sources state something that conflicts with the answer",
            "unsupported": "The answer's main claim is not in the sources at all",
        },
        "Is the answer supported by the sources? Judge only from the sources, not from outside knowledge.",
    ),
})
print("grounding", answers.choice("grounding").key)

Response

{
  "model": "ground",
  "answers": {
    "grounding": {
      "type": "choice",
      "probabilities": {
        "supported": 0.00010370661189667451,
        "partly_supported": 0.0007820830910539517,
        "contradicted": 0.9989073421734564,
        "unsupported": 0.00020686812359288533
      },
      "choice": "contradicted",
      "confidence": 0.9985431228979419
    }
  },
  "usage": {
    "input_tokens": 252,
    "output_tokens": 0,
    "orders": 1
  }
}

Results

On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters

  • ground4,160 test rows
    Qwen3.5-0.8B untrained
    28.7% · 0.061
    Jeff v1.2 0.8B alone
    49.0% · 0.237
    97.0% · 0.012

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

How sure is it, and is it right?

Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.

When this adapter says it is about 99.6% sure, it is right about 99.1% of the time (3,853 test rows).

Jeff v1.2 0.8B aloneJeff v1.2 0.8B + adapterperfectly calibrated
0%0%25%25%50%50%75%75%100%100%Stated confidenceRight answersJeff v1.2 0.8B alone: 2 rows stated 18% on average and were right 0.0% of the timeJeff v1.2 0.8B alone: 13 rows stated 23% on average and were right 30.8% of the timeJeff v1.2 0.8B alone: 57 rows stated 31% on average and were right 26.3% of the timeJeff v1.2 0.8B alone: 219 rows stated 37% on average and were right 42.0% of the timeJeff v1.2 0.8B alone: 348 rows stated 43% on average and were right 52.9% of the timeJeff v1.2 0.8B alone: 386 rows stated 50% on average and were right 51.3% of the timeJeff v1.2 0.8B alone: 366 rows stated 57% on average and were right 51.9% of the timeJeff v1.2 0.8B alone: 330 rows stated 63% on average and were right 47.6% of the timeJeff v1.2 0.8B alone: 374 rows stated 70% on average and were right 46.5% of the timeJeff v1.2 0.8B alone: 436 rows stated 77% on average and were right 43.8% of the timeJeff v1.2 0.8B alone: 482 rows stated 83% on average and were right 43.2% of the timeJeff v1.2 0.8B alone: 625 rows stated 90% on average and were right 46.2% of the timeJeff v1.2 0.8B alone: 522 rows stated 96% on average and were right 64.2% of the timeJeff v1.2 0.8B + adapter: 1 rows stated 33% on average and were right 100.0% of the timeJeff v1.2 0.8B + adapter: 2 rows stated 37% on average and were right 50.0% of the timeJeff v1.2 0.8B + adapter: 2 rows stated 44% on average and were right 50.0% of the timeJeff v1.2 0.8B + adapter: 17 rows stated 51% on average and were right 64.7% of the timeJeff v1.2 0.8B + adapter: 28 rows stated 57% on average and were right 50.0% of the timeJeff v1.2 0.8B + adapter: 31 rows stated 64% on average and were right 58.1% of the timeJeff v1.2 0.8B + adapter: 28 rows stated 70% on average and were right 67.9% of the timeJeff v1.2 0.8B + adapter: 47 rows stated 77% on average and were right 83.0% of the timeJeff v1.2 0.8B + adapter: 66 rows stated 83% on average and were right 69.7% of the timeJeff v1.2 0.8B + adapter: 85 rows stated 91% on average and were right 82.4% of the timeJeff v1.2 0.8B + adapter: 3853 rows stated 100% on average and were right 99.1% of the time

Point at a dot for its numbers. Bigger dots hold more rows.

Calibration scores
ModelCalibration error (ECE)Brier scoreLog loss
Qwen3.5-0.8B untrained0.0610.8101.937
Jeff v1.2 0.8B alone0.2370.7311.391
Jeff v1.2 0.8B + adapter0.0120.0480.099

All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.

Where it gets things wrong

  • a listed option read as none: 41 rows
  • none read as a listed option: 24 rows
  • contradicted read as partly_supported: 12 rows
  • a listed option read as another listed option: 10 rows
  • supported read as contradicted: 9 rows

The five commonest mistakes with the adapter, out of 4,160 test rows.

Accuracy per right answer
Right answerTest rowsAccuracy with the adapter
a listed option1,76897.1%
supported72898.2%
partly_supported52098.1%
unsupported41698.6%
contradicted41695.4%
none31292.3%

“A listed option” pools the rows whose right answer is one of the options listed in that request (keys such as o3 or t1, whose meaning changes from row to row); “another listed option” is a different one of them.

Against Qwen3.8-27B

About equal accuracy, 20× faster than Qwen3.8-27B alone.

Gain over Qwen3.8-27B alone: −0.3 points [−3.0, +2.3] (a tie: the 95% interval includes zero).

On 300 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 96.7% of the time at 13.22 s per query on average; with Jeff and this adapter answering first, it was right 96.3% at 659 ms. That is one question in 300, within noise, and the threshold rule was fixed in advance on the calibration rows, not tuned on the test.

Below this task's threshold of 0.64, Jeff passes the query on to Qwen3.8-27B; that happened for 1.7% of the rows, and their time is counted. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.

This is the task where routing pays off. Qwen3.8-27B is strong here (96.7% against 95.7% for Jeff + adapter alone). At the published threshold, 1.7% of queries go to Qwen3.8-27B. Raising the threshold to 0.97 sends 11.0% and reaches 98.7%, beating both models on their own while staying about 7 times faster than Qwen3.8-27B alone. The threshold chart below shows the whole dial.

Time per query and prompt length for both routes
RouteMeanMedian95th percentileTypical prompt
Qwen3.8-27B alone13.22 s7.93 s43.06 s1,025 tokens
Jeff + adapter, its own answer355 ms192 ms1.27 s992 tokens

Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off; with the queries passed on to Qwen3.8-27B counted, the mean is 659 ms. Typical prompt: the median prompt length in tokens.

This task's threshold 0.64: accuracy 96.3%, 20.1× faster, 1.7% sent on to Qwen3.8-27B.

Accuracy96.3%
95.0%97.5%100.0%0.000.250.500.751.00
Speed-up20.1×
0.0×20.0×40.0×0.000.250.500.751.00
Sent on to Qwen3.8-27B1.7%
0.0%50.0%100.0%0.000.250.500.751.00

Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 300 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.

All tasks, and how this was measured

How it was trained

Training rows
55,836
Steps
873
Training time
215 min
Size as saved
41.5 MB

One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-ground-20260930-1802.

Source: jeff-finetunes/adapters/BASELINE.md

Not yet measured: SQuAD 2.0 dev, re-ranking; HotpotQA dev (distractor), re-ranking; FEVER shared-task dev, grounding.

Data card

Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean

How the test set was held out
Requests from the 10% of source documents and articles that were never trained on, for both re-ranking and grounding.
Training data
Training data not published.
Which models made the data, counted on the 50,760 training rows
What it didModelWhere it ranTraining rows
Wrote the text text_teacherQwen3.8-Maxhosted (Alibaba Cloud DashScope)16,024
Wrote the text (named in the licence field) licenseQwen3.8-Flash-Nextlocal (own hardware)12,386
Wrote the text text_teacherQwen3.8-Flash-Nextlocal (own hardware)8,835
Checked the label check_teacherQwen3.8-Flashhosted (Alibaba Cloud DashScope)50,252
Gave a second opinion on the label second_opinionDeepSeek-V4-Flashhosted (DeepSeek)9,495
Checked the label check_teacherQwen3.8-Flash-Nextlocal (own hardware)508
Counted from each row's own record of the models that made it (the field named under each job). A row counts once under every job that names a model, so the counts do not add up to the total. Rows built from SQuAD 2.0, HotpotQA and FEVER keep their public text; the jobs above show which models wrote the generated text and checked the rows.
The terms of the hosted model providers are being checked for training and publication use.

Data and licence

The adapter is released under Apache-2.0. It was trained on:

  • SQuAD 2.0 (training split)Licence: CC-BY-SA-4.0 · Not generated by a model

    Revision 3ffb306f725f. Used for re-ranking and, with generated answers, for grounding.

  • HotpotQA, distractor setting (training split)Licence: CC-BY-SA-4.0 · Not generated by a model

    Revision 1908d6afbbea.

  • FEVER (training split)Licence: CC-BY-SA-3.0 · Not generated by a model

    Official release. Annotations include Wikipedia material under the Wikipedia licence terms. Used for grounding.

  • Generated documents, answers and near-miss passagesLicence: Released with the adapter under Apache-2.0 · Generated by Qwen3.8-Max (hosted) and Qwen3.8-Flash-Next (local); see the data card

    Documents in 12 domains (such as company policies, product manuals and API documentation) with questions and answers, plus answers and near-miss passages for SQuAD and HotpotQA questions, written by language models (which ones, and for how many rows, is in the data card). Every label confirmed by a second blind pass.

Changelog

  1. 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.

Comments

Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.