Skip to content
JeffHub
LegalOfficialReproducedVersion 0.1.0

legal-clausesContract clause types

Labels a single contract provision with its clause type, such as governing law, notices or confidentiality.

Data: Public data: LEDGAR (clauses from SEC filings)

Trained on Jeff v1.2. Will be retrained on v1.3. Roadmap

Use it when

  • You have a contract split into provisions and want each one labelled with a clause type.
  • Your contracts are commercial or employment agreements written in English, like those filed with the US SEC.
  • You can use the 100 LEDGAR clause types, or a subset of them, as your options.

Not a good fit when

  • Your contracts are not in English, or follow a legal system very different from the US. All training text comes from US SEC filings.
  • You need legal advice or a judgement on whether a clause is fair or enforceable. The adapter only names the clause type.
  • You pass a whole contract at once. It was trained on one provision at a time.
  • You need your own clause types. Training used the LEDGAR labels only.

Request format

The state is an object with these fields, in this order. Only provision changes from request to request, so it comes last and the rest can be prepared in advance.

State fieldChanges per requestWhat goes in it
documentNoOne sentence about the document the provision comes from. In training this was always "A commercial contract filed with the U.S. Securities and Exchange Commission (SEC)".
provisionYesThe text of one contract provision.
QuestionTypeWhat it decides
clause_typeChoiceWhich type of clause the provision is, judged by its main topic.

Options: 100 LEDGAR clause types, each with a snake_case key such as governing_laws and a one-line description. The full list is LEDGAR in descriptions.py in the source.

  • Use the option keys and descriptions from descriptions.py; the adapter was trained on them.
  • Several clause types are close in meaning (for example assignments and assigns). Expect probability to be shared between them.
  • Use the instructions below word for word; the adapter was trained mostly on them.

General rules for every request are in the request format guide.

Example

The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).

from jeff import Client
from jeff.client import choice_question

jeff = Client("http://localhost:8765", model="legal-clauses")

state = {
    "document": "A commercial contract filed with the U.S. Securities and Exchange Commission (SEC)",
    "provision": "This Agreement shall be governed by and construed in accordance with the laws of the State of Delaware, without regard to its conflict of laws principles.",
}

answers = jeff.ask(state, {
    "clause_type": choice_question(
        {
            "governing_laws": "Governing law: which state's or country's law governs the contract.",
            "notices": "Notices: how formal notices must be given, and to which addresses.",
            "confidentiality": "Confidentiality: keeping information secret and limits on using or disclosing it.",
            "terminations": "Termination: how and when the contract or employment can be ended, and what follows.",
            "assignments": "Assignments: whether and how a party may transfer its rights or duties under the contract to someone else.",
            "severability": "Severability: if one provision is invalid, the rest of the contract remains in force.",
            "entire_agreements": "Entire agreement: the contract is the whole agreement and replaces all earlier agreements on the subject.",
            "indemnifications": "Indemnification: one party must compensate and defend the other against losses and claims.",
            "counterparts": "Counterparts: the contract may be signed in separate copies, including electronic signatures, which together form one agreement.",
            "waiver_of_jury_trials": "Waiver of jury trial: the parties give up the right to a jury trial in disputes.",
            "amendments": "Amendments: how the contract can be amended, usually only in writing signed by the parties.",
        },
        "Which type of clause is this contract provision? Choose the clause type that best matches the main topic of the provision.",
    ),
})
print("clause_type", answers.choice("clause_type").key)

Response

{
  "model": "legal-clauses",
  "answers": {
    "clause_type": {
      "type": "choice",
      "probabilities": {
        "governing_laws": 0.9995311141955281,
        "notices": 0.0001556233523719579,
        "confidentiality": 0.0000320994245441589,
        "terminations": 0.00005575684053902212,
        "assignments": 0.000055541675953795374,
        "severability": 0.000058855819563845234,
        "entire_agreements": 0.000047198962382674635,
        "indemnifications": 0.00002217638050006831,
        "counterparts": 0.000014224856910469614,
        "waiver_of_jury_trials": 0.000012494141511863846,
        "amendments": 0.000014914350194014476
      },
      "choice": "governing_laws",
      "confidence": 0.9994842256150809
    }
  },
  "usage": {
    "input_tokens": 405,
    "output_tokens": 0,
    "orders": 1
  }
}

Results

On this adapter's held-out test set, never trained on. Measured 2026-10-01. All adapters

  • legal-clauses9,895 test rows
    Qwen3.5-0.8B untrained
    12.5% · 0.094
    Jeff v1.2 0.8B alone
    66.0% · 0.032
    85.7% · 0.011

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

How sure is it, and is it right?

Jeff gives every answer a probability. Each dot is a group of test rows with similar confidence: across, how sure the model said it was; up, how often it was right. Dots on the diagonal mean the stated confidence can be taken at face value.

When this adapter says it is about 90% sure, it is right about 91% of the time (1,310 test rows).

Jeff v1.2 0.8B aloneJeff v1.2 0.8B + adapterperfectly calibrated
0%0%25%25%50%50%75%75%100%100%Stated confidenceRight answersJeff v1.2 0.8B alone: 14 rows stated 6% on average and were right 7.1% of the timeJeff v1.2 0.8B alone: 283 rows stated 11% on average and were right 15.2% of the timeJeff v1.2 0.8B alone: 469 rows stated 17% on average and were right 26.2% of the timeJeff v1.2 0.8B alone: 496 rows stated 24% on average and were right 31.3% of the timeJeff v1.2 0.8B alone: 503 rows stated 30% on average and were right 32.8% of the timeJeff v1.2 0.8B alone: 492 rows stated 37% on average and were right 46.1% of the timeJeff v1.2 0.8B alone: 581 rows stated 43% on average and were right 47.8% of the timeJeff v1.2 0.8B alone: 561 rows stated 50% on average and were right 53.1% of the timeJeff v1.2 0.8B alone: 576 rows stated 57% on average and were right 56.4% of the timeJeff v1.2 0.8B alone: 556 rows stated 63% on average and were right 61.9% of the timeJeff v1.2 0.8B alone: 616 rows stated 70% on average and were right 65.4% of the timeJeff v1.2 0.8B alone: 712 rows stated 77% on average and were right 76.5% of the timeJeff v1.2 0.8B alone: 851 rows stated 83% on average and were right 80.7% of the timeJeff v1.2 0.8B alone: 1293 rows stated 90% on average and were right 88.9% of the timeJeff v1.2 0.8B alone: 1892 rows stated 96% on average and were right 94.3% of the timeJeff v1.2 0.8B + adapter: 16 rows stated 11% on average and were right 25.0% of the timeJeff v1.2 0.8B + adapter: 26 rows stated 17% on average and were right 11.5% of the timeJeff v1.2 0.8B + adapter: 33 rows stated 24% on average and were right 27.3% of the timeJeff v1.2 0.8B + adapter: 92 rows stated 30% on average and were right 32.6% of the timeJeff v1.2 0.8B + adapter: 173 rows stated 37% on average and were right 39.3% of the timeJeff v1.2 0.8B + adapter: 245 rows stated 43% on average and were right 49.0% of the timeJeff v1.2 0.8B + adapter: 380 rows stated 50% on average and were right 57.9% of the timeJeff v1.2 0.8B + adapter: 379 rows stated 57% on average and were right 56.2% of the timeJeff v1.2 0.8B + adapter: 410 rows stated 63% on average and were right 67.1% of the timeJeff v1.2 0.8B + adapter: 444 rows stated 70% on average and were right 71.8% of the timeJeff v1.2 0.8B + adapter: 552 rows stated 77% on average and were right 76.4% of the timeJeff v1.2 0.8B + adapter: 717 rows stated 83% on average and were right 85.6% of the timeJeff v1.2 0.8B + adapter: 1310 rows stated 90% on average and were right 90.7% of the timeJeff v1.2 0.8B + adapter: 5118 rows stated 98% on average and were right 97.6% of the time

Point at a dot for its numbers. Bigger dots hold more rows.

Calibration scores
ModelCalibration error (ECE)Brier scoreLog loss
Qwen3.5-0.8B untrained0.0940.9854.487
Jeff v1.2 0.8B alone0.0320.4791.523
Jeff v1.2 0.8B + adapter0.0110.2160.541

All three: lower is better, 0 is perfect. Brier score and log loss also reward being right.

Where it gets things wrong

  • applicable_laws read as governing_laws: 42 rows
  • defined_terms read as definitions: 32 rows
  • tax_withholdings read as withholdings: 30 rows
  • integration read as entire_agreements: 27 rows
  • withholdings read as tax_withholdings: 27 rows

The five commonest mistakes with the adapter, out of 9,895 test rows.

Accuracy per right answer
Right answerTest rowsAccuracy with the adapter
governing_laws56998.9%
counterparts483100.0%
notices41897.6%
severability40398.8%
entire_agreements37599.2%
amendments22289.6%
survival21799.1%
headings21697.7%
compliance_with_laws20798.1%
general19943.2%
Show all 100 answers
Accuracy per right answer, all answers
Right answerTest rowsAccuracy with the adapter
governing_laws56998.9%
counterparts483100.0%
notices41897.6%
severability40398.8%
entire_agreements37599.2%
amendments22289.6%
survival21799.1%
headings21697.7%
compliance_with_laws20798.1%
general19943.2%
assignments19489.2%
taxes18683.3%
expenses18287.4%
terms17675.6%
waivers16574.5%
confidentiality16295.1%
indemnifications15290.8%
further_assurances15196.7%
insurances13196.9%
litigations12895.3%
payments12587.2%
binding_effects12569.6%
definitions11989.9%
use_of_proceeds11998.3%
remedies11784.6%
terminations11788.9%
base_salary11298.2%
withholdings10670.8%
waiver_of_jury_trials10598.1%
no_conflicts10292.2%
warranties9764.9%
disclosures9788.7%
successors9462.8%
representations9464.9%
adjustments8893.2%
authorizations8868.2%
fees8881.8%
change_in_control8793.1%
vesting8695.3%
no_waivers8381.9%
financial_statements8298.8%
cooperation8282.9%
benefits8172.8%
consents8170.4%
miscellaneous7846.2%
interpretations7462.2%
tax_withholdings7445.9%
effective_dates7486.5%
organizations7384.9%
participations7097.1%
brokers6998.6%
death6695.5%
closings64100.0%
intellectual_property6498.4%
solvency6498.4%
no_defaults6393.7%
subsidiaries6396.8%
capitalization6398.4%
construction6366.7%
authority6258.1%
releases6091.7%
erisa58100.0%
duties5884.5%
integration5844.8%
modifications5536.4%
applicable_laws525.8%
forfeitures5290.4%
defined_terms5234.6%
vacations5296.2%
liens4989.8%
disability4887.5%
agreements4850.0%
employment4774.5%
arbitration4695.7%
non_disparagement46100.0%
transactions_with_affiliates4597.8%
titles4562.2%
enforceability4445.5%
sales4477.3%
specific_performance4381.4%
existence4281.0%
enforcements4238.1%
publicity4195.1%
interests4185.4%
consent_to_jurisdiction3366.7%
positions3262.5%
effectiveness3164.5%
records3096.7%
jurisdictions2934.5%
submission_to_jurisdiction2850.0%
indemnity2842.9%
approvals2552.0%
anti_corruption_laws2395.7%
venues2025.0%
costs1553.3%
sanctions1478.6%
powers1338.5%
qualifications520.0%
assigns40.0%
books250.0%

Harder test

On clauses whose wording is unlike any training clause (test-novel.jsonl), the adapter is right 83.2% of the time, against 85.7% on the full test set. This is the stricter number.

  • clauses whose wording is unlike any training clause (test-novel.jsonl)7,343 test rows
    Qwen3.5-0.8B untrained
    12.0% · 0.089
    Jeff v1.2 0.8B alone
    64.5% · 0.039
    83.2% · 0.012

Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).

Against Qwen3.8-27B

More accurate, 36× faster than Qwen3.8-27B alone.

Gain over Qwen3.8-27B alone: +12.8 points [+9.2, +16.4] (95% interval).

On 500 sampled test rows on an Apple M4 Max, 128 GB: Qwen3.8-27B alone was right 75.0% of the time at 16.53 s per query on average; with Jeff and this adapter answering first, it was right 87.8% at 466 ms.

This task's threshold is 0: Jeff stays ahead of Qwen3.8-27B without passing anything on, so it answered every query itself. Each task's threshold is the fastest one that still beats Qwen3.8-27B alone by at least 1 point on that task's calibration rows.

Time per query and prompt length for both routes
RouteMeanMedian95th percentileTypical prompt
Qwen3.8-27B alone16.53 s16.33 s19.72 s2,610 tokens
Jeff + adapter, its own answer466 ms460 ms513 ms2,577 tokens

Time per query, prompt to answer. Jeff's row is its own answer, before any hand-off. Typical prompt: the median prompt length in tokens.

This task's threshold 0.00: accuracy 87.8%, 35.5× faster, 0.0% sent on to Qwen3.8-27B.

Accuracy87.8%
75.0%87.5%100.0%0.000.250.500.751.00
Speed-up35.5×
0.0×20.0×40.0×0.000.250.500.751.00
Sent on to Qwen3.8-27B0.0%
0.0%50.0%100.0%0.000.250.500.751.00

Below the threshold, Jeff passes the query on to Qwen3.8-27B. Horizontal axis: the threshold, from 0 (Jeff answers everything) to 1 (Qwen3.8-27B answers everything). The dot and the vertical line mark the this task's threshold. On this task's 500 sampled rows; speed-up is Qwen3.8-27B's mean time divided by the route's mean time. Point at a chart to read any threshold.

All tasks, and how this was measured

How it was trained

Training rows
69,809
Steps
1,091
Training time
230 min
Size as saved
41.5 MB

One pass over the data (1 epoch) on one NVIDIA RTX PRO 6000. Run 0.8b-legal-clauses-20260930-0126.

Source: jeff-finetunes/adapters/BASELINE.md

Data card

Reproduced. Re-measured by the maintainers on a fixed 300-row sample of the test set, on a different machine and software (Apple M4 Max, MLX), within about 1.5 points of the full-test-set result. What the levels mean

How the test set was held out
The official LEDGAR test split, as distributed in LexGLUE; never trained on.
Training data
Built from public data sets, listed under Data and licence.
The source data sets are public (listed under Data and licence). A script to rebuild our rows from them will follow.

Data and licence

The adapter is released under Apache-2.0. It was trained on:

  • LEDGAR (LexGLUE ledgar configuration)Licence: CC-BY-4.0

    Provisions are from public SEC EDGAR filings (Tuggener et al. 2020). Option descriptions were written by hand from the label names.

Changelog

  1. 0.1.0 · 2026-09-30Trained on Jeff v1.2 (LoRA rank 16, one epoch). Results on the Results page. Published on Hugging Face as v1.2, with its test and calibration sets.

Comments

Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.