socSecurity alert triage
Applies a security team's playbook to a new alert - ignore, investigate, isolate, block or escalate.
Data: Alerts from real UNSW-NB15 flows, EMBER 2018 files and CERT insider data, with synthetic playbooks; code fixes every label, and GLM 5.3 rewords playbooks and alerts.
Licence, in plain words
The adapter: To be decided (trained on non-commercial and all-rights-reserved data).
- UNSW-NB15 (network alerts): non-commercial use only (Free use for academic research; commercial use prohibited (ReadMe.pdf, copyright Nour Moustafa))
- CERT Insider Threat Test Dataset r4.2 (insider alerts): all rights reserved; publishing anything built on it needs the owner’s permission (ExactData end-user agreement, "Copyright 2011 ExactData, LLC, All Rights Reserved"; no redistribution, derivative works only to the minimal extent necessary)
Trained on Jeff v1.3. Adapters are tied to the exact base they were trained on. What changed in v1.3
Use it when
- Your security operations team has a written playbook (rules in order, thresholds, an asset inventory, allowed actions) and you want each new alert sorted in milliseconds.
- You want the playbook's own answer with a calibrated probability, so the unsure alerts go to an analyst first.
- Your alerts are network detections, endpoint file alerts or user-behaviour alerts; the playbook and the alert can be your own wording.
Not a good fit when
- You want the model to find threats your playbook does not describe. It applies the rules you give it.
- Alerts span several events that must be read together (a scan followed by a login burst). Training alerts are single events or single working days.
- You plan commercial use or to publish data built on the insider alerts. Two sources restrict this (see Data and licence).
Request format
The state is an object with these fields, in this order. Only alert changes from request to request, so it comes last and the rest can be prepared in advance.
| State field | Changes per request | What goes in it |
|---|---|---|
organisation | No | One line on the organisation, its sector and which alert source this queue handles. |
playbook | No | The playbook: numbered rules (the first that matches decides), thresholds, severity table, asset inventory or office policy, allowed actions, and what to do when a rule calls for an action the team may not take. |
alert | Yes | The new alert as your tools report it - what fired, on which host or user, the evidence and any lookups. |
| Question | Type | What it decides |
|---|---|---|
action | Choice | Which action the playbook requires. Options: The playbook's allowed actions: |
- Offer only the actions your playbook allows; the playbook's own fallback rule says what to do when a rule calls for another.
- Use the instructions below word for word; every training row used them.
- Keep the organisation line and playbook the same across alerts, and put the alert last, so the unchanging part can be prepared in advance.
General rules for every request are in the request format guide.
Example
The same request three ways. It assumes a Jeff server on your machine with this adapter loaded (see Install).
from jeff import Client
from jeff.client import choice_question
jeff = Client("http://localhost:8765", model="soc")
state = {
"organisation": "Kestrel Stores, a retail organisation. This alert queue receives endpoint alerts on suspicious files.",
"playbook": "Kestrel Stores endpoint alert card. Covers suspicious files found on company computers. Use it for every new alert from the endpoint protection agents.\nDefinitions: The file reputation service returns one of: malicious (with the malware family), known good, unknown (the file has not been rated yet), or lookup failed. A file that ran was executed on the computer; otherwise it was only written to disk.\nAsset inventory (host, role, criticality): BKP-40 (backup server, critical); MTG-PC-34 (meeting-room computer, important); ERP-DB-19 (ERP database server, critical); PKI-04 (certificate authority server, important); DEV-WS-35 (developer workstation, standard); TEST-WEB-29 (test web server, standard); CRM-38 (CRM server, important); EXEC-LT-05 (executive laptop, important); INTRA-31 (intranet web server, important); HR-APP-37 (HR application server, important).\nSuspicious traits: A file is suspicious if it has no digital signature. Also if a section has entropy above 7.18. Also if it imports any of these: AdjustTokenPrivileges, CryptAcquireContextA, CryptEncrypt, InternetOpenUrlW, SetWindowsHookExW, URLDownloadToFileA, WinExec. Also if it imports no functions at all.\nRules in order, first match decides:\n(1) Unknown hosts: if the host is not in the asset inventory, escalate the alert.\n(2) Malicious files. Reputation is malicious? Check the family. Family is delf, downloadguide, high, kovter, vittalia or zamg? Escalate. File ran? Isolate the host. File did not run? Block the file. Critical hosts are never isolated automatically. Rule says isolate a critical host? Escalate instead.\n(3) Failed lookups: if the reputation lookup failed, escalate when the host is critical. Otherwise investigate.\n(4) Unknown files: reputation unknown. Count the suspicious traits. With 2 or more traits, block the file if it did not run. If it ran, investigate. With fewer traits, investigate when the host is critical. Otherwise ignore the alert.\n(5) Known good files: reputation known good. Ignore the alert.\nActions: This team may ignore, investigate, isolate, block or escalate alerts. If a rule calls for an action that is not allowed here, escalate instead.",
"alert": "Sunday 14:48 EDR ALERT host=CRM-38 user=f.jensen file=\"D:\\Shared\\Apps\\invoice_viewer.exe\" sha256=53eb6d2d440384ac885525f1af04c3688f7080a5f034e54325bfbeea21922e7b size=139264 first_seen=2018-11 execution=\"ran\" signed=no sections=\".text 5.80, .rdata 5.02, .data 3.98, .rsrc 7.43\" imports_total=54 notable_imports=\"none of note\" urls=0 reputation=\"malicious (family zbot)\"",
}
answers = jeff.ask(state, {
"action": choice_question(
{
"isolate": "Isolate: cut the affected internal computer off the network.",
"investigate": "Investigate: open a case for an analyst; no containment now.",
"escalate": "Escalate: hand the alert at once to the senior responder.",
"ignore": "Ignore: close the alert because no action at all is needed.",
"block": "Block: block the outside address, file or account involved.",
},
"You are the first-line analyst in a security operations centre. Apply the organisation's playbook to the new alert and choose the action the playbook requires. The playbook's rules are checked in order and the first rule that matches decides.",
),
})
print("action", answers.choice("action").key)import { Client, choiceQuestion } from '@jeff/client';
const jeff = new Client({ url: 'http://localhost:8765', model: 'soc' });
const state = {
organisation: 'Kestrel Stores, a retail organisation. This alert queue receives endpoint alerts on suspicious files.',
playbook: 'Kestrel Stores endpoint alert card. Covers suspicious files found on company computers. Use it for every new alert from the endpoint protection agents.\nDefinitions: The file reputation service returns one of: malicious (with the malware family), known good, unknown (the file has not been rated yet), or lookup failed. A file that ran was executed on the computer; otherwise it was only written to disk.\nAsset inventory (host, role, criticality): BKP-40 (backup server, critical); MTG-PC-34 (meeting-room computer, important); ERP-DB-19 (ERP database server, critical); PKI-04 (certificate authority server, important); DEV-WS-35 (developer workstation, standard); TEST-WEB-29 (test web server, standard); CRM-38 (CRM server, important); EXEC-LT-05 (executive laptop, important); INTRA-31 (intranet web server, important); HR-APP-37 (HR application server, important).\nSuspicious traits: A file is suspicious if it has no digital signature. Also if a section has entropy above 7.18. Also if it imports any of these: AdjustTokenPrivileges, CryptAcquireContextA, CryptEncrypt, InternetOpenUrlW, SetWindowsHookExW, URLDownloadToFileA, WinExec. Also if it imports no functions at all.\nRules in order, first match decides:\n(1) Unknown hosts: if the host is not in the asset inventory, escalate the alert.\n(2) Malicious files. Reputation is malicious? Check the family. Family is delf, downloadguide, high, kovter, vittalia or zamg? Escalate. File ran? Isolate the host. File did not run? Block the file. Critical hosts are never isolated automatically. Rule says isolate a critical host? Escalate instead.\n(3) Failed lookups: if the reputation lookup failed, escalate when the host is critical. Otherwise investigate.\n(4) Unknown files: reputation unknown. Count the suspicious traits. With 2 or more traits, block the file if it did not run. If it ran, investigate. With fewer traits, investigate when the host is critical. Otherwise ignore the alert.\n(5) Known good files: reputation known good. Ignore the alert.\nActions: This team may ignore, investigate, isolate, block or escalate alerts. If a rule calls for an action that is not allowed here, escalate instead.',
alert: 'Sunday 14:48 EDR ALERT host=CRM-38 user=f.jensen file="D:\\Shared\\Apps\\invoice_viewer.exe" sha256=53eb6d2d440384ac885525f1af04c3688f7080a5f034e54325bfbeea21922e7b size=139264 first_seen=2018-11 execution="ran" signed=no sections=".text 5.80, .rdata 5.02, .data 3.98, .rsrc 7.43" imports_total=54 notable_imports="none of note" urls=0 reputation="malicious (family zbot)"',
};
const answers = await jeff.ask(state, {
action: choiceQuestion(
{
isolate: 'Isolate: cut the affected internal computer off the network.',
investigate: 'Investigate: open a case for an analyst; no containment now.',
escalate: 'Escalate: hand the alert at once to the senior responder.',
ignore: 'Ignore: close the alert because no action at all is needed.',
block: 'Block: block the outside address, file or account involved.',
},
'You are the first-line analyst in a security operations centre. Apply the organisation\'s playbook to the new alert and choose the action the playbook requires. The playbook\'s rules are checked in order and the first rule that matches decides.',
),
});
console.log('action', answers.action.key);curl -s http://localhost:8765/v1/systemone \
-H 'content-type: application/json' \
-d '{
"model": "soc",
"state": {
"organisation": "Kestrel Stores, a retail organisation. This alert queue receives endpoint alerts on suspicious files.",
"playbook": "Kestrel Stores endpoint alert card. Covers suspicious files found on company computers. Use it for every new alert from the endpoint protection agents.\nDefinitions: The file reputation service returns one of: malicious (with the malware family), known good, unknown (the file has not been rated yet), or lookup failed. A file that ran was executed on the computer; otherwise it was only written to disk.\nAsset inventory (host, role, criticality): BKP-40 (backup server, critical); MTG-PC-34 (meeting-room computer, important); ERP-DB-19 (ERP database server, critical); PKI-04 (certificate authority server, important); DEV-WS-35 (developer workstation, standard); TEST-WEB-29 (test web server, standard); CRM-38 (CRM server, important); EXEC-LT-05 (executive laptop, important); INTRA-31 (intranet web server, important); HR-APP-37 (HR application server, important).\nSuspicious traits: A file is suspicious if it has no digital signature. Also if a section has entropy above 7.18. Also if it imports any of these: AdjustTokenPrivileges, CryptAcquireContextA, CryptEncrypt, InternetOpenUrlW, SetWindowsHookExW, URLDownloadToFileA, WinExec. Also if it imports no functions at all.\nRules in order, first match decides:\n(1) Unknown hosts: if the host is not in the asset inventory, escalate the alert.\n(2) Malicious files. Reputation is malicious? Check the family. Family is delf, downloadguide, high, kovter, vittalia or zamg? Escalate. File ran? Isolate the host. File did not run? Block the file. Critical hosts are never isolated automatically. Rule says isolate a critical host? Escalate instead.\n(3) Failed lookups: if the reputation lookup failed, escalate when the host is critical. Otherwise investigate.\n(4) Unknown files: reputation unknown. Count the suspicious traits. With 2 or more traits, block the file if it did not run. If it ran, investigate. With fewer traits, investigate when the host is critical. Otherwise ignore the alert.\n(5) Known good files: reputation known good. Ignore the alert.\nActions: This team may ignore, investigate, isolate, block or escalate alerts. If a rule calls for an action that is not allowed here, escalate instead.",
"alert": "Sunday 14:48 EDR ALERT host=CRM-38 user=f.jensen file=\"D:\\Shared\\Apps\\invoice_viewer.exe\" sha256=53eb6d2d440384ac885525f1af04c3688f7080a5f034e54325bfbeea21922e7b size=139264 first_seen=2018-11 execution=\"ran\" signed=no sections=\".text 5.80, .rdata 5.02, .data 3.98, .rsrc 7.43\" imports_total=54 notable_imports=\"none of note\" urls=0 reputation=\"malicious (family zbot)\""
},
"questions": {
"action": {
"type": "choice",
"instructions": "You are the first-line analyst in a security operations centre. Apply the organisation'\''s playbook to the new alert and choose the action the playbook requires. The playbook'\''s rules are checked in order and the first rule that matches decides.",
"criteria": {
"isolate": "Isolate: cut the affected internal computer off the network.",
"investigate": "Investigate: open a case for an analyst; no containment now.",
"escalate": "Escalate: hand the alert at once to the senior responder.",
"ignore": "Ignore: close the alert because no action at all is needed.",
"block": "Block: block the outside address, file or account involved."
}
}
}
}'A real response from this adapter appears here when it is released.
Results
On this adapter's held-out test set, never trained on, scored three ways on the same rows: the untrained model Jeff is built from, the Jeff v1.3 base alone, and the base with this adapter. Questions have 4 to 5 options. As of 2026-10-05. All adapters
| Test set | Test rows | Qwen3.5-0.8B untrained | Jeff base v1.3 alone | Jeff base v1.3 + adapter |
|---|---|---|---|---|
test | 4,929 | 22.1% · 0.011 | 33.8% · 0.032 | 94.1% · 0.014 |
test4,929 test rows- Qwen3.5-0.8B untrained
- 22.1% · 0.011
- Jeff base v1.3 alone
- 33.8% · 0.032
- Jeff base v1.3 + adapter
- 94.1% · 0.014
Each cell: accuracy · calibration error (ECE; lower is better, 0 is perfect).
By group
| Group | Rows | Qwen3.5-0.8B untrained | Jeff base v1.3 alone | Jeff base v1.3 + adapter |
|---|---|---|---|---|
| alert source: endpoint | 1,500 | 23.0% | 34.7% | 96.3% |
| alert source: insider | 1,429 | 22.3% | 29.7% | 87.8% |
| alert source: network | 2,000 | 21.4% | 36.2% | 97.0% |
| label: block | 673 | 34.5% | 34.5% | 95.4% |
| label: escalate | 1,085 | 32.4% | 15.9% | 95.2% |
| label: ignore | 1,165 | 6.9% | 63.8% | 92.0% |
| label: investigate | 1,237 | 18.6% | 15.2% | 92.5% |
| label: isolate | 769 | 25.6% | 43.3% | 97.3% |
With llama.cpp
The same test, through llama.cpp: the base GGUF plus this adapter's LoRA GGUF, with the temperature refitted for each format. Running Jeff with llama.cpp
| Test set | Full precision | Q8_0 | Q4_K_M |
|---|---|---|---|
test | 94.1% · 0.014 | 94.0% · 0.013 | 93.7% · 0.009 |
Source: results/sources/v1.3/new-adapters.table.json
Model card
soc: how it was made
How it was made
In short: Alerts from real UNSW-NB15 flows, EMBER 2018 files and CERT insider data, with synthetic playbooks; code fixes every label, and GLM 5.3 rewords playbooks and alerts.
- Test set: Whole playbooks held out (36 organisations never trained on), and the underlying data too - the network flows of UNSW-NB15's official testing set, EMBER's test-set files, and CERT employees who never appear in training.
- Training data: not published.
- Training mixed in a replay sample of the Jeff base model's own training data: 4,364 rows, about 10% on top of the adapter's 43,639 (a precaution; its effect has not been measured).
- Three alert sources: network (UNSW-NB15 flows; the true attack category enters as the detection), endpoint (EMBER 2018 files; the true label enters as the file reputation) and insider (CERT r4.2 user-days; the answer key chooses the scenario days).
- About 55% of rows are hard cases built on purpose - values just either side of a threshold, conflicting signals, missing information, and rules that call for an action the team may not take.
- The adapter learns to apply a stated policy to the facts shown. It does not judge on its own whether traffic, a file or an employee is malicious.
- In a blind check, GLM 5.3 answered 200 random test rows without the label and agreed on all 200.
- Training mixes in about 10% of the base model's own training data.
| What it did | Model | Where it ran | Rows |
|---|---|---|---|
Rewrote the playbook rules in one of five styles (checked by code to keep every number, name and action word) playbook_text.model | GLM 5.3 | local | 28,485 |
Rewrote a network connection summary (checked by code) flow_text | GLM 5.3 | local | 5,016 |
Rewrote an endpoint or insider alert in one of four styles (checked by code) alert_text | GLM 5.3 | local | 4,449 |
Data and licence
Adapter: To be decided (trained on non-commercial and all-rights-reserved data)
- UNSW-NB15 (network alerts)Free use for academic research; commercial use prohibited (ReadMe.pdf, copyright Nour Moustafa)Not generated by a modelMoustafa and Slay, "UNSW-NB15 - a comprehensive data set for network intrusion detection systems", MilCIS 2015.
- EMBER 2018 (endpoint alerts)MIT (data files; the code in the same repository is AGPL-3.0)Not generated by a modelAnderson and Roth, "EMBER - An Open Dataset for Training Static PE Malware Machine Learning Models", arXiv:1804.04637, 2018.
- CERT Insider Threat Test Dataset r4.2 (insider alerts)ExactData end-user agreement, "Copyright 2011 ExactData, LLC, All Rights Reserved"; no redistribution, derivative works only to the minimal extent necessaryNot generated by a modelLindauer, Insider Threat Test Dataset, Carnegie Mellon University, 2020. The data is synthetic (CERT's generator), from 2010-2011.
- Synthetic playbooks and alert contextBuilt for this adapter; terms follow the adapter's licenceGenerated by GLM 5.3 (own hardware), for the reworded texts; everything else by code285 fictional organisations with their playbooks, inventories and thresholds, and the alert context (hosts, directions, lookups), generated by code with fixed seeds. Playbooks and some alerts were reworded by GLM 5.3 and checked by code.
Check it yourself
Changelog
- 1.3.0 · 2026-10-03First release, trained on Jeff v1.3 with the live-last prompt layout (LoRA rank 16, one epoch, about 10% of the base model's own training data mixed in). Not yet on Hugging Face.
Comments
Comments open when JeffHub launches. They will live in the registry repository's GitHub Discussions, one thread per adapter; you sign in with GitHub, and JeffHub stores no accounts.
