$200 fine-tune beats OpenAI's Decisions API on email attacks

For about $200 of GPU time, a small open model called kev-4b was fine-tuned until it posted the highest recall at 99% precision of any model Abnormal tested, above sonnet-5.5-max, which costs $52 per thousand emails. According to Abnormal's write-up, the same run also left the model mostly unable to tell a BEC email from a scam.
At a glance
- Abnormal, which classifies billions of emails a week, benchmarked 40 models, including TypeSafe's Jev and the OpenAI Decisions API, on 1,985 synthetic emails modeled on hard production cases.
- At 99% precision, the OpenAI Decisions API caught most of what luna-6-max does at a fraction of the cost, while Jev returned the best-calibrated probabilities of any model tested.
- A roughly $200 fine-tune lifted kev-4b's recall at 99% precision from 7% to 66%, but its attack-type macro-F1 fell from 68 to 37, and its confidence interval overlaps sonnet-5.5-max's.
If you have not followed Abnormal, it connects to Microsoft 365 and Google Workspace through their APIs and builds behavioral baselines for every user, vendor and relationship, flagging deviations instead of matching signatures. Its production models learn what normal looks like from sender history, authentication and the org directory. They only see the signals engineers give them, which hurts with business email compromise, since BEC often carries no link or attachment, only text.
Every model answered three questions about 1,985 synthetic emails
The test set contains no customer data. Each email was generated to match the class, attack type and detection signals of a real hard case from production, either an attack or a false positive that looks like one, without reusing its identifiers. Attacks make up 40% of the set, far more than in any real inbox, so Abnormal notes that high precision is easier to reach here than in production.
Every model got the same record of extracted fields, including sender history, directory matches and domain age. In one request it answered whether the email was an attack, whether it was attack, spam, graymail or safe, and, for attacks, whether it was phishing, BEC, scam, malware or recon. The field covered 34 self-hosted open decision models, two decision APIs, and sonnet-5.5 and luna-6, each at low and max reasoning effort.
At 99% precision, the OpenAI Decisions API was the standout
Abnormal first ranked models on an email score: 100 times the mean of verdict macro-F1, attack-type macro-F1 and attack recall at 95% precision. Larger models generally scored higher and the frontier LLMs led, with sonnet-5.5-max at 88.3, but each step up the cost curve cost far more than the one before. Jev and the OpenAI Decisions API both sat on the Pareto frontier.
That score is not what decides whether a model is usable. When Abnormal flags an email as an attack, it pulls the email from the inbox, so every false positive is a legitimate message the customer never sees. At 99% precision, every open model Abnormal self-hosted caught only a small fraction of attacks, and Jev stayed on the frontier on price alone.
The OpenAI Decisions API caught most of what luna-6-max does, at a fraction of the cost and latency. The most telling miss was luna-6-low, the same luna-6 model prompted as an LLM. It matched the Decisions API on average precision but fell far behind at 99% precision.
Jev's stated probabilities came closest to the truth
Security systems often act on the probability, not the label. Above one threshold an email is remediated automatically, a middle band goes to review, and the rest is delivered. Those cutoffs have to mean the same thing across tenants and model upgrades, which only works if a stated 0.9 means that about 90% of those emails really are attacks.
Calibration was measured on is_attack, using probabilities exactly as returned, with no per-deployment tuning. The decision APIs came back among the best calibrated, and Jev was the best of any model tested. The frontier LLMs landed in the middle even when the prompt asked them for honest, calibrated probabilities. The small open decision models were the worst, base kev-4b among them.
A $200 LoRA run took kev-4b from 7% to 66% recall
Abnormal then fine-tuned kev-4b itself: a single-epoch r256 LoRA run of 7,500 steps over 60K email and non-email classification examples, on L40S GPUs. The benchmark's emails were never in the training data, and the benchmark encodes its text and words its questions differently from the training set.
Recall at 99% precision rose from 7% to 66%, and calibration went from among the worst in the field to among the best. With the threshold picked on a held-out 20% split, the fine-tune still held 98.6% precision at 67.7% recall. Over 637 test attacks, its recall at 99% precision is 65.5% [49.8, 80.2], against 49.9% [45.5, 83.2] for sonnet-5.5-max.
In the same run, attack-type macro-F1 fell from 68 to 37, and the model now labels most BEC and recon emails as scams. The training mix covered the four-way verdict but never attack types. The authors suspect the fine-tune sharpened the question it was trained on and degraded the one it was not.
One forward pass brings the cost from dollars to cents per thousand emails
A decision model usually reads the email and every question in a single forward pass and returns a probability for each option without generating a token. An answer therefore costs about as much as reading the input. Self-hosted decision models only prefill, so Abnormal priced them on input tokens.
Think of a multiple-choice exam where the student ticks boxes instead of writing essays, so reading the paper is the whole job. Each question is written in plain language. A new detector can therefore start zero-shot as a written question, with no feature to engineer and no model to retrain.
Abnormal says decision models moved LLM-class judgment from dollars per thousand emails to cents, and from seconds to milliseconds. Its current online detection costs about a cent per thousand emails on a mix of CPUs and GPU accelerators. That figure leaves out ingestion, parsing, storage, the portal and offline training.
The cost charts assume every GPU is busy all the time, but mail follows the workday, EU mail has to be processed in the EU, and a GPU fleet for a 4B model on every email is hard to get. Precision is also computed as if half of all mail were attacks. In our view, losing BEC, the attack that most needs a name, is a steep price, so the fine-tune likely works best as a detector for one question.
Haiku-5.5 and the prefilter split
Abnormal wrote the post before haiku-5.5 came out and plans to evaluate it the same way, but has given no date for those results. It expects the right setup to be a mix: cheap models on everything, decision models behind prefilters, frontier LLMs on the few emails that need real reasoning, and feedback loops. How much traffic the prefilters will pass on, and whether a fine-tune can keep attack types intact, has not been disclosed.
Related stories
- OpenAI's Decisions API beats its own engine by up to 10x
- OpenAI's Decisions API picks an answer in 150 milliseconds
- Jev runs 1,000 code reviews for $0.04, scoring 98%
- Mistral Large 4 tops open models at just 15% on legal tasks
- OpenAI speeds up GPT-6 models by about 50% in ChatGPT
- Codex opens up to Kimi K3 and GLM-5.3 Flash via Baseten
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
