The Phishing Score That Governs Your Inbox
Accuracy of 99.78 percent on familiar mail and 77.8 percent on mail from anywhere else describe the same filter, and only the second number tells an MSP what its technicians will be reading.

The accuracy figure on a security vendor's slide was almost certainly measured against mail that resembles the data the model trained on, which is a different population from the mail arriving in your clients' inboxes this morning. Point the same classifier at unfamiliar traffic and the numbers move in a direction the slide rarely carries.
Measured that way, accuracy on unfamiliar mail came in at 77.8 percent, with precision for the phishing class at 0.59 and recall for the phishing class at 0.85. Precision at 0.59 puts roughly four of every ten flagged messages in the false-alarm pile, and for a service business reviewing those flags with the technicians it already has, that lands as review hours on the desk and as a call from the client whose legitimate invoice sat in quarantine..
The Familiar Test and the Real Inbox
Two numbers describe the same model, and only one of them describes your environment. One recent evaluation verified its models on two separate sets, one in-distribution and one out-of-distribution. In plain terms, one set resembled the mail the models learned on and the other came from a different population. On the in-distribution set, one of the four models reached 99.78 percent accuracy with a ROC AUC of 0.9998. That is roughly 22 points of accuracy above what the same model returned on the out-of-distribution mail. Both figures are honest measurements of one system, and the second is the one that corresponds to a working mail stream, where senders and campaigns come from whatever the market is running this month. The two conditions produce two answers, and the friendlier answer is the one that usually travels with the product.
The binding constraint was generalization: on data outside the training set, the models did not hold up. Production mail permanently sits outside that boundary, because campaign language turns over faster than any corpus gets refreshed. For an MSP that turnover is a weekly fact, since the campaigns reaching a manufacturing client and a medical practice in the same week often share nothing beyond intent, and one filter with one threshold scores both.
What the Model Actually Learned
The reason for the drop is more useful than its size, because it follows from what the model was scored against and can appear in any pipeline built the same way. In the in-distribution corpus for the study, the phishing mail was heavily related to news and media, the sort of message that borrows the shape of a familiar bulletin. The phishing mail in the second corpus was characterized by action-oriented and financially motivated language, the vocabulary of a direct financial pitch. A model trained on the first can post a near-perfect score while what it has actually learned is a register of words, and that register holds only for as long as the campaigns keep using it.
The same test conditions produced a second result worth sitting with.
One of the simpler classifiers landed at 29.6 percent accuracy on the out-of-distribution mail, with precision of 0.30 and recall of 1.0 for the phishing class. That is the arithmetic of a model resolving its own uncertainty by calling nearly every message phishing. Recall of 1.0 presents as perfect detection of phishing, and it is also what a filter reports when it flags the entire inbox, so a recall figure settles nothing until precision from the same test set sits beside it.
What Precision Costs a Small Team (the important part)
Precision of 0.59 carries a labor cost, and the arithmetic is worth doing before a contract is signed. Take a client tenant receiving four thousand messages a day as an illustration, with five percent of them flagged: that is two hundred flags a day, of which roughly eighty are legitimate mail somebody has to open and release. Across a book of clients that is not a tuning detail; it is a staffing line, and it multiplies against a technician bench sized around the ticket load it already carries.
AI-assisted classification earns trust when it makes the work more observable, which means the flag carries the reason behind it and the false positive and false negative counts stay visible to whoever is accountable for them. It turns into theater when a confident headline number stands in for a control nobody has measured under the conditions it will face. The quiet version of the failure is a review load that outruns the capacity to read it, where every flag becomes background and the flag that mattered arrives in the same pile as the eighty that did not, which leaves the security outcome roughly where it started even as detection climbs.
That cost can be estimated in advance, which is what makes it a question for the buying conversation rather than a surprise after it.
Ask for the Out-of-Distribution Score
Generalization under drift is measurable, and it belongs in the procurement file alongside uptime and support response times. The question is not how accurate the model is. It is what the model scores on mail it has never seen. The conversation about AI in small-business security tends to run on capability narratives and one headline number, which leaves a buyer making a quantitative risk decision with no quantitative input. The corpus matters as much as the score, since a test set a vendor assembled from its own detections mostly reports what that vendor already catches.
It turns into theater when a confident headline number stands in for a control nobody has measured under the conditions it will face.
At the table, that standard comes down to a short list of asks.
Four things are reasonable to require before money moves. The first is the score on a corpus the model never trained on, reported as accuracy with precision and recall for the flagged class sitting beside it. The second is the measured false positive and false negative rates in production, along with the name of the person who reads those counts each week. The third is the re-evaluation cadence against fresh mail, and what the most recent re-evaluation actually changed. The fourth is the review path a flag travels before it reaches a client, including what the reviewer sees besides a score.
A vendor who can produce those has scored the model against the conditions it will run in, and the answers give an operator something concrete to hold the tool to at renewal. A vendor who cannot has measured the easier of the two conditions, and the buyer's team absorbs the difference in review hours nobody put in the budget.
The Number Measured on Mail the Model Has Never Seen
In this setup the numbers are illustrative rather than authoritative, and two corpora with four models are an illustration of a mechanism, well short of a benchmark of anything on the market. The failure mode is general, and it shows up wherever a classifier is scored only on mail that resembles its training data. The layer is still worth buying. Classification that reads linguistic and contextual signal catches manipulation that static rules and blocklists miss, and it sits on top of multifactor authentication, patching, managed email security, and the judgment of whoever reads the flag. It is one layer in that stack. What a discipline like this changes is the evidence standard: the number that governs the real inbox is the one measured on mail the model has never seen, and an operator is entitled to that number, with the reason attached, before signing.
More from Insights

Forecast Error Follows the Operating Picture
Manufacturing forecasts at 10.2 percent error, services at 14.5. The four-point spread traces to seams between the systems a growing firm runs, and it gets paid for in margin.

Owner Dependence Is an Architecture Problem
A quoting tool that answers in ninety seconds and still routes the pricing call to your desk feeds the bottleneck faster. Owner dependence is a routing problem.

Pricing the Checks in AI Code Review
Two of the three branches in the allocation rule reward restraint at the control layer, so price each check in your AI review stack before you buy orchestration on top of it.
See sooner. Decide faster. Act with confidence.
Set the Evidence Bar Before You Sign
QortexOS the operating system for the modern MSP.