
How Does Spam Filtering Actually Work?
Or: why "CONGRATULATIONS!!! YOU WON $5,000,000" isn't the clever strategy it used to be. No single clue decides the case. The system is weighing evidence.
Or: Why "CONGRATULATIONS!!! YOU WON $5,000,000" Isn't the Clever Strategy It Used to Be
Open your spam folder sometime. You'll probably find a bizarre collection of messages — fake invoices, cryptocurrency promotions, prize notifications from companies you've never dealt with, and at least one incredibly wealthy stranger who apparently needs help moving a fortune across international borders. Most of these never reached your inbox. Something looked at them first and decided: you probably don't want this.
That's a genuinely impressive decision when you think about the scale involved. Billions of legitimate emails go out every day, and many of them share the exact same characteristics spammers rely on — urgent language, links, attachments, offers, requests for action. So how does an email provider decide which message belongs in your inbox and which belongs in spam? The answer is that it doesn't rely on any one thing. Modern spam filtering is a giant evidence-gathering operation.
The Original Problem Was Much Simpler
Early spam filters worked with relatively basic rules: certain phrases, excessive punctuation, or other common spam characteristics bumped a message's spam score up. Enough suspicious characteristics, and the email landed in the junk folder. That worked reasonably well until spammers adapted. If "Viagra" got flagged, they wrote "V1agra" instead, or inserted spaces and symbols. If all-capital subjects got penalized, they stopped using all capitals. Filters improved; spammers changed their messages; filters adapted again — an arms race that's continued for decades and is why today's systems can't rely on a static list of naughty words. They need far more context.
The Message Is Only One Piece of Evidence
When an email arrives, a filtering system can examine who appears to have sent it, which servers actually handled it, whether the sending domain is properly authenticated, what links it contains, whether those destinations have bad reputations, whether millions of similar messages suddenly appeared, and how other recipients reacted to them. The words in the email still matter, but they're part of a much larger picture. "Your invoice is attached" isn't inherently suspicious. The same sentence, arriving unexpectedly while pretending to come from your bank, failing authentication checks, and linking to a freshly registered domain, is a very different story.
This is possible because email carries far more than what you see on screen. Behind the visible sender name, subject, and body sits a set of headers — technical metadata describing which mail servers handled the message, when, and how. Changing the visible "From: PayPal" line doesn't convince the receiving mail system that PayPal actually sent anything, because there's considerably more evidence underneath that a human reader never looks at.
Three Letters That Do a Lot of Quiet Work
Email authentication rests on three technologies that build on each other. SPF (Sender Policy Framework) lets a domain owner publish, in DNS, exactly which servers are authorized to send mail on its behalf — so a receiving system can check whether a message claiming to be from a given domain actually arrived from infrastructure that domain approved. DKIM (DomainKeys Identified Mail) goes further, letting the sending system attach a cryptographic signature to parts of the message using a private key, verifiable by anyone with the corresponding public key published in DNS — the same kind of hash-and-signature machinery that shows up everywhere else in security, confirming the signed content hasn't been altered and really was authorized by whoever holds that key.
DMARC (Domain-based Message Authentication, Reporting and Conformance) ties the two together. It requires that the domains verified by SPF and DKIM actually align with the domain shown to the user in the visible From address, and lets a domain owner publish a policy telling receivers what to do with messages that fail — monitor them, quarantine them, or reject them outright. Google's own guidance for email senders lays out exactly this three-part structure as a baseline expectation for anyone sending bulk mail to Gmail users today, which is a strong signal of how central this authentication stack has become to whether a message gets delivered at all.
None of this, importantly, proves a sender is trustworthy — only that they are who their technical records say they are. A scammer can register a brand-new domain, configure SPF and DKIM flawlessly, publish a clean DMARC policy, and still send something malicious from it. They're proving, quite convincingly, that they own their own criminal domain. That's why authentication functions as one signal among many rather than a final verdict — and it does nothing at all against a lookalike domain registered and configured entirely legitimately by someone hoping you simply won't notice it's the wrong one.
Reputation Is Evidence, Not a Verdict
Email providers track reputation around sending infrastructure the way lenders track credit history. A server with years of legitimate, expected mail that recipients engage with normally builds a positive track record. A server that suddenly sends two million nearly identical messages, which recipients immediately mark as spam, builds a very different one. A brand-new domain sending enormous volumes has no history yet — that doesn't make it malicious, since every legitimate sender was new once, but it does mean the system treats it more cautiously until a pattern emerges.
Recipient behavior itself becomes part of that evidence. When thousands of people receive a message and immediately hit "report spam," that's informative. When people from a different sender regularly open, reply to, and rescue messages from spam, that's informative too. This is why marking a wrongly filtered message as "Not Spam," and reporting unwanted mail instead of just deleting it, genuinely feeds back into how future messages get classified — not dramatically, not from one click, but as one more data point in a constantly learning system.
From Keyword Lists to Probability
A major historical leap in spam filtering was Bayesian filtering — a statistical approach where the system learns, from a large set of already-classified messages, how much seeing a particular word or feature shifts the probability that a new message is spam, rather than treating any single word as an automatic trigger. As Wikipedia's overview of naive Bayes spam filtering describes, this lets a system adapt: a word strongly associated with spam in general use can behave completely differently for a recipient whose legitimate mail happens to use it constantly (a bank employee seeing "mortgage" everywhere, for instance). The system evaluates combinations of evidence rather than hunting for forbidden words, which was a real improvement over rigid rule lists — and modern machine-learning systems extend the same basic philosophy across far more signals: language, sender behavior, authentication results, URL reputation, message structure, and historical patterns all at once, often detecting combinations no handwritten rule would ever describe.
Phishing Doesn't Need to Be Sloppy Anymore
Spam and phishing overlap but aren't identical — spam is unwanted bulk mail, while phishing is specifically designed to deceive you into handing over information, money, or access. For years the advice was to watch for bad grammar and spelling, but that's an increasingly unreliable signal. Scammers have translation tools, grammar checkers, stolen legitimate templates, and generative AI at their disposal, so a phishing email today can be flawlessly written. CISA's own guidance on recognizing phishing reflects this shift, pointing people toward the actual substance of a message rather than its polish: who really sent it, where a link actually leads, whether the message was expected, and what action it's pushing you to take right now.
Urgency in particular is doing a lot of the manipulative work in these messages — your account will close, your payment failed, your package can't be delivered — precisely because urgency pushes people to act before thinking it through. Filters can recognize some patterns associated with these tactics, but urgency itself can't simply be blocked, since real banks occasionally do need you to verify something quickly too. Context has to carry the weight that keyword-matching can't.
When the Attacker Doesn't Need Malware at All
The hardest cases for filters involve no malicious link, no dangerous attachment, and sometimes no forged authentication whatsoever. In Business Email Compromise (BEC), an attacker impersonates an executive, vendor, or colleague — sometimes through a genuinely compromised real account — to convince someone to wire money or change payment details. The message might simply read: "please send the updated invoice payment to our new account." According to the FBI's Internet Crime Complaint Center, BEC has remained one of the most financially devastating categories of cybercrime for years precisely because it exploits trust and routine business process rather than any technical vulnerability a filter can scan for. When an attacker compromises a real, trusted account with a real sending history, the filter has to notice behavioral anomalies instead — an unusual request, an unfamiliar link from someone who's never sent one, a sudden blast to hundreds of contacts — because every traditional authentication and reputation signal will look perfectly clean.
It's Fundamentally a Confidence Calculation
The easiest way to understand a modern spam filter is to stop picturing one big rule and start picturing evidence accumulating in both directions. A good sender reputation, a passing DKIM check, aligned DMARC, and an established relationship with the recipient all push toward the inbox. A brand-new domain, a failed authentication check, a link pointing at known malicious infrastructure, and a sudden flood of nearly identical messages all push toward spam. At some point the accumulated weight tips the decision one way or the other — and providers deliberately keep the exact thresholds and weightings secret, because publishing them would hand spammers a precise target to just barely avoid.
That secrecy doesn't stop the experimentation, though. Spammers run their own version of A/B testing — changing subject lines, domains, servers, and wording to see what slips through — and defenders adjust their models in response. It's an ongoing adversarial system, not a problem that ever gets solved once and left alone. Spam also persists for a straightforward economic reason: sending email costs almost nothing, so a campaign that only convinces 0.001% of ten million recipients still nets a hundred victims. Every defensive layer that makes delivery harder, reputations harder to build, or domains harder to reuse raises the attacker's cost — and in security, making something unprofitable is often just as effective as making it impossible.
The Bard's Take
Spam filtering used to be easy to picture: look for suspicious words, find too many, throw the message in junk. Modern filtering looks nothing like that anymore. Today's systems weigh a web of evidence — language and structure, sender and infrastructure reputation, SPF, DKIM, and DMARC results, link and attachment analysis, sending patterns, user reports, recipient relationships, and statistical and machine-learning predictions — because almost every individual clue, taken alone, also shows up constantly in completely legitimate mail. Banks really do send urgent messages. Retailers really do say "free." New companies really do use new domains for the first time.
The filter's job is to judge whether all of those pieces make sense together, while facing an adversary actively trying to produce evidence that looks legitimate. Every time filters learn a new trick, spammers search for a way around it; every time attackers find something that works, providers gather evidence and adjust. It's a quiet, constant arms race happening behind the Send button, and the most impressive part of it isn't that an occasional phishing message slips through — it's everything you never see at all, the hundreds of messages rejected, quarantined, or classified away before you ever opened your inbox to find twenty waiting and nothing more.
Sources
- Email Sender Guidelines — Gmail Help, Google
- Naive Bayes Spam Filtering — Wikipedia
- Anti-spam Protection — Microsoft Learn
- Recognize and Report Phishing — CISA