How Discord Spam Detection Works: From Filters to Honeypots
Written by RiskyMH
From keyword filters to honeypot traps: how eight layers of Discord spam detection work and why you should have several at once.
Discord spam is not one problem. It shows up as repeated text, as text that changes every time, as images with no text at all, as automated accounts, and as messages from real accounts that were stolen. No single detector sees all of that, so a whole toolbox of approaches grew up over time.
This guide walks through that toolbox from the simplest signal to the more specialised ones: what people see, what the message contains, how it is sent, who sent it, and eventually why they interacted with something at all. For each one I cover what it checks, what it is good at, where it stands today, where it is heading, and some well known bots that use it. The goal is to show what is possible. It is not a ranking, because the honest answer is that you will probably want several of these running at once.
Spam and its defences grew up together. Early raids were easy to catch because many accounts sent the same text. As spammers started varying text, using images and eventually compromising legitimate accounts, defenders had to look at more than the message itself.
1. People: reports and moderator pings
The first defence was a human. A member sees spam, pings the moderator role or uses a report command, and a moderator deals with it. Discord's best practices keeping server safe recommends this. Good examples include moderator pings, report commands, modmail channels and modmail bots. My own /report bot is an example of the report command style.
This layer understands context in a way no automated system does. A person can tell a scam from a joke. The limit is speed. Someone has to be awake and watching, and the spam has already been seen by the time they act. In a big server that delay is the whole problem.
Where it is now
Reports are still the safety net for everything automation misses. They also feed the platform. Discord says community reports raised its ability to identify bad actors by 1000%, as described in How We're Fighting Spammers on Discord.
Where it is heading
Reporting stays, but automation is taking the obvious cases first so moderators only see what is left.
Examples
- Discord's built-in report options for messages and DMs.
- Report command and modmail bots, which collect reports in one place for your moderators.
2. Keyword and link filters: AutoMod
Next came automatic rules. The system checks each message against a list of words, phrases or link patterns and acts on a match.
If you take one thing from this guide, make it this: before adding another moderation bot, turn on Discord's own AutoMod and set up two rules. Block invite links, and block @everyone pings.
Those two rules fit how a lot of spam works. Most spam is built around an invite link or a scam link. A mass ping is what turns a single message into a notification for the whole server. Remove both and many of the lazy attempts fail on the spot.
AutoMod is free and needs no bot to install. Discord describes it as detecting and blocking content before it is ever posted. A moderation bot generally has to see the message first and delete it afterwards.
A good starting setup:
- A custom keyword rule for invites, using wildcard patterns such as
*discord.gg/*and*discord.com/invite/*. - A custom keyword rule for
@everyoneand@here. - Exempt rules for your moderator roles and any announcement channels.
- The built-in Block Mention Spam and Block Spam Content rules switched on.
- Separately, remove the Mention Everyone permission from the default role so only trusted roles can ping everyone at all.
Filters are also the first thing spammers adapt to. Lookalike Unicode characters can slip past word matching, and an image-only message has no text to match. That is why this layer is the base of a setup and not the whole of it.
Where it is now
Invite and mass ping rules still stop a large amount of common spam for no cost, which makes them the best first step.
Where it is heading
Discord keeps adding built-in rule types, so it is worth checking what AutoMod already does before adding a bot for the same job.
Examples
3. Content matching: have we seen this before?
Filters match what you tell them to look for. Content matching goes further and keeps a growing collection of known bad things, then checks new messages against it.
For text that means known phrases and shared blocklists of scam domains. For images it means fingerprints. A normal cryptographic hash is useless here, because changing one pixel changes the whole hash. A perceptual hash (pHash) works the other way around. Similar images produce similar hashes, so a resized, recompressed or lightly cropped copy of a known scam image lands close to the original. The system compares hashes by distance and flags anything under a threshold.
This is a good fit for how image spam works. A campaign tends to reuse a small set of images thousands of times with minor edits, so one confirmed sample can protect every server that shares the database.
Where it is now
Image-only spam is the reason this layer matters. A text filter has nothing to read when the message is only a picture, and fingerprinting is one of the cheapest ways to handle it.
Where it is heading
As spammers get better at producing endless variants, near-duplicate matching has to cover a wider range of edits. I expect it to be paired more and more with classifiers rather than used alone.
Examples
- Discord's AutoMod spam content rule looks for messages that resemble reported spam.
- RaidProtect has an image analysis module called ScamLens for scam images.
4. Behaviour detection: does this look like spam?
Content detection asks what was sent. Behaviour detection asks how it was sent.
The signals include message frequency, repeated messages, how many channels an account posts in, how many accounts join in a short window, and bursts of activity across a server. Automation tends to be consistent in ways people are not. It sends fast, it sends the same thing, and it ignores what the conversation is about.
This approach does not depend on the wording of the spam, which helps because wording is cheap to change. Thresholds still need tuning. A limit strict enough to stop a raid can also catch a busy moderator or a fast chat during an event.
Where it is now
A raid of fresh accounts and a single compromised account look very different. A raid is loud. A compromised account might send only a handful of messages. The same thresholds do not fit both.
Where it is heading
More systems combine several weak signals into one score instead of relying on a single rule. A join a few seconds ago adds a little suspicion. Posting in many channels adds more. Crossing the total triggers action.
Examples
- Discord's AutoMod mention spam rule, which limits unique mentions per message.
- ProBot can flag users who send more than five messages in five seconds.
- Wick describes its anti spam as based on a heat system.
- Security Bot uses a score based anti raid system that can lock channels when a limit is reached.
5. Verification: can we stop them before they start?
Verification is preventative. Instead of detecting bad activity after it happens, it tries to make getting in more expensive.
The common tools are CAPTCHAs, button verification, role gates, restrictions on new members, account age requirements, and extra checks that only switch on when something looks suspicious. Think of it as access control. A raid that needs a hundred accounts to each solve a challenge costs far more than one that does not.
Discord has described the same idea at the platform level. In its post on fighting spam, it talked about testing a server safe mode that watches for inauthentic behaviour from new members and requires CAPTCHAs for a while when it sees it.
Where it is now
Verification is strongest against fresh automated accounts. It is weaker against a stolen account that joined and passed the gate long ago.
Where it is heading
Adaptive verification, where the gate only appears during suspicious activity, keeps the cost low for real members while still raising it for raids.
Examples
- Wick offers CAPTCHA verification for everyone or only for suspicious joins.
- Security Bot supports hCaptcha, in app captcha codes and one click verification.
- Captcha.bot is a dedicated captcha option.
6. Reputation and account signals: who is sending this?
Reputation is separate from behaviour. Behaviour looks at what an account does right now. Reputation looks at what is known about the account.
Signals include account age, moderation history, membership across servers, whether the account has been flagged elsewhere, and other trust indicators. Some systems share that information across servers so a spammer caught in one place is recognised in another.
The important point is that reputation is not identity. A compromised account can be years old with a normal history and still be malicious right now. Discord's own fighting spam article says compromised accounts cause some of the highest user impact spam because the messages come from accounts that look legitimate.
Where it is now
Reputation works well as one signal among several and badly as the only one.
Where it is heading
Shared intelligence only helps if it arrives in time. In the data I have looked at, the same spam wave can reach different servers within seconds of each other, so a shared list has almost no window to warn anyone first.
Examples
- Discord's own account level detection. In its Q2 2022 transparency report, Discord said it disabled over 27 million accounts for spam in one quarter, with 90% caught proactively.
- Report Spam signals from members, which feed that same detection.
7. Honeypots: why did they interact with this?
Honeypots take a different approach from everything above. Most systems look at what someone sends. A honeypot creates something legitimate users have no reason to touch, then watches for anyone who does.
The idea is old in security, and honeypot channels have existed on Discord for a while, but they have become more popular recently as spam increasingly targets every channel it can reach. Image-only messages made content analysis expensive, and stolen accounts made account checks unreliable. A decoy channel can sidestep all three. It does not read the message, it does not care how old the account is, and it does not need to understand the image. The interaction itself is the signal, so it detects intent instead of content.
Where it is now
When the trap is placed somewhere legitimate members have no reason to use, it can be cheap, fast and have very few false positives. For raw image spam with no caption, it works without any image parsing.
Where it is heading
Spam is slowly adapting. Some bots now pick channels more carefully, and some target voice channel chat, which older tools often ignore. Expect setups with several honeypots, a voice channel honeypot, and channels that look active. The goal is not a perfect trap but making the cheapest attacks not worth it. Honeypots also suit some servers better than others, mainly large public ones.
Examples
- RaidProtect includes a HoneyPot feature as part of its wider suite.
- Carl-bot can be set up the same way with a rule that acts on any message in one channel.
- Honeypot is a dedicated bot built around this one idea.
8. AI and machine learning: can we understand what this is?
AI and machine learning are increasingly being used for spam and scam classification, particularly where simple rules cannot keep up with changing content. Traditional filters are told what to look for, while machine learning systems learn patterns from examples — useful when spam is too varied or too new for hand-written rules.
In practice this covers text classification, image classification, and multimodal models that read the text inside an image and look at the picture together. It also covers scam detection that depends on context, such as a fake giveaway that never uses a banned word.
AI is not a shortcut, and it is easy to oversell. Running a model on every message and every image costs real money at Discord scale, and it adds delay. It can also be fooled. A spammer can crop an image, add noise, reword the text or wrap a scam in something that looks normal, and a model that was confident yesterday can be wrong today.
It is also hard to judge spam without context. A giveaway, a trade offer or a link to a download can be a scam in one server and perfectly normal in another, and a model rarely knows which server it is looking at. So the question is never simply whether something is a scam. That is why human judgement still matters, and why AI works best as one layer that flags things for the other layers or for a person, not as the final decision.
Where it is now
Discord runs its own detection at platform scale, and a growing number of newer bots are experimenting with text and image classification for scams and similar content. As an emerging layer it is less established than the others, so I have not linked specific examples here.
Where it is heading
Spam written and generated with AI will vary more, which pushes defenders toward models that generalise instead of rules that enumerate. Cheaper inference should make that more practical. It will stay an arms race though, because models can be fooled just as filters can, and the hardest calls will still need a person.
Examples
- Discord's own proactive detection, as in the transparency report above.
- Newer multimodal moderation bots that read text and images together.
Why you probably want several at once
Asking which technique is best is the wrong question. Each one is strong where another is weak.
| Layer | Asks | Best at catching | Weak spot |
|---|---|---|---|
| People | Does this look wrong? | Anything needing context | Speed and coverage |
| Filters (AutoMod) | Does it match a rule? | Invites, mass pings, known phrases | Lookalike text, images |
| Content matching | Have we seen this? | Reused spam and images | Brand new content |
| Behaviour | How was it sent? | Bursts, raids, automation | Low volume spam |
| Verification | Can they get in? | Fresh automated accounts | Accounts already inside |
| Reputation | Who sent it? | Known bad actors | Stolen accounts with good history |
| Honeypots | Why touch this? | Spam that hits every channel | Spam that avoids the trap |
| AI | What is this? | Novel or complex content | Expensive, easy to fool, lacks context |
Here is how a single wave can play out. A message with an @everyone ping and an invite link is blocked by AutoMod before it posts. A known scam image is caught by content matching. A raid of thirty fresh accounts is caught by behaviour or stopped at the verification gate. A stolen account with years of history slips past reputation, but not past a honeypot if it posts in every channel. A brand new image that nothing has seen before is where AI or a human report picks up the slack. Each layer covers what the others miss.
What a good setup looks like depends on the server, and most servers end up with more layers than this. As one example, my own server runs an AutoMod rule for @everyone pings, a way for members to report problems, a honeypot, and scam keywords added to AutoMod by hand as new ones show up. That is not meant to suggest it is this easy. Bigger or more targeted servers usually need more, and the right mix changes as the spam does.
There is no single fix
The interesting part of Discord spam detection is that the message itself is not always the best thing to inspect. Sometimes the useful signal is the frequency of messages, sometimes the account sending them, and sometimes simply the fact that someone interacted with something they had no reason to touch.
That is why spam prevention has grown into a toolbox rather than a single detector. The more different the signals are, the less likely one change from a spammer is enough to get around everything.
The strongest setups do not choose between these approaches. They combine them.
Further reading
Adjacent notes and related topics.
Why I'm Not Worried About Spam Bots Getting Smarter
Honeypot does not need to win the arms race forever. It just needs to keep making the cheapest attacks not worth it.
My Discord Account Got Hacked: How to Recover It and Lock It Down for Good
Token theft, phishing, and social engineering are how Discord accounts actually get taken over. Here is how to recover one and lock it down for good.