← Back to blog

How Discord Spam Detection Works: From Filters to Honeypots

Written by RiskyMH

SecurityOctober 7, 202614 min read

From keyword filters to honeypot traps: how eight layers of Discord spam detection work and why you should have several at once.

Discord spam is not one problem. It shows up as repeated text, as text that changes every time, as images with no text at all, as automated accounts, and as messages from real accounts that were stolen. No single detector sees all of that, so a whole toolbox of approaches grew up over time.

This guide walks through that toolbox from the simplest signal to the more specialised ones: what people see, what the message contains, how it is sent, who sent it, and eventually why they interacted with something at all. For each one I cover what it checks, what it is good at, where it stands today, where it is heading, and some well known bots that use it. The goal is to show what is possible. It is not a ranking, because the honest answer is that you will probably want several of these running at once.

Spam and its defences grew up together. Early raids were easy to catch because many accounts sent the same text. As spammers started varying text, using images and eventually compromising legitimate accounts, defenders had to look at more than the message itself.

1. People: reports and moderator pings

The first defence was a human. A member sees spam, pings the moderator role or uses a report command, and a moderator deals with it. Discord's best practices keeping server safe recommends this. Good examples include moderator pings, report commands, modmail channels and modmail bots. My own /report bot is an example of the report command style.

This layer understands context in a way no automated system does. A person can tell a scam from a joke. The limit is speed. Someone has to be awake and watching, and the spam has already been seen by the time they act. In a big server that delay is the whole problem.

Where it is now

Reports are still the safety net for everything automation misses. They also feed the platform. Discord says community reports raised its ability to identify bad actors by 1000%, as described in How We're Fighting Spammers on Discord.

Where it is heading

Reporting stays, but automation is taking the obvious cases first so moderators only see what is left.

Examples

  • Discord's built-in report options for messages and DMs.
  • Report command and modmail bots, which collect reports in one place for your moderators.

3. Content matching: have we seen this before?

Filters match what you tell them to look for. Content matching goes further and keeps a growing collection of known bad things, then checks new messages against it.

For text that means known phrases and shared blocklists of scam domains. For images it means fingerprints. A normal cryptographic hash is useless here, because changing one pixel changes the whole hash. A perceptual hash (pHash) works the other way around. Similar images produce similar hashes, so a resized, recompressed or lightly cropped copy of a known scam image lands close to the original. The system compares hashes by distance and flags anything under a threshold.

This is a good fit for how image spam works. A campaign tends to reuse a small set of images thousands of times with minor edits, so one confirmed sample can protect every server that shares the database.

Where it is now

Image-only spam is the reason this layer matters. A text filter has nothing to read when the message is only a picture, and fingerprinting is one of the cheapest ways to handle it.

Where it is heading

As spammers get better at producing endless variants, near-duplicate matching has to cover a wider range of edits. I expect it to be paired more and more with classifiers rather than used alone.

Examples

  • Discord's AutoMod spam content rule looks for messages that resemble reported spam.
  • RaidProtect has an image analysis module called ScamLens for scam images.

4. Behaviour detection: does this look like spam?

Content detection asks what was sent. Behaviour detection asks how it was sent.

The signals include message frequency, repeated messages, how many channels an account posts in, how many accounts join in a short window, and bursts of activity across a server. Automation tends to be consistent in ways people are not. It sends fast, it sends the same thing, and it ignores what the conversation is about.

This approach does not depend on the wording of the spam, which helps because wording is cheap to change. Thresholds still need tuning. A limit strict enough to stop a raid can also catch a busy moderator or a fast chat during an event.

Where it is now

A raid of fresh accounts and a single compromised account look very different. A raid is loud. A compromised account might send only a handful of messages. The same thresholds do not fit both.

Where it is heading

More systems combine several weak signals into one score instead of relying on a single rule. A join a few seconds ago adds a little suspicion. Posting in many channels adds more. Crossing the total triggers action.

Examples

  • Discord's AutoMod mention spam rule, which limits unique mentions per message.
  • ProBot can flag users who send more than five messages in five seconds.
  • Wick describes its anti spam as based on a heat system.
  • Security Bot uses a score based anti raid system that can lock channels when a limit is reached.

5. Verification: can we stop them before they start?

Verification is preventative. Instead of detecting bad activity after it happens, it tries to make getting in more expensive.

The common tools are CAPTCHAs, button verification, role gates, restrictions on new members, account age requirements, and extra checks that only switch on when something looks suspicious. Think of it as access control. A raid that needs a hundred accounts to each solve a challenge costs far more than one that does not.

Discord has described the same idea at the platform level. In its post on fighting spam, it talked about testing a server safe mode that watches for inauthentic behaviour from new members and requires CAPTCHAs for a while when it sees it.

Where it is now

Verification is strongest against fresh automated accounts. It is weaker against a stolen account that joined and passed the gate long ago.

Where it is heading

Adaptive verification, where the gate only appears during suspicious activity, keeps the cost low for real members while still raising it for raids.

Examples

  • Wick offers CAPTCHA verification for everyone or only for suspicious joins.
  • Security Bot supports hCaptcha, in app captcha codes and one click verification.
  • Captcha.bot is a dedicated captcha option.

6. Reputation and account signals: who is sending this?

Reputation is separate from behaviour. Behaviour looks at what an account does right now. Reputation looks at what is known about the account.

Signals include account age, moderation history, membership across servers, whether the account has been flagged elsewhere, and other trust indicators. Some systems share that information across servers so a spammer caught in one place is recognised in another.

The important point is that reputation is not identity. A compromised account can be years old with a normal history and still be malicious right now. Discord's own fighting spam article says compromised accounts cause some of the highest user impact spam because the messages come from accounts that look legitimate.

Where it is now

Reputation works well as one signal among several and badly as the only one.

Where it is heading

Shared intelligence only helps if it arrives in time. In the data I have looked at, the same spam wave can reach different servers within seconds of each other, so a shared list has almost no window to warn anyone first.

Examples

  • Discord's own account level detection. In its Q2 2022 transparency report, Discord said it disabled over 27 million accounts for spam in one quarter, with 90% caught proactively.
  • Report Spam signals from members, which feed that same detection.

7. Honeypots: why did they interact with this?

Honeypots take a different approach from everything above. Most systems look at what someone sends. A honeypot creates something legitimate users have no reason to touch, then watches for anyone who does.

The idea is old in security, and honeypot channels have existed on Discord for a while, but they have become more popular recently as spam increasingly targets every channel it can reach. Image-only messages made content analysis expensive, and stolen accounts made account checks unreliable. A decoy channel can sidestep all three. It does not read the message, it does not care how old the account is, and it does not need to understand the image. The interaction itself is the signal, so it detects intent instead of content.

Where it is now

When the trap is placed somewhere legitimate members have no reason to use, it can be cheap, fast and have very few false positives. For raw image spam with no caption, it works without any image parsing.

Where it is heading

Spam is slowly adapting. Some bots now pick channels more carefully, and some target voice channel chat, which older tools often ignore. Expect setups with several honeypots, a voice channel honeypot, and channels that look active. The goal is not a perfect trap but making the cheapest attacks not worth it. Honeypots also suit some servers better than others, mainly large public ones.

Examples

  • RaidProtect includes a HoneyPot feature as part of its wider suite.
  • Carl-bot can be set up the same way with a rule that acts on any message in one channel.
  • Honeypot is a dedicated bot built around this one idea.

8. AI and machine learning: can we understand what this is?

AI and machine learning are increasingly being used for spam and scam classification, particularly where simple rules cannot keep up with changing content. Traditional filters are told what to look for, while machine learning systems learn patterns from examples — useful when spam is too varied or too new for hand-written rules.

In practice this covers text classification, image classification, and multimodal models that read the text inside an image and look at the picture together. It also covers scam detection that depends on context, such as a fake giveaway that never uses a banned word.

AI is not a shortcut, and it is easy to oversell. Running a model on every message and every image costs real money at Discord scale, and it adds delay. It can also be fooled. A spammer can crop an image, add noise, reword the text or wrap a scam in something that looks normal, and a model that was confident yesterday can be wrong today.

It is also hard to judge spam without context. A giveaway, a trade offer or a link to a download can be a scam in one server and perfectly normal in another, and a model rarely knows which server it is looking at. So the question is never simply whether something is a scam. That is why human judgement still matters, and why AI works best as one layer that flags things for the other layers or for a person, not as the final decision.

Where it is now

Discord runs its own detection at platform scale, and a growing number of newer bots are experimenting with text and image classification for scams and similar content. As an emerging layer it is less established than the others, so I have not linked specific examples here.

Where it is heading

Spam written and generated with AI will vary more, which pushes defenders toward models that generalise instead of rules that enumerate. Cheaper inference should make that more practical. It will stay an arms race though, because models can be fooled just as filters can, and the hardest calls will still need a person.

Examples

  • Discord's own proactive detection, as in the transparency report above.
  • Newer multimodal moderation bots that read text and images together.

Why you probably want several at once

Asking which technique is best is the wrong question. Each one is strong where another is weak.

LayerAsksBest at catchingWeak spot
PeopleDoes this look wrong?Anything needing contextSpeed and coverage
Filters (AutoMod)Does it match a rule?Invites, mass pings, known phrasesLookalike text, images
Content matchingHave we seen this?Reused spam and imagesBrand new content
BehaviourHow was it sent?Bursts, raids, automationLow volume spam
VerificationCan they get in?Fresh automated accountsAccounts already inside
ReputationWho sent it?Known bad actorsStolen accounts with good history
HoneypotsWhy touch this?Spam that hits every channelSpam that avoids the trap
AIWhat is this?Novel or complex contentExpensive, easy to fool, lacks context

Here is how a single wave can play out. A message with an @everyone ping and an invite link is blocked by AutoMod before it posts. A known scam image is caught by content matching. A raid of thirty fresh accounts is caught by behaviour or stopped at the verification gate. A stolen account with years of history slips past reputation, but not past a honeypot if it posts in every channel. A brand new image that nothing has seen before is where AI or a human report picks up the slack. Each layer covers what the others miss.

What a good setup looks like depends on the server, and most servers end up with more layers than this. As one example, my own server runs an AutoMod rule for @everyone pings, a way for members to report problems, a honeypot, and scam keywords added to AutoMod by hand as new ones show up. That is not meant to suggest it is this easy. Bigger or more targeted servers usually need more, and the right mix changes as the spam does.

There is no single fix

The interesting part of Discord spam detection is that the message itself is not always the best thing to inspect. Sometimes the useful signal is the frequency of messages, sometimes the account sending them, and sometimes simply the fact that someone interacted with something they had no reason to touch.

That is why spam prevention has grown into a toolbox rather than a single detector. The more different the signals are, the less likely one change from a spammer is enough to get around everything.

The strongest setups do not choose between these approaches. They combine them.

Further reading

Adjacent notes and related topics.

Looking for the core product docs? Try the documentation or the setup guide.