What filters catch
Several pages on this site end by telling you to report something. This page is about what can happen next: a report that comes back with no action taken, about something that is plainly manipulative to read but contains nothing an automated filter has a hook for.
It is written for two kinds of reader: someone who reported one of the tactics documented on this site and heard nothing back, and someone who moderates and wants to know why a report like that so often goes nowhere. It is not a guide to what gets past automated detection. The tactics here are already named and explained on every other page of this site, in the open, and nothing below is written to help anyone use them more effectively.
Two of the examples measured below are the kind of thing that needs more than a platform report: a stated threat, and a stated intention to publish someone's home address and workplace. If either is genuinely your situation right now, treat this page as background reading, not as the guide for what to do next. Read the red lines first — and support services lists organisations by country.
What was measured
The messages were written for this test rather than drawn from anywhere. A set sitting clearly across this site's red lines was written first: an explicit threat, a veiled threat that names a location, a stated intention to publish someone's home address and employer, and a message organising other people to go after someone's job. Alongside them sat examples of tactics this site documents elsewhere — sealioning, gaslighting, concern trolling, DARVO — and one message written as a control: plain, unambiguous swearing, with no manipulative structure at all.
Each message was run through two filters: PurgoMalum, a keyword profanity filter, which checks text against a list of terms and reports whether it altered anything, and Detoxify (toxic-bert), a machine-learning toxicity classifier built on the unitary/toxic-bert model, which scores each message across 6 categories and treats it as flagged once the highest of those scores crosses a fixed threshold of 0.5.
PurgoMalum
keyword profanity filter
Detoxify (toxic-bert)
machine-learning toxicity classifier — unitary/toxic-bert
Two filters, not a platform's moderation stack
Neither of these is what Meta, X, TikTok or Reddit actually run. Both are freely available, representative examples of one layer of moderation — the kind of filter a small site or an independent moderator can add without a budget. The systems the major platforms run themselves are proprietary, layered, and combine automated text classification with human review and behavioural and account-level signals that a text classifier never sees at all.
That sets the edge of what this measurement can support. It supports a narrow claim: automated classification of this kind keys on abusive language — most reliably profanity and slurs — and the tactics this site documents are, by design, largely not abusive language. It does not support, and is not offered as, a claim that platforms cannot detect these tactics, or that reporting them does no good. No test run against two open filters can tell you what a platform's full moderation system does, because that system is not open to testing.
What came back
The pattern was the same across every message that was not the control.
1 of the 9 example messages in this set was flagged by either tool.
None of the 4 examples written to sit across a red line — the threat, the veiled threat, the stated intention to dox, the call to organise a pile-on — were flagged by either tool.
None of the 4 tactic examples — sealioning, gaslighting, concern trolling, DARVO — were flagged by either tool.
The control was the one exception: both tools flagged it. The keyword filter altered it, and the classifier's score crossed the threshold.
The distinction the filters were drawing, in other words, was not manipulative-or-not. It was sworn-at-or-not.
“I am merely asking for a source. Are you unable to defend your statement, or simply unwilling?”
“You absolute ******* idiot, get lost.”
What this does not show
This is a demonstration, not a study, and the difference is worth spelling out rather than leaving implied.
The set has 9 example messages, each run once against 2 tools. That is not a sample size that supports a general claim about how often these filters catch or miss anything, and no claim about false-positive or false-negative rates is made here.
Every message in the set was written for this test. None is a real report, a message someone actually sent, or drawn from a sample of anything. They were written to be unambiguous on purpose, so the pattern would be legible — which is a different thing from being representative of how these tactics read across real, messier conversations.
The classifier returns a score per category and draws no line of its own; the line above which this test counted a message as flagged — 0.5 on the model's own scoring scale — is the conventional default for a classifier of this kind, and is what was sent with each request. It was not tuned to produce this result, but it is still a choice made here rather than a property of the model, and a different threshold would change which messages count as flagged, in both directions.
The measurement was taken once, on 2 August 2026. Filters are software; a later run of these same 9 messages against updated versions of either tool could return different scores.
What this means if your report went nowhere
One consequence follows from the pattern above. A tactic built from language that reads as ordinary — a repeated polite question, a claim about your memory, a claim about who started it — is unlikely to trigger an automated system on its own, however manipulative it plainly is to someone reading the whole exchange in context. A report about it is then read, if it is read closely at all, by a person weighing it against a written policy, and a single message rarely looks like much set against that policy on its own.
That is why a contemporaneous record — what was said, by whom, when, and what came immediately before it — is often the only evidence that ends up existing for a pattern like this. Not because it is the strongest kind of evidence in the abstract, but because it may be the only kind these tools were never going to produce for you.
None of this is a reason not to report. It is a reason not to be surprised when a report about a pattern like this comes back with nothing, and a reason the record you keep matters more, not less, for tactics that read as polite.
More on where to send a report and what to expect from it: How to report.