
Sometimes you have to fight fire with fire. But when it comes to the flood of AI-generated slop and hateful content threatening the safety and value of social media platforms, adding more AI can make the problem worse. At its best, social media is a haven for people who want to share experiences and knowledge. It gets closest to this ideal when users contribute authentic, valuable content. Relying primarily on AI tools to preserve that authenticity misses what makes social media worthwhile in the first place: the people behind it.
Erroneous Erasures
In April, the moderation team of the AskHistorians subreddit noticed something alarming. Their usually busy Slack channel, which automatically receives links to moderation mail, was flooded with alerts after dozens of comments and posts dating back 10 years were automatically removed from the community. “And there was nothing we or the experts could do about it,” said Dr. Sarah Gilbert, one of the moderators.
The deletions were especially damaging because AskHistorians users treat the community as an archive of detailed answers that continue to educate people long after they are posted. Some moderators and contributors spend hours or even days researching and writing responses. Reddit’s recently revamped AI moderation tools were apparently responsible, the moderators believe. After recovering the text of some posts, they noticed all removed content linked to Rare Historical Photos, a historical image-sharing website. They think Reddit may have designated that domain as spam, causing any post using its images for illustrative purposes to be swept up in an automated purge. Reddit has not responded to a request for comment.
This incident demonstrates a core problem: automated moderation, when applied too broadly, can erase valuable information without any meaningful appeal. And it may have contributed to metrics designed to show how effective AI moderation has become.
Questionable Metrics
Reddit says that thanks to AI, it has increased enforcement actions on hate and violent content by more than 200 percent. It also claims AI has helped reduce exposure to potentially harmful content by more than 40 percent. Reddit uses large language models to catch subtle, coordinated patterns of fake behavior and artificial hype. But the AskHistorians ordeal shows that more enforcement does not necessarily mean better enforcement. A higher number of automated actions can simply mean a higher number of mistakes.
The False Positives Problem
The growth of generative AI has created new obstacles for moderation. Large language models have made spam detection harder because they can mimic real human voices. One moderator noted that in the past two to three months, their community was absolutely flooded by LLM-powered spambots. Marketing agencies are also creating social media content designed to get brands cited by generative AI chatbots. Some startups, such as ReachLLM, focus specifically on chatbot marketing and have created and moderate subreddits to boost client visibility.
These challenges have led social media companies to explore new AI-based moderation techniques. Yet many platforms have become overly reliant on tools that are quick to penalize users for innocuous content.
Recently, Discord admitted that its AI moderation system wrongfully banned about 8,400 accounts between May and early July. The AI mistakenly labeled images containing square grids, such as chessboards or spreadsheets, as CSAM and issued permanent bans. Discord says all affected accounts have since been reinstated. The company said the AI was never intended to operate without human supervision; a human employee was supposed to review flagged content before action, but a bug caused the AI to bypass that step.
The mishap highlights why human guardrails remain essential. Without meaningful oversight, an AI-based moderation system can make thousands of mistakes in a matter of weeks, with lasting consequences. Since 2025, many Facebook and Instagram users have complained about mass bans they blame on AI moderation. The lack of human moderators has fueled frustration because there is no way to speak with a company employee about the cause of the ban or how to get an account reinstated. Meta has not said whether AI is behind those bans, but the company has increasingly relied on generative-AI-based moderation in recent years.
Tumblr has also suffered from automated moderation failures. In March, a Tumblr spokesperson said the platform’s automated systems wrongfully banned fewer than 200 accounts in one afternoon. In 2025, users complained that automatic content moderation systems inaccurately flagged content as “mature,” reducing its visibility. Tumblr has not confirmed that AI caused these problems, but says it uses a mix of machine-learning classification and human moderation.
Platforms often argue that AI is needed to keep up with the sheer volume of posts. Yet the most successful moderation systems have always relied on human volunteers and specialists who understand the culture of each community. On Reddit, moderators like those at AskHistorians are unpaid and often work long hours to enforce nuanced rules. They know that context matters: a historical photo might be appropriate in one subreddit and unacceptable in another. AI models, by design, are trained on patterns, not community values.
AI moderation can save money and help remove harmful content faster. But until these systems can eliminate basic mistakes, such as labeling a checkerboard picture as CSAM, they need human oversight. A moderator from AskHistorians noted that even reporting hateful content often results in an automated response saying it is not in violation of Reddit’s rules, forcing users into an appeals process. “It’s hard to trust the numbers because it’s hard to trust the judgment of Reddit’s systems.” False positives are a huge problem on Reddit.
AI’s Biases
Typical social media AI moderation systems use machine learning classifiers to analyze posts and flag content that breaks platform rules. But machines have difficulty understanding the nuances of sarcasm, satire, and slang. Research has increasingly suggested that marginalized groups can be disproportionately affected by AI moderation. Without human oversight, AI can end up penalizing the very communities most vulnerable to the hateful content these systems are designed to combat.
Dr. Gilbert, who is also research director of Cornell’s Citizens and Technology Lab, says marginalized and vulnerable populations experience the highest rates of moderation, typically due to false positives. These are often driven by instances of counter-speech, language reclamation, and responses to hateful content. “False positives are an equity issue,” she says. “They mean that groups that are already marginalized are further silenced and censored.”
AI moderators can also make human communities less effective at policing themselves. On Reddit, for example, subreddit moderators often prefer to ban users who engage in hateful or violent rhetoric. But if AI removes that content before a human moderator sees it, the moderator loses the ability to assess whether a ban is warranted. That might protect the platform from liability, but it weakens the community’s own standards and judgment.
AI Is A Tool, Not The Solution
Moderators across platforms have repeatedly blamed the generative AI boom for a spike in content that breaks community-specific or platform rules. This is a serious problem for social media sites that rely on user contributions. Low-effort AI-generated content is changing the challenges moderation teams face, but that makes stronger approaches more necessary, not less. Machine-scale detection should be combined with human judgment and expertise.
Some platforms are beginning to give human moderators more control. Reddit recently announced it is expanding testing for Rules Hub, a suite of tools that allows human mods to choose which rules should be automatically enforced, decide what happens when a rule is triggered, preview the experience before enabling it, and review logs and insights. Reddit expects Rules Hub to eventually replace Automod, a keyword-based tool that has long been a blunt instrument for enforcement.
The push for AI moderation is driven by economics, but the cost of false positives is borne by users. Wrongful bans can cut people off from communities, destroy years of content, and create an atmosphere of fear. If platforms continue to reduce human oversight, they will not only undermine trust but also make the online world more hostile for the people they claim to protect.
AI will certainly play a role in moderating the vast amount of content generated every minute. But AI is not a substitute for human judgment. Just as social media has no value without people, content moderation cannot succeed without human judgment at the forefront.
Source:Ars Technica News
