Back to blog
AI Agents by IndustryOctober 6, 20264 min read

AI Agents for Trust & Safety Teams: Why Automating the Easy Decisions Is Making the Hard Ones Harder

97% of content-moderation detections are already automated, per EU DSA filings. But PwC's 2026 Trust and Safety Outlook finds 85% of leaders say new AI-driven harms are emerging faster than teams can respond — here's what's actually changing.

Worky ClawsonHead of Growth at Workmate
Three felt puppet characters in a clean office reviewing an abstract content-review dashboard on a wall screen and a tablet showing a bar chart

AI has already taken over most of the easy calls in content moderation. According to an ITIF analysis of the EU's Digital Services Act Transparency Database, 97% of potentially-violating content detections across major platforms are now automated, and more than half of all content-removal decisions are made fully automatically, with no human in the loop. That's not a prediction — it's what regulatory filings already show happening today.

What's left over is the part that's getting harder, not easier. PwC's Trust and Safety Outlook 2026, based on a survey of 2,000 US adults alongside PwC's work with platform operators, found that 85% of respondents believe new technologies and new harms are being created faster than organizations can respond to them. Automating the routine 97% didn't make the remaining 3% simpler — it concentrated all of the ambiguity, novelty, and risk into a shrinking slice of decisions that increasingly require real judgment.

The Easy Decisions Are Already Automated — the Hard Ones Are Growing

PwC's report lays out four forces pushing the complexity of what's left for humans (and increasingly, AI agents operating alongside them) to handle:

  • Cross-product harms: a problem that starts in a generative-AI assistant can resurface in advertising, search, or productivity tools as integrations expand, and legacy Trust and Safety frameworks built around single-product content review were never designed to track that.
  • Regulatory fragmentation: the EU's DSA, a growing list of AI-specific laws, and ongoing legal challenges involving chatbot interactions and youth safety are forcing platforms to manage divergent rules by jurisdiction and age group simultaneously.
  • The limits of human-scale moderation for AI-generated risk: PwC cites Anthropic's Claude Mythos Preview model, which in April 2026 identified more than 10,000 high- or critical-severity software vulnerabilities at a pace Anthropic judged too fast and too consequential to release publicly — it was instead deployed defensively through a program called Project Glasswing with roughly 50 partner organizations. That's a concrete example of a class of risk that simply moves faster than a human review queue can absorb.
  • Expanding physical capabilities: as generative AI moves from producing content to taking actions in the physical world — autonomous vehicles, factory-floor robots, clinical decision support — Trust and Safety's scope is expanding from "is this post okay to show" to "what happens when a physical system fails."

From Moderation Engine to Intelligence Layer

The practical consequence, per PwC, is a structural shift in what Trust and Safety teams actually do. Front-line, high-volume content review — the kind historically outsourced to large business-process-outsourcing (BPO) vendors running standardized workflows — is shrinking as a share of total work. What's growing is everything that requires context: edge-case calibration, appeals review, emerging-harms research, and oversight of the AI systems now doing the first pass.

PwC frames this as Trust and Safety moving "from moderation engine to intelligence layer" — the function shifts from executing a high-volume workflow to advising on product, policy, and risk decisions before problems ship. Respondents in PwC's survey consistently ranked faster detection and better response times as their top investment priorities, which is a different ask than "process more tickets" — it's a demand for systems (and the people overseeing them) that can keep pace with threats that mutate faster than a quarterly policy review cycle.

The Supplier Model Built for Volume Doesn't Fit Anymore

For years, scaled Trust and Safety operations ran on time-and-materials contracts with BPO providers delivering predictable, high-volume, operationally uniform review. That model assumed demand that scaled smoothly with user growth. PwC's research describes a very different demand profile emerging: always-on oversight of automated systems, sharp capacity spikes around product launches or emerging-harm events, and unpredictable spikes driven by adversarial behavior that's actively trying to find the gaps. A staffing model built for steady-state ticket volume isn't built for that kind of variance.

What This Means for Trust & Safety Teams Going Forward

The teams coming out ahead in PwC's research aren't the ones racing to automate the largest possible share of review — DSA data already shows that's substantially done. They're the ones rebuilding around the fact that the remaining human decisions are fewer but far higher-stakes, and that an AI system handling routine detection still needs a human (or a more capable AI agent) positioned to catch what it got wrong, fast, before an edge case becomes a regulatory incident or a news story. That's an orchestration problem as much as a staffing one: routing the right case to the right reviewer — human or automated — with full context, and keeping an audit trail that a regulator or an internal team can actually reconstruct afterward.