In one minute

Layer your detectionUser reports, rules, hash matching, classifiers, language models and behavior signals each catch what the others miss. Know which layers cover each of your top harms.
Detect patterns, not messagesGrooming, raids and scams build over days or weeks. Watch the account or the relationship over time, and let the response build as signals stack: log it, add friction, then restrict and send it to a trained reviewer.
Prevent before you removeSafer defaults, rate limits and limits on new accounts stop most harm before anything needs reviewing, and most users never feel them.
Set thresholds from data, then keep checking themTest against past confirmed cases and false alarms, and revisit every quarter, because offenders learn what triggers friction.
Watch rates, not volumeWhen one category's rate climbs while volume is normal, something coordinated is usually happening.
The mistake to avoidWaiting for users to report. Children rarely report, and much of the worst harm happens where nobody else is watching.

Why it matters

A lot of harmful content is never reported. Children rarely report what happens to them. Victims of a pile-on are outnumbered. Offenders who groom or scam move their targets to another app as fast as they can, so by the time anyone reports, the harm has often moved somewhere you can't see. If user reports are your only detection, you'll find harm late, and you'll only find the kinds people choose to report.

Law sets a floor here, not a ceiling. In the US, providers that become aware of apparent child sexual abuse material must report it to NCMEC (18 U.S.C. § 2258A), but the same section says it doesn't require you to monitor users or scan content (subsection (f)). Looking for harm is a choice you make, not one the law makes for you. In the UK, Ofcom's illegal content codes of practice recommend perceptual hash matching for known child sexual abuse material on larger or higher-risk user-to-user services, including file-sharing and file-storage services at high risk of it.

Detection has a cost when it's wrong, too. A system that catches more by acting on more will push up your proactive detection rate while good users pay for it. Detection is only working if it's measured.

What good looks like

What you have, and what you can show, at each stage

Early

A founder or first safety hire covers trust and safety, usually with under a million users.

What you have

Reporting on every place users meet or post, block lists for the words and links behind your top harms, hash matching of every image and video upload against known child sexual abuse material, rate limits and limits on new accounts, and one owner for detection.

What you can show

The share of actioned violations you found before any user report, by policy area. Time from report to action.

Growing

A dedicated safety team, millions of users, and new markets or features on the way.

What you have

Classifiers for your top harms with measured precision, signals tracked per account and per relationship, responses that build as signals stack, rate-based alerts that page whoever's on call, thresholds tested on past cases with an owner and a change log, and membership of signal-sharing programs you're eligible for.

What you can show

Precision and recall per policy area and model version. The share of alerts that turned out to be real coordination. How long new kinds of abuse run before you notice.

At scale or regulated

Tens of millions of users, a heavily regulated sector, or extra duties as a very large platform under EU or UK law.

What you have

Detection of networks, not just posts; red teams probing for gaps; thresholds reviewed every quarter; signals shared with and contributed to industry programs; a privacy review for each detection system; and settings tuned for each market's rules.

What you can show

Harmful reach before action, time from first signal to protective action, the share of cases caught before the move to another app, and a record of every threshold change with the numbers behind it.

How to do it

8 steps

  1. Step 01

    Map the detection stack against your harms

    No single layer catches everything. Each one has a job and a blind spot:

    LayerWhat it catchesWhere it falls short
    User reportsWhat users see and care about, with context only they haveChildren rarely report. Report volume measures how easy reporting is as much as how much harm exists. Groups can mass-report a target.
    RulesKnown patterns: keywords, bad links, posting velocity. Fast, cheap and easy to explain.Easy to evade with misspellings, symbols and coded language.
    Hash matchingCopies of known illegal images and video, matched by digital fingerprintOnly finds material that's already known. New material needs other layers.
    ClassifiersNew content that looks like what the model was trained on, at scaleAccuracy varies by language and context, and drifts as behavior changes.
    Language modelsRules applied with context, and new policies without new training dataSlower and costlier per item. Need testing before they act (chapter 18).
    Behavior signalsAccounts, relationships and networks over time: account age, devices, who contacts whom, gifts and paymentsNeed data, engineering and a privacy review.

    If you host images or video, hash matching is usually the first proactive layer to add, because the lists already exist. A text-only product might start with rules and rate limits instead. Microsoft's PhotoDNA is free to qualified organizations for known child sexual abuse material. Members of GIFCT, the Global Internet Forum to Counter Terrorism, share hashes of terrorist and violent extremist content through its hash-sharing database. Chapter 9 covers choosing tools, including free and open-source ones.

    Rate each of your top harms against these layers with the coverage radar. Usually detection engineering owns the systems and policy owns what they're trying to find, though on a small team that may be one person. You're done when every high-risk harm has at least one layer that doesn't depend on a user report.

  2. Step 02

    Detect the pattern, not the message

    Many serious harms don't show up in a single message. They show up as a sequence:

    • Grooming: an adult account friends many younger players it has no connection to, gifts currency early, asks about parents, moves to private voice, then pushes for another app (chapter 6).
    • A raid: dozens of new accounts post the same kind of abuse in one room within minutes.
    • A scam: a new account sends first messages to many strangers, builds trust, then asks for money, gift cards or a move off the platform.

    Each step can be innocent on its own. Together they're a pattern. Reviewing messages one at a time misses most of it.

    So decide what you're watching for each harm. Sometimes it's a single item. Often it's an account, a relationship between two accounts, or a group of accounts acting together. Keep a short signal history for each one, within a retention window you've agreed with Legal, so the third signal can be read next to the first two.

  3. Step 03

    Let the response build as signals stack

    Because each signal is weak on its own, the response should get stronger as signals add up:

    SignalsResponseWho notices
    OneLog it. Nothing visible happens.Nobody
    Two or more within a short windowAdd friction: rate limits, a safety prompt, no new private channels, limits on gifts or paymentsThe user
    The high-risk step after that: a request to move off-platform, for images or for money, or a threatRestrict the account or the contact, and send the full history to a trained reviewerA reviewer

    This is the graduated approach in chapter 4 applied to detection. Act automatically where the evidence is clear and the action can be undone. Use friction and limits where it's less certain. Gather more signals where you can't support a decision yet. Send it to a person where an error would cost too much.

    When flagged cases outnumber reviewers, protective actions stay automatic and the queue is worked by risk. A restricted account waiting a day for review is a cost worth paying (chapter 7). Chapter 6 sets out this ladder in full for adult-to-minor contact.

  4. Step 04

    Prevent before you remove

    The protections most likely to work act before any harmful message is sent. They're often cheaper than detection, and most users never notice them:

    • Safer defaults. Minors' accounts are private, and adults can't message minors they aren't connected to (chapter 6).
    • Rate limits. Cap posting, messaging, invites and friend requests, with stricter caps for new accounts.
    • Limits on new accounts. No links, no messages to strangers and no payouts until an account has built some trust.
    • Friction at the risky moment. A warning when someone asks for money, gift cards or a move to another app. A prompt before sending something your filters flag.
    • Proportionate responses to spikes. When a raid hits, slow down the accounts causing it rather than shutting the space for everyone (step 6).

    Put friction where the risk is, not everywhere. A limit that hits every user will be removed at the first growth review. A limit that hits new accounts messaging strangers will survive.

    These protections are easiest to add before launch, when changing a default is a quick edit to the spec. After launch, it means taking something away from users. That's why the abuse review belongs at design review (chapter 2, chapter 15). And measure what they prevent, such as contact that never happened, not just what you removed.

  5. Step 05

    Set thresholds from data, and keep them current

    Every rule, classifier and signal ladder has a threshold. Set it from evidence, not instinct:

    1. Gather past confirmed cases and past false alarms for the harm.
    2. Replay the signals and see where each candidate threshold would have fired.
    3. Pick the point where precision holds and reviewers can keep pace. A threshold that sends ten times more cases than your team can review isn't protecting anyone.
    4. Write it down: the threshold, the numbers behind it, the owner and who signed it off. A threshold is policy (chapter 4).
    5. Revisit it every quarter. Offenders learn what triggers friction and route around it.

    Treat every change like a release. Take a baseline before the change, check again a few weeks after, and roll it back if precision or overturns move past a limit you agreed in advance (chapter 18).

    Measure precision on a regular cadence, for example weekly, by having experts re-check a random sample of what each system acted on. Estimate recall by labeling a sample of content the system didn't flag. Report both per policy area and per model version. A blended number hides failures in small, high-severity areas, and a rising proactive detection rate means nothing if precision fell at the same time.

  6. Step 06

    Use rates as baselines, and alert on the jumps

    Volume swings with your audience. A big match, a launch or a holiday can multiply activity several times over in an evening. The rate of a category, such as abuse reports per 1,000 messages or one category's share of enforcement actions, usually holds steadier. So the rate makes a better baseline than the volume.

    When volume goes up and the rate stays put, that's a busy night. When one category's rate climbs while volume is normal, something coordinated is usually happening: a raid, or a group targeting one person.

    I'd set up an alert on that rate. Something simple works to start: if a category sits well above its usual level for that hour for 15 minutes or so, page whoever's on call.

    When it fires, respond in proportion. Locking the room punishes thousands of people for what a few hundred accounts are doing. Slowing down posting for accounts less than a week old can catch most of the people causing it and leave everyone else alone.

    Judge the alert by two numbers: how fast the team acts once it fires, and how many alerts turn out to be real coordinated activity. If most turn out to be reactions to something that happened in the game or the news, the threshold is too low. For new kinds of abuse that no alert covers yet, track how long they run before you notice, and log it in every incident review (chapter 13).

  7. Step 07

    Share signals across platforms

    Harm that starts on one service often lands on another. Offenders make first contact in a game and move the child to a private messaging app. Nudify apps show the same shape: the image is made on one service, the tool is promoted on another and the harm lands on a third. No single platform sees the whole thing, so the response has to be coordinated across them.

    Programs that exist for this:

    • Lantern, run by the Tech Coalition since 2023, lets qualifying tech companies and financial institutions share signals about online child sexual exploitation and abuse. That includes content signals such as hashes and URLs, and incident signals such as attempts to move conversations with minors off-platform.
    • GIFCT's hash-sharing database lets member companies match content against hashes of known terrorist and violent extremist material. Members share hashes, not the content.
    • StopNCII.org lets adults create a hash of their intimate images on their own device, so participating platforms can find and remove matching copies. The image never leaves the device.
    • Take It Down, from NCMEC, does the same for nude or sexually explicit images taken when the person was under 18.

    By stage: an early team matches against the lists these programs provide. A growing team joins the ones it's eligible for. At scale, you contribute signals back. Even before you're eligible, design your case records so you could share well-structured signals later. If you'd rather not build detection infrastructure from scratch, ROOST, a nonprofit, publishes free, open-source tools, including the Osprey rules engine and Coop review console.

  8. Step 08

    Set privacy limits before launch

    Detection reads people's data, so decide its limits before it runs, not after a complaint:

    • Narrow scope. Detect where the risk is. For grooming, that means adult-to-minor contact, not every conversation.
    • Metadata before content. Who contacts whom, how often and from how new an account often tells you enough to add friction. Read message content only when the pattern justifies it.
    • Agreed retention. Agree with Legal how long signals and evidence are kept, and delete on schedule.
    • Regional rules. Rules on scanning private messages vary by market, especially in the EU. Check each one before launch.

    In the EU, the GDPR's principles of data minimisation and storage limitation (Article 5(1)(c) and (e)) apply to detection data like any other personal data. Processing likely to result in a high risk to people's rights and freedoms needs a data protection impact assessment before it starts (Article 35). For each detection system, write down what it reads, why, who sees the results, how long anything is kept and which markets it runs in. That record is what Privacy and Legal will ask for, and what a regulator will ask for later.

Mistakes to avoid

And what to do instead

  1. 01

    Relying on user reports

    Add at least one proactive layer for every high-risk harm, starting with hash matching if you host images or video.

  2. 02

    Reviewing messages one at a time

    Track accounts and relationships over time, and let the response build as signals stack.

  3. 03

    Treating removal as the only response

    Use friction, rate limits and feature limits where the evidence is weaker, and keep removal for when it's clear.

  4. 04

    Setting thresholds by instinct

    Test them on past confirmed cases and false alarms, write them down, and revisit them every quarter.

  5. 05

    Celebrating a rising proactive detection rate on its own

    Read it next to precision. More catches can mean more good users caught by mistake.

  6. 06

    Alerting on volume

    Alert on the rate of each category for that hour, and judge the alert by how many fires were real.

  7. 07

    Locking the room

    Slow down the accounts causing the problem instead of punishing everyone in the space.

  8. 08

    Scanning everything because you can

    Narrow the scope, use metadata first, agree retention and check each market's rules.

Start from this template

Copy it, fill it in, make it yours

Template

Detection inventory

One row per harm. Review it every quarter.

HarmLayers in placeWhat you watch (item, account, relationship, network)Threshold and ownerLast tested on past casesPrecisionPrivacy notes (what it reads, retention, markets)

Template

Signal ladder

Fill in one for each pattern-based harm, and agree it with Legal and Privacy before launch.

SignalsWindowAutomatic responseWho reviews, and how fast
OneLog onlyNo review
Two or more
The high-risk step

Template

Rate alert

One per category you alert on: the baseline rate for each hour of the week; what counts as "well above" it; how long it has to stay there; who gets paged; the first proportionate response; and, each month, the share of alerts that were real coordination.

Do it with

Free tools and metrics that go with this chapter

Workbench tool

Classifier eval

Build a labeled set of hard cases from a rule, run a classifier against it and see where it fails. Open content

Workbench tool

Coverage radar

Rate five layers of defense for eight kinds of harm, against your products' risk. Open content

Workbench tool

Incident tabletop

Rehearse The sock puppet surge (40,000 new accounts pushing the same line overnight) and The breath-hold challenge (a dangerous trend spreading on a Friday evening).

Metric

Proactive detection rate

The share of actioned violations your systems found before any user reported them. Always read it with precision.

Metric

Precision and recall by policy area

When your systems act, how often they're right, and how much they miss.

Metric

Time to detect emerging trends

How long a new kind of abuse runs before you notice it, with the gap from first signal to triage shown on its own.

Metric

Harmful reach before action

How many people see a violating item before you act on it, which links detection speed to harm.

Further reading

Steven's posts on this topic, and sources worth the time

From Steven's writing · 6 posts

  1. Commentary

    Watch the rate, not the volume

    In Stream's live sports chat data, volume swung 7.5x but the rate of racist content held steady. Alert on rate jumps to spot raids, and slow down new accounts instead of locking the room.

    Also in: 7. Standing up review operations, 13. Crisis response

  2. Essay

    Grooming is a pattern, not a message

    Responses build as signals stack on adult-to-minor contact, with thresholds tested on past cases, a queue worked by risk, privacy limits agreed up front, and four numbers that show it works.

    Also in: 6. Child safety and age assurance, 7. Standing up review operations

  3. Share

    Nudify apps: harm no single platform sees in full

    The image is made on one service, the tool promoted on another and the harm lands on a third. That's why cross-platform signal sharing through Lantern matters.

    Also in: 6. Child safety and age assurance

  4. Essay

    Games are where kids socialize now, and regulators know it

    Grooming builds over weeks and usually moves off-platform. Protections that work act earlier and limit who can reach a child. Measure prevented contact, not just removals.

    Also in: 6. Child safety and age assurance, 16. Regulation and compliance

  5. Essay

    Automation rate isn't a measure of maturity

    Automate as much as the evidence supports. Where a wrong decision can't be reversed or someone's safety is at risk, automation prepares the case and a person closes it.

    Also in: 7. Standing up review operations, 14. Moderator wellbeing, 18. AI in Trust & Safety

Show 1 more post
  1. Essay

    You can't moderate your way out of a systems problem

    Treating Trust & Safety mainly as an operations function is a mistake. Reputation, history, age and behavior signals belong in one risk model, automation needs clear limits, and safety belongs in the product architecture from the start.

    Also in: 1. What Trust & Safety is for, 8. Hiring and structuring the team, 15. Working with Product, Legal, Comms and leadership

Outside sources

Recent changes

4 changes to this chapter, newest first

The updates page has every change to the handbook, by date.

  1. AddedAlso: 1. What Trust & Safety is for, 8. Hiring and structuring the team, 15. Working with Product, Legal, Comms and leadership

    Linked Steven's post on why treating Trust & Safety as an operations function is a mistake, and added it to chapter 8's view on where the team should report.

    Read the post: You can't moderate your way out of a systems problem →

  2. RevisedWith 18 other chapters

    Practical advice that varies by platform, such as cadences, sample sizes, targets and who owns what, is now set out as options with examples, so each team can choose what fits.

  3. DraftedWith 17 other chapters

    First full drafts of the other 18 chapters, built on Steven's posts, the handbook's principles and the Workbench's open content, with every legal and factual claim checked against its source. Stories from Steven's own work come next.

  4. AddedWith 14 other chapters

    Linked the first 16 posts to the chapters they inform, and set out the ten principles behind the handbook.