In one minute

Trust & Safety has four jobsKeep people from harm, make fair decisions fast, meet your legal duties, and keep users' trust that the first three are happening. Write them down before you build anything.
Draw the borders on purposeSecurity, fraud, legal, support and integrity teams all touch the same problems. Give every seam between them an owner, or the worst cases land with nobody.
Treat safety as a retention and revenue questionIn raw data, harassed users can look like your best-retained users, because the most engaged people run into the most abuse. A matched cohort shows the real cost.
Report outcomes next to operationsOperational metrics show the machine is running. Outcome metrics show it's working. Report both, and say so when they disagree.
Hold a point of viewThis handbook rests on ten principles. Use them to settle the arguments you'll have every week.
The mistake to avoidJudging the program by how much it removes. More removals can mean better detection, more harm, or more good users caught by mistake.

Why it matters

Any product where people meet, talk, trade or create will be misused by some of them. Trust & Safety (T&S) is the function that decides what's allowed, finds what isn't, acts on it, and answers for those decisions.

Without a clear purpose, the function drifts into one of two shapes. It becomes a support queue that reacts to whatever gets reported, or a removal machine judged by throughput. Neither can tell leadership whether users are safer.

The cost shows up in the business, not just in user harm. In a 2023 white paper from Take This, using survey data Nielsen collected from 2,328 players in North America, 61% said they had at least once decided not to spend money in a game because of how other players treated them, and 60% had quit a match or a game because of harassment or hate (Take This, Toxic Gamers Are Alienating Your Core Demographic).

Parts of the job are also law now. In the EU, the Digital Services Act requires hosting services to let anyone flag illegal content (Article 16) and to explain each restriction to the user with a statement of reasons (Article 17) (Regulation (EU) 2022/2065). In the UK, the Online Safety Act requires user-to-user services to assess the risk of illegal content on their service (section 9) and to let users report it easily (section 20). The question has moved from "do you have a policy?" to "can you prove it works?"

What good looks like

What you have, and what you can show, at each stage

Early

A founder or first safety hire covers trust and safety, usually with under a million users.

What you have

A one-page statement of what T&S is for and what it owns, a named owner for each seam with security, fraud, legal and support, a log of every report and decision, and one outcome number next to the activity counts.

What you can show

Who owns each kind of harm. Time to action on the most severe reports. One rough outcome number, such as a weekly labeled sample of what users see.

Growing

A dedicated safety team, millions of users, and new markets or features on the way.

What you have

A charter agreed with Product, Legal and Security, a scorecard that pairs operational and outcome metrics, safety exposure as a standard cut on the retention dashboard, and a quarterly review with leadership.

What you can show

How much of what users see breaks the rules, with a confidence interval. 7- and 30-day return for users exposed to harm against a matched group. How often appeals overturn decisions, by policy.

At scale or regulated

Tens of millions of users, a heavily regulated sector, or extra duties as a very large platform under EU or UK law.

What you have

A charter tied to your legal duties and risk assessments, a target, an owner and a trend for every metric, safety metrics as a required input to launch and business reviews, and a written record behind every threshold.

What you can show

Whether exposure to severe harm is falling, whether bad actors come back after enforcement, how early high-risk behavior is caught, and what harm costs in retention, with public reporting that matches what enforcement does.

How to do it

7 steps

  1. Step 01

    Write the job down on one page

    Most T&S teams can list what they do. Fewer can say what it's for. Start with four jobs:

    JobWhat it meansHow you'd know it's happening
    Keep people from harmPrevent, find and stop abuse between users, with the most effort on the most severe harmsExposure to severe harm goes down, and high-risk behavior is caught earlier
    Make fair decisions fastApply clear rules the same way every time, act fastest on the worst cases, and fix mistakesTime to action on severe cases, appeal overturn rate, good users wrongly actioned
    Meet your legal dutiesReport what the law says to report, keep the records, and answer regulators and law enforcementEvery duty has an owner, a process and evidence
    Keep users' trustPeople can report, get an answer, understand decisions, and feel safe enough to stayUsers who say they feel safe, and retention of people who were targeted

    The jobs sometimes pull against each other. Speed can cost fairness, and a cautious legal reading can slow protection. Decide in advance which one wins where. One rule this handbook holds to: where a wrong decision can't be reversed or someone's safety is at risk, automation can prepare the case, but a person closes it.

    Usually the T&S lead writes the page and an executive sponsor signs it, which in a small company may be a founder. It's done when a product manager, a lawyer and a support lead can each say what T&S does, and what it doesn't.

  2. Step 02

    Draw the borders with neighboring teams

    T&S shares edges with at least five other functions. The names vary by company, so describe the work, not the org chart. Where the lines fall varies too, so treat the splits below as a common starting point, not the only one:

    TeamWhat it usually ownsWhere it meets T&SWhat to write down
    SecurityAttacks on your systems and data: breaches, vulnerabilities, misuse of internal accessAccount takeover, scraping, staff looking up user dataSecurity owns the breach. T&S owns what was done to users through the stolen accounts.
    Fraud and payments riskThe company's own losses: stolen cards, chargebacks, promotion abuseScams where one user tricks another into payingWho owns a scam that starts in chat and ends in a payment
    LegalReading the law, legal process, privilegeReporting duties, law enforcement requests, regulator questionsLegal makes the legal call. T&S runs the process, including outside office hours.
    Customer supportFirst contact with usersReports that arrive as tickets, appeals, users angry after enforcementSupport routes and explains. T&S decides.
    IntegrityVaries most: often fake accounts, coordinated manipulation, spam and misinformation. At some companies it's another name for T&S.Bots, coordinated campaigns, fake engagementWhether integrity is a separate team, and who owns coordinated abuse

    The rule: every harm in your risk register (chapter 2) has exactly one owner. Shared work is fine. Shared ownership means nobody is accountable when it goes wrong.

    The T&S lead drafts the table with each partner team. It's done when each partner has agreed to it and the on-call list matches it.

  3. Step 03

    Make the safety case with your own data

    Pull raw retention data at most game studios and harassed players look like some of the best-retained users. They chat more, queue more and play longer, so they run into more abuse. A team that stops there concludes toxicity doesn't hurt retention, and the safety budget conversation ends before it starts.

    Industry surveys like the one above help, but they rarely move a budget. Your own data does, if it's cut correctly. The cut that works is a matched cohort:

    1. Take new users whose early sessions included an actioned incident: they were the target of something you confirmed and acted on.
    2. Match each one to a new user with clean sessions and similar activity. In a game, match on playtime, mode, region and platform.
    3. Compare their 7- and 30-day return.

    Without the matching, the analysis measures engagement instead of harm. Present the result as an association unless you've run a proper causal analysis.

    The same blind spot affects enforcement on high-value users. Banning a top spender hits the revenue report the next day. The players they drove away stay invisible unless someone builds the cohort.

    Numbers worth tracking, from a game:

    • Share of new players whose first five matches include an actioned incident
    • D7 and D30 return (the share still active 7 and 30 days later) for exposed new players, against the matched group
    • Voice chat opt-out rate in a player's first week

    Other products need their own cut. On a dating app, focus on users harassed in their first week, when a bad experience makes people delete the app. On a social platform, match on audience size, because larger accounts are targeted more and churn differently. The churn after toxic exposure page has the method and a starter query.

    By stage: an early team without a data team can compare the 30-day retention of users who filed a harassment report with everyone else. It's rough, but directional. A growing team builds the matched cohort once and reruns it on a regular cadence, such as every quarter. At scale, safety exposure is a standard cut on the retention dashboard, owned jointly by T&S and the data team, and reviewed with every other churn driver. Chapter 19 turns this into a budget case.

  4. Step 04

    Separate the machine from the outcome

    If your team took down 40% more harmful content last quarter than the quarter before, is that good news? You don't know yet. Detection may have improved. Harm on the platform may have grown. Or automation got more aggressive, and legitimate users are paying for it. Volume alone can't tell you which.

    Most T&S reporting measures how much work the system does, because those numbers are easy to produce and they almost always go up. But what gets measured gets resourced. A program rewarded for activity optimizes for activity.

    Operational metrics: the machine is runningOutcome metrics: it's working
    Items removedIs exposure to severe harm going down?
    Reports actionedAfter enforcement, do bad actors stop, or come back on new accounts?
    Automation rateAre we catching high-risk behavior earlier?
    Time to decisionWhat do appeals say about where our policy or automation is wrong?

    You need both columns. Operational metrics tell you where the process is breaking this week. Outcome metrics tell you whether any of it matters.

    Some numbers don't belong on an executive slide on their own. Total items removed rises with volume and with over-enforcement. Reports received measures how easy reporting is as much as how much harm exists. Automation rate counts the decisions people didn't make, not whether they were right. Accounts banned is easy to inflate with throwaway spam accounts. The running the program guide lists the rest.

  5. Step 05

    Report both, and say when they disagree

    A metric only changes behavior when a named meeting looks at it on a schedule. One rhythm that works:

    • Weekly operations review: this week's problems. Time to action, backlog, quality, reports.
    • Monthly health review: trends and owners. Detection, appeals, repeat offending, harm rates by area.
    • Quarterly executive review: outcomes, risk and investment. Exposure to harm, retention, legal readiness.

    A small team might fold the first two into one meeting, and a live product with fast-moving harms might watch some numbers daily. What matters is that every metric has a meeting that reads it.

    Put operational and outcome numbers on the same page, and read them together. When they disagree, say so. It takes discipline to walk into a business review and say, "Our numbers are up, and I'm not convinced we're safer." That's the conversation that earns credibility, sends investment to the right places, and keeps T&S from being judged like a cost center measured by throughput.

    Know the limits of each signal. Appeal overturns only measure over-enforcement, because nobody appeals the harm you missed, so pair them with a random sample of what users actually see. Measure for a while before you commit to a target, for example six to eight weeks, because early targets are guesses.

    Keep what the company says in public in line with what enforcement actually does. Public safety claims are evidence in litigation. Chapter 11 covers measurement in depth, and chapter 15 covers reporting to executives and the board.

  6. Step 06

    Use the principles to settle arguments

    This handbook takes positions. The ten principles are listed in full in the README, each with the post it comes from. Every chapter holds to them. In practice, each one answers an argument you'll have:

    The argumentThe principle that answers itWhere to go deeper
    "Removals are up 40%. Good quarter?"1. Measure impact, not activityChapter 11
    "Can we automate this whole policy area?"2. Automate as far as the evidence supports, and 3. Automation earns its scopeChapter 18, chapter 10
    "Each message looked fine on its own."4. Harm is a pattern, not a messageChapter 5
    "Should we lock the whole room?"5. Put friction where the risk isChapter 5, chapter 13
    "We banned them. Case closed?"6. A ban is one move, not a closed caseChapter 12
    "Do we really need age checks?"7. Age assurance is the foundationChapter 6
    "Product ships DMs next month. Can you look at it?"8. Get in at design review, and earn the inviteChapter 2, chapter 15
    "A regulator asks why this account wasn't restricted."9. Be able to prove it worksChapter 16
    "Finance asks what safety returns."10. Safety is a retention and revenue questionChapter 19

    Of the ten, the one I've found hardest to hold to is that automation earns its scope. The pressure to automate usually runs ahead of the evidence, so agree the rules for expanding automation before that pressure arrives (chapter 18).

    Agree the principles with your leadership early, while nothing is on fire. A principle everyone signed up to in a calm week is far easier to apply in the middle of an incident than one you have to argue for on the spot.

  7. Step 07

    Know which outcomes you're aiming at

    Every later chapter ends its "What you can show" row with numbers. They roll up to a short list of outcomes, each with a metric page in the T&S Metrics Framework:

    OutcomeMetrics that show it
    Less exposure to harmViolating-content prevalence, plus the north star for your kind of product, such as toxicity per 1,000 match-hours or the unsafe-contact rate for minors
    The worst cases handled fastestTime to action by severity (p90), time to report child sexual exploitation
    Fair decisionsAppeal overturn rate, good users wrongly actioned
    Harm caught earlierProactive detection rate, and contact prevented before it happens
    Bad actors stopRepeat-offender rate
    Users feel safe and stayUsers who feel safe, churn after toxic exposure
    Legal duties met, with evidenceStatement-of-reasons coverage, systemic-risk assessment currency

    You don't need all of them at once. An early team picks the one north star that fits its product and tracks time to action on the most severe cases. A growing team adds decision quality and repeat offending. A team at scale or under regulation tracks all seven, each with a target and an owner.

Mistakes to avoid

And what to do instead

  1. 01

    Judging the program by removals

    Volume can't tell better detection from more harm or over-enforcement. Report an outcome next to every activity number.

  2. 02

    Leaving the borders unwritten

    When two teams think they share a harm, nobody owns it. Give every harm and every cross-team request one owner, in writing.

  3. 03

    Stopping at raw retention data

    Harassed users look engaged because they are. Compare them with a matched group before you conclude anything.

  4. 04

    Treating automation rate as maturity

    It counts the decisions people didn't make, not whether they were right. Judge automation on decision accuracy, cost per decision and appeal overturn rate.

  5. 05

    Smoothing over numbers that disagree

    If activity is up and outcomes aren't, say so in the review. It's the conversation that earns credibility.

  6. 06

    Promising more in public than enforcement does

    Public safety claims are litigation evidence. Have T&S check every safety claim before Comms makes it.

  7. 07

    Running T&S only as a queue

    Waiting for reports leaves you blind to the harm nobody reports, and much of it never is. Get into design review (chapter 2) and measure what users actually see.

Start from this template

Copy it, fill it in, make it yours

Template

One-page T&S charter

Fill in each row, then have your executive sponsor, Product, Legal and Security agree to it.

SectionWhat to write
Why we existThe four jobs, in your words, and which one wins when they conflict
What we ownThe harms and decisions T&S owns end to end
What we don't ownThe neighboring work, and the team that owns it
Decisions we make aloneFor example: removing content under a written policy, suspending an account
Decisions we make with othersFor example: reporting to law enforcement with Legal, public statements with Comms
Outcomes we reportOne north star, time to action on severe cases, and one decision-quality number
Who we answer to, and how oftenThe executive sponsor, and the weekly, monthly and quarterly reviews

Template

Seams table

One row for each harm or request that crosses teams.

Harm or requestPartner teamWhat T&S doesWhat the partner doesWho decidesWho's on call
Account takeoverSecurity
Scam that ends in a paymentFraud
Law enforcement requestLegal
Appeal that arrives through supportSupport
Coordinated fake accountsIntegrity or Security

Template

Metrics pair

For every operational number you report, name the outcome it should move and the number that keeps it honest.

Operational numberOutcome it should moveRead it with
Items removedViolating-content prevalenceAppeal overturn rate
Time to decisionHarmful reach before actionQA agreement rate
Accounts bannedRepeat-offender rateAppeal overturn rate
Reports receivedUser-report rate, by reasonViolating-content prevalence
Automation ratePrecision by policy areaAppeal overturn rate

Do it with

Free tools and metrics that go with this chapter

Workbench tool

Demo company

See every tool filled in for a fictional app.

Workbench tool

Metrics framework

Pick your platform and stage, and build a scorecard that puts outcomes next to operations. Open content

Metric

Churn after toxic exposure

The retention gap between users who were targeted and matched users who weren't.

Metric

Violating-content prevalence

Out of everything people see, how much breaks your rules.

Reference

North star metrics

The outcomes to measure first, with the question each one answers.

Reference

Running the program

Which meeting looks at which metric, how to set targets, and the numbers to keep off the executive slide.

Further reading

Steven's posts on this topic, and sources worth the time

From Steven's writing · 3 posts

  1. Essay

    Harassed players look like your best-retained users

    Raw retention data hides the cost of toxicity because harassed players are the most engaged. A matched cohort shows it, and safety exposure belongs on the retention dashboard.

    Also in: 11. Measuring what matters, 19. Budgets, roadmaps and making the case

  2. Essay

    Activity is easy to measure. Impact is harder.

    A 40% jump in removals could mean better detection, more harm or over-enforcement. Report outcome metrics next to operational ones, and tell leadership when they disagree.

    Also in: 11. Measuring what matters, 19. Budgets, roadmaps and making the case

  3. Essay

    You can't moderate your way out of a systems problem

    Treating Trust & Safety mainly as an operations function is a mistake. Reputation, history, age and behavior signals belong in one risk model, automation needs clear limits, and safety belongs in the product architecture from the start.

    Also in: 5. Detection and prevention, 8. Hiring and structuring the team, 15. Working with Product, Legal, Comms and leadership

Outside sources

Recent changes

5 changes to this chapter, newest first

The updates page has every change to the handbook, by date.

  1. AddedAlso: 5. Detection and prevention, 8. Hiring and structuring the team, 15. Working with Product, Legal, Comms and leadership

    Linked Steven's post on why treating Trust & Safety as an operations function is a mistake, and added it to chapter 8's view on where the team should report.

    Read the post: You can't moderate your way out of a systems problem →

  2. RevisedWith 18 other chapters

    Practical advice that varies by platform, such as cadences, sample sizes, targets and who owns what, is now set out as options with examples, so each team can choose what fits.

  3. RevisedWith 14 other chapters

    Added Steven's own calls from an interview: where Trust & Safety should report, what to automate first, the one number to track from day one, who makes the 2am call, and more. Practical choices that vary by platform are now laid out as options.

  4. DraftedWith 17 other chapters

    First full drafts of the other 18 chapters, built on Steven's posts, the handbook's principles and the Workbench's open content, with every legal and factual claim checked against its source. Stories from Steven's own work come next.

  5. AddedWith 14 other chapters

    Linked the first 16 posts to the chapters they inform, and set out the ten principles behind the handbook.