The T&S Handbook · Part 1: Before the first hire
What Trust & Safety is for
What is this function for, and how do you know it's working?
All chapters
Part 1: Before the first hire
Part 2: Build
- 4Writing policy and an enforcement ladder
- 5Detection and prevention
- 6Child safety and age assurance
- 7Standing up review operations
- 8Hiring and structuring the team
- 9Choosing vendors and tools
Part 3: Run
- 10Quality, calibration and appeals
- 11Measuring what matters
- 12Severe harm escalations
- 13Crisis response
- 14Moderator wellbeing
- 15Working with Product, Legal, Comms and leadership
Part 4: Scale and govern
General information, not legal advice. Laws differ by country, change often and apply differently to each service. Check with your own legal team or outside counsel before acting on anything here.
In one minute
Why it matters
Any product where people meet, talk, trade or create will be misused by some of them. Trust & Safety (T&S) is the function that decides what's allowed, finds what isn't, acts on it, and answers for those decisions.
Without a clear purpose, the function drifts into one of two shapes. It becomes a support queue that reacts to whatever gets reported, or a removal machine judged by throughput. Neither can tell leadership whether users are safer.
The cost shows up in the business, not just in user harm. In a 2023 white paper from Take This, using survey data Nielsen collected from 2,328 players in North America, 61% said they had at least once decided not to spend money in a game because of how other players treated them, and 60% had quit a match or a game because of harassment or hate (Take This, Toxic Gamers Are Alienating Your Core Demographic).
Parts of the job are also law now. In the EU, the Digital Services Act requires hosting services to let anyone flag illegal content (Article 16) and to explain each restriction to the user with a statement of reasons (Article 17) (Regulation (EU) 2022/2065). In the UK, the Online Safety Act requires user-to-user services to assess the risk of illegal content on their service (section 9) and to let users report it easily (section 20). The question has moved from "do you have a policy?" to "can you prove it works?"
What good looks like
What you have, and what you can show, at each stage
Early
A founder or first safety hire covers trust and safety, usually with under a million users.
What you have
A one-page statement of what T&S is for and what it owns, a named owner for each seam with security, fraud, legal and support, a log of every report and decision, and one outcome number next to the activity counts.
What you can show
Who owns each kind of harm. Time to action on the most severe reports. One rough outcome number, such as a weekly labeled sample of what users see.
Growing
A dedicated safety team, millions of users, and new markets or features on the way.
What you have
A charter agreed with Product, Legal and Security, a scorecard that pairs operational and outcome metrics, safety exposure as a standard cut on the retention dashboard, and a quarterly review with leadership.
What you can show
How much of what users see breaks the rules, with a confidence interval. 7- and 30-day return for users exposed to harm against a matched group. How often appeals overturn decisions, by policy.
At scale or regulated
Tens of millions of users, a heavily regulated sector, or extra duties as a very large platform under EU or UK law.
What you have
A charter tied to your legal duties and risk assessments, a target, an owner and a trend for every metric, safety metrics as a required input to launch and business reviews, and a written record behind every threshold.
What you can show
Whether exposure to severe harm is falling, whether bad actors come back after enforcement, how early high-risk behavior is caught, and what harm costs in retention, with public reporting that matches what enforcement does.
How to do it
7 steps
Jump to a step
- Step 01
Write the job down on one page
Most T&S teams can list what they do. Fewer can say what it's for. Start with four jobs:
Job What it means How you'd know it's happening Keep people from harm Prevent, find and stop abuse between users, with the most effort on the most severe harms Exposure to severe harm goes down, and high-risk behavior is caught earlier Make fair decisions fast Apply clear rules the same way every time, act fastest on the worst cases, and fix mistakes Time to action on severe cases, appeal overturn rate, good users wrongly actioned Meet your legal duties Report what the law says to report, keep the records, and answer regulators and law enforcement Every duty has an owner, a process and evidence Keep users' trust People can report, get an answer, understand decisions, and feel safe enough to stay Users who say they feel safe, and retention of people who were targeted The jobs sometimes pull against each other. Speed can cost fairness, and a cautious legal reading can slow protection. Decide in advance which one wins where. One rule this handbook holds to: where a wrong decision can't be reversed or someone's safety is at risk, automation can prepare the case, but a person closes it.
Usually the T&S lead writes the page and an executive sponsor signs it, which in a small company may be a founder. It's done when a product manager, a lawyer and a support lead can each say what T&S does, and what it doesn't.
- Step 02
Draw the borders with neighboring teams
T&S shares edges with at least five other functions. The names vary by company, so describe the work, not the org chart. Where the lines fall varies too, so treat the splits below as a common starting point, not the only one:
Team What it usually owns Where it meets T&S What to write down Security Attacks on your systems and data: breaches, vulnerabilities, misuse of internal access Account takeover, scraping, staff looking up user data Security owns the breach. T&S owns what was done to users through the stolen accounts. Fraud and payments risk The company's own losses: stolen cards, chargebacks, promotion abuse Scams where one user tricks another into paying Who owns a scam that starts in chat and ends in a payment Legal Reading the law, legal process, privilege Reporting duties, law enforcement requests, regulator questions Legal makes the legal call. T&S runs the process, including outside office hours. Customer support First contact with users Reports that arrive as tickets, appeals, users angry after enforcement Support routes and explains. T&S decides. Integrity Varies most: often fake accounts, coordinated manipulation, spam and misinformation. At some companies it's another name for T&S. Bots, coordinated campaigns, fake engagement Whether integrity is a separate team, and who owns coordinated abuse The rule: every harm in your risk register (chapter 2) has exactly one owner. Shared work is fine. Shared ownership means nobody is accountable when it goes wrong.
The T&S lead drafts the table with each partner team. It's done when each partner has agreed to it and the on-call list matches it.
- Step 03
Make the safety case with your own data
Pull raw retention data at most game studios and harassed players look like some of the best-retained users. They chat more, queue more and play longer, so they run into more abuse. A team that stops there concludes toxicity doesn't hurt retention, and the safety budget conversation ends before it starts.
Industry surveys like the one above help, but they rarely move a budget. Your own data does, if it's cut correctly. The cut that works is a matched cohort:
- Take new users whose early sessions included an actioned incident: they were the target of something you confirmed and acted on.
- Match each one to a new user with clean sessions and similar activity. In a game, match on playtime, mode, region and platform.
- Compare their 7- and 30-day return.
Without the matching, the analysis measures engagement instead of harm. Present the result as an association unless you've run a proper causal analysis.
The same blind spot affects enforcement on high-value users. Banning a top spender hits the revenue report the next day. The players they drove away stay invisible unless someone builds the cohort.
Numbers worth tracking, from a game:
- Share of new players whose first five matches include an actioned incident
- D7 and D30 return (the share still active 7 and 30 days later) for exposed new players, against the matched group
- Voice chat opt-out rate in a player's first week
Other products need their own cut. On a dating app, focus on users harassed in their first week, when a bad experience makes people delete the app. On a social platform, match on audience size, because larger accounts are targeted more and churn differently. The churn after toxic exposure page has the method and a starter query.
By stage: an early team without a data team can compare the 30-day retention of users who filed a harassment report with everyone else. It's rough, but directional. A growing team builds the matched cohort once and reruns it on a regular cadence, such as every quarter. At scale, safety exposure is a standard cut on the retention dashboard, owned jointly by T&S and the data team, and reviewed with every other churn driver. Chapter 19 turns this into a budget case.
- Step 04
Separate the machine from the outcome
If your team took down 40% more harmful content last quarter than the quarter before, is that good news? You don't know yet. Detection may have improved. Harm on the platform may have grown. Or automation got more aggressive, and legitimate users are paying for it. Volume alone can't tell you which.
Most T&S reporting measures how much work the system does, because those numbers are easy to produce and they almost always go up. But what gets measured gets resourced. A program rewarded for activity optimizes for activity.
Operational metrics: the machine is running Outcome metrics: it's working Items removed Is exposure to severe harm going down? Reports actioned After enforcement, do bad actors stop, or come back on new accounts? Automation rate Are we catching high-risk behavior earlier? Time to decision What do appeals say about where our policy or automation is wrong? You need both columns. Operational metrics tell you where the process is breaking this week. Outcome metrics tell you whether any of it matters.
Some numbers don't belong on an executive slide on their own. Total items removed rises with volume and with over-enforcement. Reports received measures how easy reporting is as much as how much harm exists. Automation rate counts the decisions people didn't make, not whether they were right. Accounts banned is easy to inflate with throwaway spam accounts. The running the program guide lists the rest.
- Step 05
Report both, and say when they disagree
A metric only changes behavior when a named meeting looks at it on a schedule. One rhythm that works:
- Weekly operations review: this week's problems. Time to action, backlog, quality, reports.
- Monthly health review: trends and owners. Detection, appeals, repeat offending, harm rates by area.
- Quarterly executive review: outcomes, risk and investment. Exposure to harm, retention, legal readiness.
A small team might fold the first two into one meeting, and a live product with fast-moving harms might watch some numbers daily. What matters is that every metric has a meeting that reads it.
Put operational and outcome numbers on the same page, and read them together. When they disagree, say so. It takes discipline to walk into a business review and say, "Our numbers are up, and I'm not convinced we're safer." That's the conversation that earns credibility, sends investment to the right places, and keeps T&S from being judged like a cost center measured by throughput.
Know the limits of each signal. Appeal overturns only measure over-enforcement, because nobody appeals the harm you missed, so pair them with a random sample of what users actually see. Measure for a while before you commit to a target, for example six to eight weeks, because early targets are guesses.
Keep what the company says in public in line with what enforcement actually does. Public safety claims are evidence in litigation. Chapter 11 covers measurement in depth, and chapter 15 covers reporting to executives and the board.
- Step 06
Use the principles to settle arguments
This handbook takes positions. The ten principles are listed in full in the README, each with the post it comes from. Every chapter holds to them. In practice, each one answers an argument you'll have:
The argument The principle that answers it Where to go deeper "Removals are up 40%. Good quarter?" 1. Measure impact, not activity Chapter 11 "Can we automate this whole policy area?" 2. Automate as far as the evidence supports, and 3. Automation earns its scope Chapter 18, chapter 10 "Each message looked fine on its own." 4. Harm is a pattern, not a message Chapter 5 "Should we lock the whole room?" 5. Put friction where the risk is Chapter 5, chapter 13 "We banned them. Case closed?" 6. A ban is one move, not a closed case Chapter 12 "Do we really need age checks?" 7. Age assurance is the foundation Chapter 6 "Product ships DMs next month. Can you look at it?" 8. Get in at design review, and earn the invite Chapter 2, chapter 15 "A regulator asks why this account wasn't restricted." 9. Be able to prove it works Chapter 16 "Finance asks what safety returns." 10. Safety is a retention and revenue question Chapter 19 Of the ten, the one I've found hardest to hold to is that automation earns its scope. The pressure to automate usually runs ahead of the evidence, so agree the rules for expanding automation before that pressure arrives (chapter 18).
Agree the principles with your leadership early, while nothing is on fire. A principle everyone signed up to in a calm week is far easier to apply in the middle of an incident than one you have to argue for on the spot.
- Step 07
Know which outcomes you're aiming at
Every later chapter ends its "What you can show" row with numbers. They roll up to a short list of outcomes, each with a metric page in the T&S Metrics Framework:
Outcome Metrics that show it Less exposure to harm Violating-content prevalence, plus the north star for your kind of product, such as toxicity per 1,000 match-hours or the unsafe-contact rate for minors The worst cases handled fastest Time to action by severity (p90), time to report child sexual exploitation Fair decisions Appeal overturn rate, good users wrongly actioned Harm caught earlier Proactive detection rate, and contact prevented before it happens Bad actors stop Repeat-offender rate Users feel safe and stay Users who feel safe, churn after toxic exposure Legal duties met, with evidence Statement-of-reasons coverage, systemic-risk assessment currency You don't need all of them at once. An early team picks the one north star that fits its product and tracks time to action on the most severe cases. A growing team adds decision quality and repeat offending. A team at scale or under regulation tracks all seven, each with a target and an owner.
Mistakes to avoid
And what to do instead
- 01
Judging the program by removals
Volume can't tell better detection from more harm or over-enforcement. Report an outcome next to every activity number.
- 02
Leaving the borders unwritten
When two teams think they share a harm, nobody owns it. Give every harm and every cross-team request one owner, in writing.
- 03
Stopping at raw retention data
Harassed users look engaged because they are. Compare them with a matched group before you conclude anything.
- 04
Treating automation rate as maturity
It counts the decisions people didn't make, not whether they were right. Judge automation on decision accuracy, cost per decision and appeal overturn rate.
- 05
Smoothing over numbers that disagree
If activity is up and outcomes aren't, say so in the review. It's the conversation that earns credibility.
- 06
Promising more in public than enforcement does
Public safety claims are litigation evidence. Have T&S check every safety claim before Comms makes it.
- 07
Running T&S only as a queue
Waiting for reports leaves you blind to the harm nobody reports, and much of it never is. Get into design review (chapter 2) and measure what users actually see.
Start from this template
Copy it, fill it in, make it yours
Template
One-page T&S charter
Fill in each row, then have your executive sponsor, Product, Legal and Security agree to it.
| Section | What to write |
|---|---|
| Why we exist | The four jobs, in your words, and which one wins when they conflict |
| What we own | The harms and decisions T&S owns end to end |
| What we don't own | The neighboring work, and the team that owns it |
| Decisions we make alone | For example: removing content under a written policy, suspending an account |
| Decisions we make with others | For example: reporting to law enforcement with Legal, public statements with Comms |
| Outcomes we report | One north star, time to action on severe cases, and one decision-quality number |
| Who we answer to, and how often | The executive sponsor, and the weekly, monthly and quarterly reviews |
Template
Seams table
One row for each harm or request that crosses teams.
| Harm or request | Partner team | What T&S does | What the partner does | Who decides | Who's on call |
|---|---|---|---|---|---|
| Account takeover | Security | ||||
| Scam that ends in a payment | Fraud | ||||
| Law enforcement request | Legal | ||||
| Appeal that arrives through support | Support | ||||
| Coordinated fake accounts | Integrity or Security |
Template
Metrics pair
For every operational number you report, name the outcome it should move and the number that keeps it honest.
| Operational number | Outcome it should move | Read it with |
|---|---|---|
| Items removed | Violating-content prevalence | Appeal overturn rate |
| Time to decision | Harmful reach before action | QA agreement rate |
| Accounts banned | Repeat-offender rate | Appeal overturn rate |
| Reports received | User-report rate, by reason | Violating-content prevalence |
| Automation rate | Precision by policy area | Appeal overturn rate |
Do it with
Free tools and metrics that go with this chapter
Metrics framework
Pick your platform and stage, and build a scorecard that puts outcomes next to operations. Open content
Churn after toxic exposure
The retention gap between users who were targeted and matched users who weren't.
Running the program
Which meeting looks at which metric, how to set targets, and the numbers to keep off the executive slide.
Further reading
Steven's posts on this topic, and sources worth the time
From Steven's writing · 3 posts
- Essay
Harassed players look like your best-retained users
Raw retention data hides the cost of toxicity because harassed players are the most engaged. A matched cohort shows it, and safety exposure belongs on the retention dashboard.
- Essay
Activity is easy to measure. Impact is harder.
A 40% jump in removals could mean better detection, more harm or over-enforcement. Report outcome metrics next to operational ones, and tell leadership when they disagree.
- Essay
You can't moderate your way out of a systems problem
Treating Trust & Safety mainly as an operations function is a mistake. Reputation, history, age and behavior signals belong in one risk model, automation needs clear limits, and safety belongs in the product architecture from the start.
Outside sources
- TSPA: Trust & Safety Fundamentalsthe Trust & Safety Professional Association's free curriculum, from policy and operations to law enforcement and safety by design.
- Digital Trust & Safety Partnership: Best Practices Frameworkfive industry commitments covering product development, governance, enforcement, improvement and transparency.
- Take This: Toxic Gamers Are Alienating Your Core Demographicthe 2023 white paper, with Nielsen data, on how harassment changes what players spend and whether they stay.
- European Commission: the Digital Services Acta plain-language overview of the EU's rules on reporting, explaining decisions and appeals.
Recent changes
5 changes to this chapter, newest first
The updates page has every change to the handbook, by date.
- AddedAlso: 5. Detection and prevention, 8. Hiring and structuring the team, 15. Working with Product, Legal, Comms and leadership
Linked Steven's post on why treating Trust & Safety as an operations function is a mistake, and added it to chapter 8's view on where the team should report.
- RevisedWith 18 other chapters
Practical advice that varies by platform, such as cadences, sample sizes, targets and who owns what, is now set out as options with examples, so each team can choose what fits.
- RevisedWith 14 other chapters
Added Steven's own calls from an interview: where Trust & Safety should report, what to automate first, the one number to track from day one, who makes the 2am call, and more. Practical choices that vary by platform are now laid out as options.
- DraftedWith 17 other chapters
First full drafts of the other 18 chapters, built on Steven's posts, the handbook's principles and the Workbench's open content, with every legal and factual claim checked against its source. Stories from Steven's own work come next.
- AddedWith 14 other chapters
Linked the first 16 posts to the chapters they inform, and set out the ten principles behind the handbook.