In one minute

Report outcomes next to operations, and say so when they disagreeRemovals and automation rate show the machine is running. Exposure to harm, repeat offending and prevented contact show it's working.
Adopt metrics a few at a timeOne order that works: north stars, then health metrics, then diagnostics. Each needs an owner, a target band and a meeting that looks at it on a schedule.
Measure exposure with a random sample of what users seePrevalence sampling is the only way to count the harm nobody reported.
Pair appeals with prevalenceOverturned appeals show what you over-enforced. Sampling shows what you missed.
Prove the retention case with a matched cohortIn raw data, harassed users can look like your best-retained users, because they're your most engaged ones.
The mistake to avoidCelebrating a jump in removals. 40% more could mean better detection, more harm, or good users caught by mistake.

Why it matters

What gets measured gets resourced. When a program is rewarded for activity, it optimizes for activity: more removals, more reports closed, a higher automation rate. Those numbers are easy to produce and almost always go up and to the right. None of them says whether users are safer, so a safety team measured only by throughput gets judged like a cost center.

The outcomes are real and measurable. In a 2023 Take This report based on a Nielsen poll of 2,328 teens and adults in North America, 61% of players said they had at least once decided not to spend money in a game because of how other players treated them, and 60% had quit a match or a game because of harassment and hate (Take This, 2023). But surveys rarely move a budget. Your own data does, if it's cut correctly.

Regulators are asking too. The EU's Digital Services Act requires most services' transparency reports to include indicators of the accuracy and possible error rate of automated moderation, and the number of complaint decisions that were reversed (Article 15(1)(d) and (e)). Counting what you did is no longer enough. You need to show whether it was right and whether it worked.

What good looks like

What you have, and what you can show, at each stage

Early

A founder or first safety hire covers trust and safety, usually with under a million users.

What you have

Every decision logged with its policy, source and timestamps; one north star for your platform, even from a small weekly sample; appeals tagged upheld or overturned; and a monthly one-page review.

What you can show

Your north star with an honest range, the overturn rate, how fast the most severe cases are handled, and what changed this month.

Growing

A dedicated safety team, millions of users, and new markets or features on the way.

What you have

Health metrics for detection, decision quality and operations, each with an owner and a target band; prevalence from a properly sized random sample; repeat-offender tracking; a weekly, monthly and quarterly review rhythm; and safety exposure as a cut on the retention dashboard.

What you can show

Prevalence with a 95% confidence interval by policy area, overturns read next to prevalence, the share of new users whose early sessions include an actioned incident, and outcome and operational metrics side by side on every leadership slide.

At scale or regulated

Tens of millions of users, a heavily regulated sector, or extra duties as a very large platform under EU or UK law.

What you have

Diagnostics across languages, markets, emerging trends and people; matched-cohort retention analysis every quarter; safety metrics required in launch and business reviews; and numbers that reconcile with the transparency report and the company's own figures.

What you can show

A trend, target and owner for every north star, 7- and 30-day return for exposed new users against a matched group, evidence that specific safety changes moved outcomes, and accuracy and error rates for automated tools that hold up in a transparency report or an audit.

How to do it

8 steps

  1. Step 01

    Separate activity from impact

    If your team took down 40% more harmful content last quarter than the quarter before, is that good news? The honest answer is that you don't know yet. It could mean detection improved. It could mean harm on the platform grew. It could mean automation got more aggressive, and legitimate users are paying for it. Volume alone can't tell you which.

    Activity metrics measure how much work the system does: removals, reports actioned, automation rate and time to decision. Impact metrics answer four harder questions. Map each one to something you can measure:

    The questionWhat to measure
    Is exposure to severe harm going down?Violating-content prevalence, harmful reach before action, or your platform's own harm rate
    After enforcement, do bad actors stop, or come back on new accounts?Repeat-offender rate, with ban evasion counted
    Are we catching high-risk behavior earlier?Time to action by severity, time to detect emerging trends
    What are appeals telling us about where policy or automation is wrong?Appeal overturn rate by policy area and enforcement source (chapter 10)

    Some numbers belong in the weekly operations review but not on the executive slide: total items removed, reports received, automation rate on its own, moderator headcount and accounts banned. Each can rise for good or bad reasons, and none says what users experience. The running the program guide explains why.

  2. Step 02

    Adopt north stars, then health metrics, then diagnostics

    If you can track only one number from day one, I'd make it harmful reach before action: how many people saw something harmful before you acted on it. It ties your speed to the harm people actually experienced, and it's hard to game.

    The metrics framework sorts 36 metrics into three tiers, in a suggested order of adoption.

    TierWhat it tells youExamplesWho looks, and when
    North starWhether users are actually safer. A few top-level numbers.Violating-content prevalence, fraud loss rate, toxicity per 1,000 match-hours, unsafe-contact rate for minors, users who feel safeExecutives every quarter, T&S leadership more often
    HealthWhether the system that produces safety is workingProactive detection, precision and recall, appeal overturn rate, QA agreement, time to action by severityT&S leadership monthly, operations weekly
    DiagnosticWhy a number moved, and where to lookChurn after toxic exposure, repeat-offender rate, consistency across languages, backlog age, cost per decisionWhoever is investigating, plus a quarterly look

    Pick one or two north stars that fit your platform. A social or user-generated content service usually starts with prevalence. A marketplace or payments product starts with fraud loss. A game starts with toxicity per match-hour. A product used by children adds the unsafe-contact rate for minors (chapter 6).

    Don't adopt everything at once. Each metric page says which program stage it suits, from "early and later" to "mature and later". An early team can run on one north star, the overturn rate, QA agreement, time to action for the most severe cases and proactive detection. Add more only when someone will act on them.

    Every metric on the framework names another metric to read it with, because any number can be gamed. A rising proactive detection rate with falling precision means you're catching more innocent content. Keep the pairs together on the same page.

  3. Step 03

    Measure exposure with prevalence sampling

    Prevalence is the share of what users see that breaks a rule. It measures exposure, not effort, and it's the only number that counts harm nobody reported.

    The method matters more than the formula, which is on the violating-content prevalence page:

    • Sample views, not posts. A post seen a million times should be a million times more likely to be picked.
    • Label blind, against the current policy, with two labels per item and an expert to break ties.
    • Report a confidence interval every time. Rare harms need large samples, and a change inside the interval isn't a change.
    • Split by surface and audience. Report feeds, search and recommendations separately, and what young users see separately from adults.

    Ownership varies, but data science often runs the sampling pipeline and Policy the labeling guidelines.

    You don't need a data team to start. Pull a random sample of views each week from your logs, 200 for example, and have two people label them in a spreadsheet. It's imprecise, but it's real, and it's more honest than a removal count.

    Recall works the same way. To estimate how much your systems miss, label a sample of content nothing flagged, and scale up what you find (precision and recall by policy area).

  4. Step 04

    Measure what bad actors do next, and what you prevented

    Removal is one move. Measure what happens after it, and what never happened because of a safeguard.

    Repeat offending. Track the share of actions against accounts with a recent prior violation. Split it by the earlier penalty, so you learn which penalties actually change behavior. Link accounts by device, payment and contact signals where your privacy policy allows, or ban evasion will hide your repeat offenders behind new accounts. A ban is one move, not a closed case.

    Earlier detection. Time to action matters most at the slow end, so report the 90th percentile by severity, not the average. Harmful reach before action ties speed to harm: removing a viral post fast matters more than removing an unseen one fast. Report the tail, because a few viral items usually carry most of the reach.

    Prevention. Count risky contact prevented alongside content removed. For a product with young users, that means adults found in teen spaces, the unsafe-contact rate per 10,000 young users and the share of cases caught before the move to another app (chapter 6). Safer defaults and targeted limits stop most harm before anything needs removing, so if you only count removals, your best work is invisible.

  5. Step 05

    Pair appeals with prevalence

    Appeals and prevalence measure opposite errors, so read them together.

    Overturned appeals measure over-enforcement. They're the early warning that automation or guidance has drifted. But nobody appeals the harm you missed, so overturns need to sit next to prevalence sampling, which measures under-enforcement. Appeal rates also vary by market, so a low overturn rate where few people appeal isn't automatically good news.

    Read the two together:

    PrevalenceOverturnsWhat it probably means
    DownFlat or downReal progress
    DownUpYou're catching more, and good users are paying for it
    UpFlat or downHarm is growing, or detection is missing it. Check recall.
    FlatUpAutomation or guidance has drifted. Break overturns out by source and policy area.

    Then estimate the people hurt by mistakes. Most users who are wrongly actioned never appeal. They leave. Good users wrongly actioned estimates them from QA error rates as well as overturns, and follows what happens to them next.

  6. Step 06

    Show safety's effect on retention with a matched cohort

    Pull the raw retention data at most game studios and harassed players look like some of the best-retained users. They chat more, queue more and play longer, so they run into more abuse. A team that stops there concludes toxicity doesn't hurt retention, and the safety budget conversation ends before it starts.

    The cut that works is a matched cohort: a comparison group chosen to look like the group you're studying, so the main difference is the thing you're measuring. Take new players whose early sessions included an actioned incident, and compare their 7- and 30-day return against new players with clean sessions, matched on playtime, mode, region and platform. Without the matching, the analysis measures engagement instead of harm.

    The same blind spot affects enforcement on high-value users. Banning a top spender hits the revenue report the next day, while the players they drove away stay invisible unless someone builds the cohort.

    Numbers worth tracking:

    • Share of new players whose first five matches include an actioned incident
    • D7 and D30 return (the share still active 7 and 30 days later) for exposed new players against the matched group
    • Voice chat opt-out rate in a player's first week

    Safety exposure belongs as a standard cut on the retention dashboard, owned jointly by T&S and data teams and reviewed alongside every other churn driver.

    A few cautions. Define exposure as being the target of a confirmed incident, not just filing a report. Present the result as a range, and call it an association unless you've run a proper causal analysis. Outside games, match on what drives exposure on your platform, such as audience size on a social app. With children's data, involve your privacy team and report only aggregates. The churn after toxic exposure page has the steps and a starter query. Chapter 19 turns the result into a budget case.

  7. Step 07

    Log the data from day one

    Most teams can't measure well because a timestamp or a field was never logged. You can't go back and add it later. The data to log guide lists the event logs every metric depends on. The essentials:

    • Every decision with its policy, severity, action, source (automated, proactive or user report), when the item was first reported or detected, when it was decided, who decided, which vendor and queue, the language, and the views at the time of action. The first-seen time and the source are the two fields teams most often forget.
    • Every automated detection with its model version and threshold, or you can never compare before and after a model change.
    • Every appeal joined back to its original decision.
    • A random exposure sample, labeled.
    • Blind QA re-reviews, with experts who can't see the original decision.
    • Usage denominators from the same source the company reports to the board, so your rates reconcile with theirs.

    Decide once when each clock starts, at creation, report or detection, and never change it quietly. A silently changed definition breaks every trend line.

    If your review console is ROOST's open-source Coop, running on Coop shows where each field lives and which ones you'll need to log yourself (chapter 9).

  8. Step 08

    Run a review rhythm, and tell leadership the truth

    A metric only changes behavior when a named meeting looks at it on a schedule. Here is one rhythm that works, which a small team might fold into fewer meetings:

    MeetingWhoWhat it's for
    Weekly operations reviewOperations leads and vendor managersThis week's problems: QA agreement, time to action, response-time targets, backlog, graphic exposure
    Monthly health reviewT&S leadership, policy, detection and qualityTrends and owners: harm rates, proactive detection, precision, overturns, repeat offenders, cost per decision
    Quarterly executive reviewExecutives, legal and financeNorth stars, retention impact, risk and investment
    Every releaseModel safety and product, for AI featuresWhether to ship (chapter 18)

    Set targets carefully. Measure for several weeks, for example 6 to 8, before you commit to a number. Set a band, not a point, so the gap between "on track" and "off track" is your early warning. Pair every target with a guardrail: a response-time target without a quality floor rewards rushing. Set thresholds by severity, and show the uncertainty on anything from a sample.

    Then report operational and outcome metrics side by side, and be willing to say when they disagree. It takes discipline to walk into a business review and say, "Our numbers are up, and I'm not convinced we're safer." That's the conversation that earns credibility, directs investment to the right places and keeps Trust & Safety from being judged like a cost center measured by throughput.

    Activity is easy to measure. Impact is harder.

Mistakes to avoid

And what to do instead

  1. 01

    Leading with removals

    Put outcome metrics first, and keep removal counts in the operations review.

  2. 02

    Adopting every metric at once

    Start with one or two north stars and a few health metrics, each with an owner and a meeting.

  3. 03

    Counting items instead of views

    Sample what users see, so popular content counts more.

  4. 04

    Reporting a sample without its interval

    Show the confidence interval, and don't celebrate a change inside it.

  5. 05

    Reading retention data raw

    Build a matched cohort, or you'll measure engagement instead of harm.

  6. 06

    Treating overturns as the whole quality picture

    Pair them with prevalence and recall sampling.

  7. 07

    Changing a definition quietly

    Write down every change to a metric or a clock, with the date, and restate the trend.

  8. 08

    Hiding the disagreement

    When the numbers are up and users aren't safer, say so.

Start from this template

Copy it, fill it in, make it yours

Template

Leadership one-pager

One row per outcome question, outcome rows first.

QuestionMetricThis periodTarget bandTrendWhat we're doing about it
Is exposure to severe harm going down?
Do bad actors stop after enforcement?
Are we catching high-risk behavior earlier?
What do appeals say about where we're wrong?
Is the machine running? (operations)

Template

Matched-cohort brief

Agree it with the data team before anyone pulls numbers.

ItemYour answer
Exposure: what counts as an actioned incident
Cohort: which new users, and which window
Matched onPlaytime, mode, region and platform, or your platform's equivalents
OutcomeD7 and D30 return, plus spend or opt-outs if available
OwnersT&S and data, jointly
How it will be describedAn association, with a range, unless a causal analysis is run

Template

Metric register

One row per metric you adopt.

MetricTierOwnerMeetingTarget bandRead it withDefinition last changed
North starQuarterly
HealthMonthly

Do it with

Free tools and metrics that go with this chapter

Workbench tool

Metrics framework

36 metrics with formulas, measurement steps and SQL, and a one-pager for leadership. Open content

Workbench tool

Program maturity

Rate your program's metrics and reporting, and seven other areas, against the targets for its stage. Open content

Metric

Churn after toxic exposure

The retention gap between users targeted by a confirmed incident and matched users who weren't.

Metric

Violating-content prevalence

Out of everything people see, how much breaks your rules, from a random sample of views.

Metric

Repeat-offender rate

Whether people stop after enforcement or keep going, and which penalties deter.

Metric

Appeal overturn rate

How often appealed decisions are reversed, read next to prevalence.

Reference

Running the program

Review rhythm, target setting and vanity metrics to avoid.

Reference

Data to log

The event logs every metric depends on.

Further reading

Steven's posts on this topic, and sources worth the time

From Steven's writing · 3 posts

  1. Essay

    Harassed players look like your best-retained users

    Raw retention data hides the cost of toxicity because harassed players are the most engaged. A matched cohort shows it, and safety exposure belongs on the retention dashboard.

    Also in: 1. What Trust & Safety is for, 19. Budgets, roadmaps and making the case

  2. Essay

    Appeal overturns: the early warning for automation

    Overturned appeals show drift weeks before anything else. Automation expands only where its overturn rate matches human review, and changes roll back when the rate moves.

    Also in: 10. Quality, calibration and appeals, 18. AI in Trust & Safety

  3. Essay

    Activity is easy to measure. Impact is harder.

    A 40% jump in removals could mean better detection, more harm or over-enforcement. Report outcome metrics next to operational ones, and tell leadership when they disagree.

    Also in: 1. What Trust & Safety is for, 19. Budgets, roadmaps and making the case

Outside sources

Recent changes

4 changes to this chapter, newest first

The updates page has every change to the handbook, by date.

  1. RevisedWith 18 other chapters

    Practical advice that varies by platform, such as cadences, sample sizes, targets and who owns what, is now set out as options with examples, so each team can choose what fits.

  2. RevisedWith 14 other chapters

    Added Steven's own calls from an interview: where Trust & Safety should report, what to automate first, the one number to track from day one, who makes the 2am call, and more. Practical choices that vary by platform are now laid out as options.

  3. DraftedWith 17 other chapters

    First full drafts of the other 18 chapters, built on Steven's posts, the handbook's principles and the Workbench's open content, with every legal and factual claim checked against its source. Stories from Steven's own work come next.

  4. AddedWith 14 other chapters

    Linked the first 16 posts to the chapters they inform, and set out the ten principles behind the handbook.