In one minute

Know your cost drivers before Finance models them for youPeople, vendors and tools each have a driver: volume, handle time, hours of coverage, languages and the mix of severities. Change the driver and the cost moves.
Don't let the program be judged by throughputRemovals and decisions per hour show the machine is running, not that it's working. Report outcomes next to operations, and say so when they disagree.
Make the revenue case with your own data, cut correctlyIn raw retention data, harassed users look like your best-retained users. A matched cohort shows what exposure costs, and Finance's own numbers turn that into money.
Tie every ask to a roadmap leadership can followRate the program area by area against the target for your stage, close the biggest gaps first, and name the risk and the metric behind each ask.
Plan the cuts before a tight year forces themKnow what you'd cut first, what you never cut, and who signs off on the risk each cut accepts.
The mistake to avoidDefending the budget with volume. When volume is the argument, efficiency becomes the only question, and the program gets judged as a cost center.

Why it matters

What gets measured gets resourced. When a program is rewarded for activity, it optimizes for activity, and when it asks for money with activity, the answer is usually a question about doing the same work more cheaply. Activity numbers almost always go up and to the right, so on a budget slide the program looks like a growing cost with no visible return.

The return is real, but it hides. In raw retention data, harassed players often look like some of the best-retained users, because they chat more, queue more and play longer, so they run into more abuse. A team that stops there concludes toxicity doesn't hurt retention, and the safety budget conversation ends before it starts. Players say otherwise: in a 2023 Take This report based on a Nielsen poll of 2,328 teens and adults in North America, 61% said they had at least once decided not to spend money in a game because of how other players treated them (Take This, 2023). But surveys rarely move a budget. A company's own data does, if it's cut correctly.

What good looks like

What you have, and what you can show, at each stage

Early

A founder or first safety hire covers trust and safety, usually with under a million users.

What you have

A one-page view of what you spend on people, vendors and tools, updated monthly; a short list of top risks and what covers each; one outcome metric shown next to the operational ones; and a regular update to the founders or executive team.

What you can show

Cost per decision by queue next to QA agreement, the share of new users whose early sessions include an actioned incident, and what last quarter's spending changed.

Growing

A dedicated safety team, millions of users, and new markets or features on the way.

What you have

A cost model with its drivers and a capacity forecast; a matched-cohort retention analysis owned jointly with the data team; a roadmap by maturity area with an owner, a cost and a metric on every line; a quarterly executive review; and a cut plan agreed in advance.

What you can show

7- and 30-day return for exposed new users against a matched group, translated into revenue at risk; maturity ratings against stage targets each quarter; and whether last quarter's asks delivered.

At scale or regulated

Tens of millions of users, a heavily regulated sector, or extra duties as a very large platform under EU or UK law.

What you have

Safety exposure as a standard cut on the retention dashboard; an investment case for every major ask, reviewed afterwards against what it promised; a budget tied to the risk register and to legal duties; and accepted risks signed off at the right level.

What you can show

Revenue at risk from exposure as a trend; outcome metrics that moved after specific investments; every legal duty funded and owned; and a record of accepted risks and what became of them.

How to do it

7 steps

  1. Step 01

    Build a cost model you can explain

    Finance will model the program whether or not you do. Build the model first, so the drivers are yours. Split the spend into a few lines, and for each one name what drives it and what you can change:

    CostWhat drives itLevers you control
    In-house people: reviewers, investigators, policy, data and engineeringVolume times handle time, divided by productive hours; hours of coverage, including round-the-clock cover for the most severe harms; the number of languagesDetection and routing that send cases to the right queue, friction that prevents harm before anyone reviews it, and clearer guidance that cuts handle time
    Vendors: review staffing, AI moderation, age assurance, translationPrice per decision or per hour, minimum commitments, surge terms, and wellbeing and quality standardsContract terms, price adjusted for quality, and consolidating overlapping vendors
    Tools and infrastructureClassifiers, hash matching, case management and computeFree and open-source tools, and building only what's specific to you
    ComplianceRisk assessments, transparency reports, audits and legal adviceLogging that turns a report into a query rather than a project (chapter 17)

    The unit that ties the lines together is cost per decision: the fully loaded cost of review divided by the decisions made, by queue and vendor. Never show it alone. Optimizing it on its own lowers quality, so put it next to QA agreement on every slide. In the metrics framework's worked example, one vendor costs $0.38 a decision against another's $0.52, but agrees with expert reviewers 8 points less often, and rework and appeals make the cheaper vendor the more expensive choice.

    Capacity is the biggest driver, and it's forecastable from volume trends (chapter 7). Before you ask for headcount, check the levers that reduce volume or handle time. Free tools such as those from ROOST can cover hash matching and review workflow without a license fee (chapter 9).

    By stage: an early team divides each vendor invoice by the decisions it covered, and tracks people and tools in a spreadsheet. A growing team builds the model with Finance and forecasts on Finance's planning cycle, quarterly for example. At scale, every queue and vendor has a cost per decision, a quality number and a forecast, owned by T&S operations with Finance.

  2. Step 02

    Stop being judged by throughput

    If your team took down 40% more harmful content last quarter than the quarter before, is that good news? You don't know yet. It could mean detection improved, harm grew, or automation got more aggressive and legitimate users are paying for it. Volume can't tell you which, and a budget argued on volume invites one response: do the same volume for less.

    Keep these off the executive slide: total items removed, reports received, automation rate on its own, moderator headcount and accounts banned. Each can rise for good or bad reasons, and none says what users experience. Automation rate is an efficiency metric, not a measure of maturity. It counts the decisions people didn't make, not whether they were right.

    Lead with the outcome questions instead:

    • Is exposure to severe harm going down?
    • After enforcement, do bad actors stop, or come back on new accounts?
    • Are we catching high-risk behavior earlier?
    • What are appeals telling us about where our policy or automation is wrong?

    Then frame the budget the same way. "Child safety: 12 reviewers" is a cost line. "Child safety: unsafe contact per 10,000 young users, down from last quarter, and what it takes to keep it falling" is an outcome with a price. Chapter 11 covers the metrics.

    Report operational and outcome numbers side by side, and say so when they disagree. It takes discipline to walk into a business review and say, "Our numbers are up, and I'm not convinced we're safer." That's the conversation that earns credibility, directs investment to the right places and keeps Trust & Safety from being judged like a cost center measured by throughput.

  3. Step 03

    Make the revenue case with a matched cohort

    The cut that works is a matched cohort. Take new players whose early sessions included an actioned incident, and compare their 7- and 30-day return against new players with clean sessions, matched on playtime, mode, region and platform. Without the matching, the analysis measures engagement instead of harm. Chapter 11 and the churn after toxic exposure page cover the method. Outside games, match on what drives exposure on your platform, such as audience size on a social app.

    A rough comparison, such as people who reported harassment against everyone else, can tell you where to look. A budget case needs the matched version.

    Then turn the gap into money, using Finance's numbers rather than your own:

    1. Users lost. The retention gap times the number of exposed new users each month.
    2. Revenue at risk. Users lost times the value of a retained user, from the same figures Finance reports to the board.
    3. Skipped spending. The difference in spend between exposed users and their matched group over the same window.

    Here's an illustration of the arithmetic, from the metrics framework's worked example: exposed players retain at 52% after 30 days against 61% for matched players. A 9-point gap across 20,000 exposed players a month is about 1,800 players a month, and Finance can put a value on each one.

    Be honest about what it is. Present the result as a range, and call it an association unless you've run a proper causal analysis. With children's data, involve your privacy team and report only aggregates. An overclaimed number is the fastest way to lose the room the second time.

    Numbers worth tracking:

    • Share of new players whose first five matches include an actioned incident
    • D7 and D30 return for exposed new players against the matched group
    • Voice chat opt-out rate in a player's first week

    Safety exposure belongs as a standard cut on the retention dashboard, owned jointly by T&S and the data team and reviewed alongside every other churn driver. Once it's there, the revenue case is updated every month without anyone having to argue for it.

  4. Step 04

    Count the players a top spender drives away

    The same blind spot affects enforcement on high-value users. Banning a top spender hits the revenue report the next day. The players they drove away stay invisible unless someone builds the cohort.

    So build it before the argument starts:

    • When a high-value account is up for enforcement, pull the users it targeted in confirmed incidents over a set window.
    • Compare their return and spend afterwards against matched users who weren't targeted.
    • Put both numbers in the decision record: the account's own revenue, and the return and spend gap among the people it targeted.

    The cohort completes the revenue argument. It doesn't change the rule. Enforcement follows the same policy for every account, because consistency means the same penalty rules for everyone. If someone wants an exception for a spender, the question goes to the executive who owns safety risk, with both numbers in front of them.

  5. Step 05

    Build a roadmap leadership can follow

    Leadership says yes to plans it can check. The program maturity model gives you the structure: eight areas (policy, detection and prevention, review operations, quality and appeals, crisis response, metrics and reporting, regulatory readiness and team wellbeing), each rated from 1 (reactive) to 5 (leading).

    The model suggests a target for each stage:

    StageTarget
    EarlyLevel 2, and level 3 for crisis response, regulatory readiness and team wellbeing
    GrowingLevel 3 in every area
    At scale or regulatedLevel 4 in every area

    Crisis response, compliance and wellbeing stay at level 3 even for early teams, because the harm is the same whatever your size.

    Build the roadmap from the gaps, biggest first. When two gaps are the same size, one order that works is crisis response, then regulatory readiness, then detection and prevention, though your own risks may argue for another. Each level has two concrete steps to the next level, so every line on the roadmap is something specific.

    Then check the roadmap against risk. The harm coverage radar compares each harm area's risk with its coverage across policy, detection, enforcement, appeals and measurement. Without that view, budget follows whoever argues loudest rather than where the exposure is.

    Give each line the same six things: the gap, the step, the owner, the cost, the risk it reduces (from the abuse pre-mortem or incident data) and the metric that will show it worked. Legal deadlines go on with their fixed dates (chapter 16). Keep it short, for example one page in three columns: now, next and later. Take a snapshot of the maturity ratings every quarter, so progress is visible without a long explanation.

  6. Step 06

    Run the quarterly executive review

    A metric only changes behavior when a named meeting looks at it on a schedule. A quarterly executive review, or whatever cadence matches your planning cycle, brings together executives, Legal and Finance to look at outcomes, risk and investment. Send a pre-read, and spend the meeting on decisions.

    PartWhat's in it
    OutcomesNorth-star metrics against last quarter and the target band, with ranges
    Where the numbers disagreeOperational metrics next to outcomes, and what you think explains the difference
    Retention and revenueThe exposure cohort: retention gap, revenue at risk and the trend
    RiskTop risks, accepted risks and what happened to them, and legal duties coming due
    RoadmapWhat shipped, what slipped and why, and the maturity snapshot
    Asks and decisionsEach ask with the risk it reduces, its cost and the metric that will show it worked

    Set targets carefully. Measure long enough to see the normal range before you commit to a number, for example 6 to 8 weeks. Set a band, not a point, so the gap between "on track" and "off track" is your early warning. Pair every target with a guardrail, and show the uncertainty on anything from a sample.

    Close the loop on past asks. Each quarter, show what last year's investments promised and what they delivered, including the ones that didn't work. Nothing earns the next yes like a record of reporting honestly on the last one.

  7. Step 07

    Plan the cuts before you need them

    Tight years come. A team that has already decided what it would cut, and what it never would, makes better choices than one deciding under a deadline. The first column below is an example to adapt to your own risks.

    Cut firstNever cut
    Response-time targets on low-severity queues: relax them, don't abandon themChild safety and severe-harm escalation, including round-the-clock cover for the worst harms (chapter 12)
    Overlapping tools and vendors, and minimum commitments you don't useDuties the law requires, such as child sexual abuse reporting, notices and required risk assessments (chapter 16)
    Full reviews of low-risk launches, replaced with a short checklistReviewer wellbeing: exposure limits and specialist support (chapter 14)
    Manual work that automation can take over, only where its overturn rate is at or below human reviewQuality sampling and prevalence measurement, because without them you can't see what the other cuts did
    Reports and dashboards nobody usesAppeals, because they're how you find out a cut went too far

    Every cut accepts a risk, so treat it like any accepted risk. Write down what it saves, what risk it accepts, the metric to watch and a date to revisit. The person who owns the product outcome signs off, and the most serious risks go to the executive who owns safety risk. If the metric moves past an agreed limit, reverse the cut. A cut that nobody measured is a risk nobody accepted.

Mistakes to avoid

And what to do instead

  1. 01

    Defending the budget with volume

    Lead with outcomes and risk, and keep volume in the operations review.

  2. 02

    Showing raw retention data

    Build the matched cohort first, or harassment will look like it improves retention.

  3. 03

    Claiming causation from a correlation

    Present a range, and call it an association unless you've run a causal analysis.

  4. 04

    Optimizing cost per decision alone

    Show it next to QA agreement, and count rework and appeals.

  5. 05

    Asking for headcount without a roadmap

    Tie each ask to a gap, a risk and a metric.

  6. 06

    Letting a spender's revenue decide enforcement

    Apply the same rules to everyone, with the cohort of the people they drove away in the record.

  7. 07

    Cutting what lets you see

    Measurement and quality sampling go last, because they show what every other cut did.

Start from this template

Copy it, fill it in, make it yours

Template

Investment case

One page per ask.

SectionWhat to write
The riskWhat harm, to whom and how often, from the pre-mortem or incident data
The evidenceCurrent metric and trend, cohort result and maturity gap
The askPeople, money or engineering time, and when
Expected effectWhich metric should move, by how much and by when, as a range
How we'll knowThe metric, its owner and the review date
If we don'tThe risk accepted, and who accepts it

Template

Revenue at risk worksheet

Fill it in with the data team and Finance.

InputSourceValue
Exposed new users per monthT&S and data
D30 return, exposed against matchedMatched cohort
Users lost per month in association with exposureGap times exposed users
Value of a retained userFinance
Revenue at risk per month, as a rangeUsers lost times value
Spend gap, exposed against matchedMatched cohort

Template

Cut plan

One row per cut, agreed before you need it.

ItemSavingRisk acceptedMetric to watchSigned off byRevisit date

Do it with

Free tools and metrics that go with this chapter

Workbench tool

Program maturity

Rate the program in eight areas against the targets for your stage, and get a phased roadmap. Open content

Workbench tool

Metrics framework

36 metrics with formulas, measurement steps and SQL, and a one-pager for leadership. Open content

Workbench tool

Vendor scorecard

Choose a moderation vendor on evidence: weighted criteria, rubrics, minimums and RFP questions. Open content

Workbench tool

Harm coverage radar

See where each harm area's risk outruns your defenses, so budget follows exposure. Open content

Metric

Churn after toxic exposure

The retention gap between users targeted by a confirmed incident and matched users who weren't, the starting point for revenue at risk.

Metric

Cost per decision

The fully loaded cost of each review decision by queue and vendor, always read next to QA agreement.

Reference

Running the program

The weekly, monthly and quarterly review rhythm, target setting, and the numbers to keep off the executive slide.

Further reading

Steven's posts on this topic, and sources worth the time

From Steven's writing · 2 posts

  1. Essay

    Harassed players look like your best-retained users

    Raw retention data hides the cost of toxicity because harassed players are the most engaged. A matched cohort shows it, and safety exposure belongs on the retention dashboard.

    Also in: 1. What Trust & Safety is for, 11. Measuring what matters

  2. Essay

    Activity is easy to measure. Impact is harder.

    A 40% jump in removals could mean better detection, more harm or over-enforcement. Report outcome metrics next to operational ones, and tell leadership when they disagree.

    Also in: 1. What Trust & Safety is for, 11. Measuring what matters

Outside sources

Recent changes

3 changes to this chapter, newest first

The updates page has every change to the handbook, by date.

  1. RevisedWith 18 other chapters

    Practical advice that varies by platform, such as cadences, sample sizes, targets and who owns what, is now set out as options with examples, so each team can choose what fits.

  2. DraftedWith 17 other chapters

    First full drafts of the other 18 chapters, built on Steven's posts, the handbook's principles and the Workbench's open content, with every legal and factual claim checked against its source. Stories from Steven's own work come next.

  3. AddedWith 14 other chapters

    Linked the first 16 posts to the chapters they inform, and set out the ten principles behind the handbook.