In one minute

Hire for your risks, not an org chartThe harms your product invites decide which skills you need first. For a long time, one person will cover several roles, so write down who wears which hat.
Make the first hire someone who can write a rule, run a queue and handle an escalationThen add operations, investigations, data and engineering in the order your risks call for. Support for anyone who sees harmful content comes before any of them.
Report where you can be heard before launch, and can't be overruled quietlyEvery reporting line has a trade-off. Whichever you pick, name the executive who owns safety risk and write down who decides what.
Interview for judgment, not knowledge of the rulesGive candidates a case with no clean answer. Listen for what they'd want to know, what they'd do now that can be undone, and how they'd explain the decision later.
Treat vendor staff as part of the teamSame guidance, same calibration, same wellbeing standards, and career paths that don't run only through the worst queues.
The mistake to avoidHiring a large review team before anyone owns policy, data and escalations. Reviewers can only be as consistent as the rules and the system around them.

Why it matters

Many Trust & Safety teams start by accident. Someone in support starts handling the worst reports, a founder makes the hard calls at night, and an engineer adds a ban button. That works until it doesn't: decisions depend on who handled the case, nobody owns the thresholds, and the person making 2am calls burns out or leaves with everything they knew.

Regulators now expect named accountability. In the UK, Ofcom's illegal content codes of practice recommend that every user-to-user service names a person accountable to its most senior governance body for complying with its illegal content duties and its reporting and complaints duties. In the EU, the Digital Services Act requires very large online platforms and search engines to set up a compliance function that is independent of their operational functions, with a head who reports directly to the management body (Article 41). How you structure the team is now part of how you show you're in control of the risk.

What good looks like

What you have, and what you can show, at each stage

Early

A founder or first safety hire covers trust and safety, usually with under a million users.

What you have

A safety lead who owns policy, the queue and escalations, a vendor or part-time reviewers, outside counsel you can call, a clinical partner for anyone who sees harmful content, and a named executive who owns safety risk.

What you can show

A one-page list of who decides what. Every person who reviews harmful content has support.

Growing

A dedicated safety team, millions of users, and new markets or features on the way.

What you have

Separate leads for policy and operations, an investigator for severe harm, an analyst who owns safety metrics, engineers assigned to safety tooling, a legal or compliance partner, and vendor staff calibrated with the in-house team.

What you can show

An owner for every policy area, threshold and legal duty. Quality and attrition reported for in-house and vendor teams on the same page.

At scale or regulated

Tens of millions of users, a heavily regulated sector, or extra duties as a very large platform under EU or UK law.

What you have

Specialist teams by harm area and region, a compliance function where the law requires one, a wellbeing lead, written career levels for reviewers, specialists and managers, and workforce planning tied to the risk assessment.

What you can show

The accountability regulators expect, written down. How many leads and specialists came up from the review team. Attrition by team and exposure level.

How to do it

7 steps

  1. Step 01

    Start from your risks

    The harms your product invites decide which skills you need first, so start from your risk assessment (chapter 2), not a template org chart. A few product traits change the first specialist hire more than anything else:

    If your product...You'll need early
    Lets strangers talk to each other, especially if children use itChild safety investigation, trained reviewers and a restricted escalation path
    Moves money or items between usersFraud and scam investigation, and access to payments data
    Hosts uploads at scaleOperations and vendor management, and hash matching for known illegal images
    Shares or exposes locationInvestigation of stalking and real-world threats, with a route to the authorities
    Generates content with AIPolicy that can be turned into prompts and evaluations, and people who can attack your own model
    Serves users in the EU or UK, or many childrenA legal or compliance partner who maps each duty to an owner

    Early on, one person covers several of these. That's fine, as long as you write down which hats each person wears. A gap is only dangerous when nobody knows it's there.

  2. Step 02

    Know the roles

    These are the roles a mature team ends up with. At first they're responsibilities, not job titles.

    RoleWhat they ownYou need it covered when
    PolicyThe written rules, internal guidance for reviewers, the enforcement ladder and the change log (chapter 4)Decisions depend on who handles the case
    OperationsQueues, response-time targets, capacity, vendors and getting guidance changes to every reviewer (chapter 7)There's more review than one person can do, or you sign a vendor
    InvestigationsNetworks of accounts, severe harm cases, evidence, referrals to law enforcement and cross-platform signal sharing (chapter 12)You see repeat offenders, organized abuse or child safety cases
    DataMetrics, thresholds, prevalence sampling and the analysis that shows safety's effect on retention (chapter 11)You need to set a threshold from evidence or prove the program works
    EngineeringReview tooling, detection, logging and the safety features inside the product (chapter 5)Your tools limit what you can do, for example when banning is the only lever
    Legal liaison or complianceEach legal duty mapped to an owner, reporting duties, and requests from regulators and law enforcement (chapter 16)A safety law applies to you, or a regulator writes
    WellbeingExposure limits, counseling, support tools and wellbeing standards for vendors (chapter 14)Anyone reviews harmful content: from day one, usually as a contracted clinical partner first

    Data and engineering are easy to borrow and hard to keep. A data analyst on loan for a quarter can build a dashboard. Nobody on loan owns a threshold. If the program has to prove it works, someone has to own the numbers. Since every threshold decision gets tested the day a regulator, a parent or a court asks why a specific account was or wasn't restricted, that someone needs to be on your team, or permanently assigned to it.

  3. Step 03

    Choose the order of your first five hires

    There's no single right order: it's a choice your risks should make. For a consumer product where users can contact each other, one order that works well is:

    1. A Trust & Safety lead. A generalist who can write a policy, run a queue, handle a severe escalation and explain a trade-off to an executive. Not a pure specialist: in the first year they'll do all of it. Chapter 3 covers what they do first.
    2. An operations lead or senior reviewer. Someone who builds the queues, targets and vendor relationship, so the lead isn't the queue.
    3. An investigator for severe harm. Trained in your highest risk, whether that's child safety, threats or organized fraud, and in preserving evidence and meeting reporting duties.
    4. A data analyst. At least half dedicated, owning the metrics and the evidence behind every threshold.
    5. A safety engineer. Or an engineer assigned for a long stretch, who owns the review tooling and builds safety features into the product.

    Then a policy specialist and a legal or compliance partner, as your markets and laws demand.

    Other orders are just as defensible, depending on what you face first:

    • Engineer earlier when safety tooling and product changes would do more than headcount, for example when reviewers are slowed by poor tools or the biggest risks can be designed out.
    • Investigator before operations when severe harm can't wait and a vendor can cover queue volume at first.
    • Data earlier when leadership needs proof before it will invest, or when you can't yet tell which harms are growing.

    Change the order when your risks are different. A marketplace or payments product is likely to need a fraud investigator before an operations lead. An AI product needs policy and evaluation skills early. A product launching in the EU or UK, or built for children, may need compliance help before an analyst, even if it's shared with Legal.

    Some things are better bought than hired at first: review volume and language coverage from a vendor, outside counsel for unusual legal questions, open-source tools such as ROOST's Coop, and a clinical partner for wellbeing. Put the clinical partner in place before anyone starts reviewing harmful content, not after the first person struggles.

  4. Step 04

    Decide where the team reports

    Every option has trade-offs, but I have a view. Trust & Safety should report into Product or directly to the executive level, and work across every team: Product, Engineering, Legal, Operations, Support and Comms. Making safety a function of Operations is a mistake, and I've written about why: You can't moderate your way out of a systems problem. At scale, you can't moderate your way out of a systems problem. Under Operations, safety gets judged on throughput and cost per ticket, and the work that prevents harm, policy, detection and getting into design reviews, gets starved. Here are the options, with their trade-offs:

    Reports intoStrengthsRisksWorks best when
    ProductIn the room at design review, with access to engineers and roadmapsSafety competes with growth goals set by the same leader, and safeguards that cost a feature its metric get traded awayThe product leader has safety outcomes in their own goals
    LegalRegulatory weight, close to reporting duties and law enforcementDrifts toward the legal minimum and away from product decisionsYou're heavily regulated, or building out compliance fast
    Operations or supportProcess discipline, vendor management and cost controlJudged on throughput and cost per ticket, like a cost center, with policy and detection starvedThe work is mostly high-volume review, and risk ownership sits elsewhere
    CEO or COOAuthority, fast escalation and a clear signal that safety mattersLittle executive time, and distance from daily product workSafety is core to the business, or you're rebuilding trust after a crisis

    Whichever you choose, three things matter more than the box on the org chart:

    • Name the executive who owns safety risk, and give the Trust & Safety lead a direct line to them for escalations, whatever the reporting line.
    • Write down decision rights. Who can change a policy, move a threshold, report to law enforcement, approve a public statement, or accept a known risk at launch. When a launch goes ahead with a known risk, the person who owns the product outcome signs off, with a date to revisit. Trust & Safety makes sure the risk is in front of them, and takes it to the executive who owns safety risk when it's too serious for one person.
    • Keep the team out of a throughput-only scorecard. If your leader judges you on tickets closed and cost per ticket, report outcomes beside them (chapter 11), or the program will be run as a cost center.

    Revisit the reporting line when you change stage. What suited five people may not suit fifty.

  5. Step 05

    Interview for judgment

    Most Trust & Safety work is a decision made with incomplete information, under time pressure, that someone will question later. Interview for that, not for knowledge of your rules, which anyone can learn.

    • Use a case, not trivia. Give a short written scenario with no clean answer: an account that's probably a minor on an adult feature, a report that could be a credible threat or a joke, a top creator breaking a rule. A decision from the incident tabletop works well.
    • Add a fact halfway through. Good candidates change their answer when the evidence changes, and explain why.
    • Score against a rubric agreed before the interviews, with at least two interviewers scoring separately.
    • For reviewer roles, use a work sample. A set of cases near the line, ten for example, with your written guidance open. You're testing how they read guidance, not what they remember.
    • Test languages in the language, with a native speaker.
    • For leadership roles, hand them a quarter where the numbers went up. The answer you want sounds like "Our numbers are up, and I'm not convinced we're safer," followed by what they'd check.

    You don't need graphic material to test judgment, so leave it out of interviews. But tell candidates plainly what the job involves, what they may see, and what support exists, before they accept.

  6. Step 06

    Build career paths, so people can grow without staying in the worst queues

    Expertise in this work is slow to build and quick to lose. Every leaver takes their training and judgment with them. Career paths are how you keep it.

    • Make the paths visible. Which ones you can offer depends on your size. Common ones: reviewer to quality and calibration specialist, to harm-area specialist, to investigator. Reviewer to lead, to operations manager. Sideways into policy, data, vendor management or product.
    • Offer a specialist track. Pay and titles for expertise, such as child safety investigation or a language market, without forcing people to manage.
    • Never make the worst queue the only route up. Time on high-exposure queues should be limited and rotated, and promotion shouldn't depend on staying there (chapter 14).
    • Pay for scarce skills. Languages, investigation and engineering that understands abuse are hard to hire, so pay to keep them.
    • Write down the levels. What's expected at each one, so promotion doesn't depend on who your manager is.

    Track attrition by team and exposure level, using the attrition and wellness-support usage metric. If people leave high-exposure queues much faster than others, the problem is the work design, not the hiring.

  7. Step 07

    Treat vendor staff as part of the team

    Where you use a vendor, its reviewers may make most of the decisions your users actually experience. If they're treated as a ticket count, the quality shows it.

    • Same guidance, at the same moment. A change reaches vendor teams when it reaches yours, not a week later through an account manager (chapter 7).
    • Same calibration. Vendor reviewers join the same calibration sessions, and their quality is measured the same way as yours.
    • Same escalation paths. A vendor reviewer who finds a child at risk at 3am reaches your on-call specialist as fast as an employee would.
    • Same wellbeing standards, written into the contract and checked (chapter 14).
    • A way to flag what they see. Vendor reviewers may spot a new abuse pattern before anyone else. Give them a route to report it, and tell them what happened.
    • Least access, logged. Give each reviewer access only to what the case needs, log every lookup, and get audit rights in the contract. The tabletop scenario The DMs nobody reported shows why.

    How you treat vendor staff in a crisis shapes quality afterwards. If a site goes dark, keep paying, ask how you can help, and expect experienced people back sooner. Chapter 9 covers choosing and contracting vendors.

Mistakes to avoid

And what to do instead

  1. 01

    Hiring reviewers before anyone owns policy

    Write the rules and the escalation path first, or a larger team just makes inconsistent decisions faster.

  2. 02

    Making the first hire a narrow specialist

    The first lead needs to write policy, run a queue and handle escalations. Add specialists second.

  3. 03

    Borrowing data and engineering indefinitely

    Somebody on your team has to own the numbers and the tools.

  4. 04

    Reporting into a line that measures only throughput

    Report outcomes beside activity, or the program will be run as a cost center.

  5. 05

    Leaving safety risk without an executive owner

    Name one, and write down who decides what.

  6. 06

    Interviewing for knowledge of the rules

    Test judgment with a case that has no clean answer.

  7. 07

    Making the worst queue the path to promotion

    Rotate exposure and offer a specialist track.

  8. 08

    Treating vendor reviewers as a ticket count

    Same guidance, calibration, escalation and wellbeing as your own team.

  9. 09

    Waiting for scale to arrange wellbeing support

    Put a clinical partner in place before the first person reviews harmful content.

Start from this template

Copy it, fill it in, make it yours

Template

Who wears which hat

Fill it in today, and again on a regular rhythm, every quarter for example.

RoleWho covers it nowShare of their timeBiggest gapHire, buy or borrow, and by when
Policy
Operations
Investigations
Data
Engineering
Legal liaison or compliance
Wellbeing

Template

Decision rights

Agree it with your executive sponsor and Legal.

DecisionWho proposesWho decidesWho must be consultedWhere it's recorded
New or changed policy
Threshold change
Accepting a known risk at launch
Report to law enforcement
Public statement about an incident

Template

Judgment interview scorecard

Score each separately on a short scale, for example 1 to 4, before you discuss the candidate: asks for the facts that matter; separates what to do now from what to decide later; weighs harm against fairness and privacy; changes their answer when the evidence changes; explains the decision so a user or regulator could follow it; knows when to escalate.

Do it with

Free tools and metrics that go with this chapter

Workbench tool

Program maturity

Rate the program in eight areas against the targets for your stage, and get a phased roadmap. Open content

Workbench tool

Incident tabletop

Use a scenario's decisions as interview cases, and rehearse The DMs nobody reported (a contractor on the review team misusing access).

Workbench tool

Vendor scorecard

Choose a moderation vendor on evidence, with reviewer wellness as a minimum. Open content

Metric

QA agreement rate

How often in-house and vendor reviewers make the same call an expert would, measured the same way for both.

Metric

Attrition and wellness-support usage

Whether you're keeping the people who keep users safe, by team and exposure level.

Further reading

Steven's posts on this topic, and sources worth the time

From Steven's writing · 1 post

  1. Essay

    You can't moderate your way out of a systems problem

    Treating Trust & Safety mainly as an operations function is a mistake. Reputation, history, age and behavior signals belong in one risk model, automation needs clear limits, and safety belongs in the product architecture from the start.

    Also in: 1. What Trust & Safety is for, 5. Detection and prevention, 15. Working with Product, Legal, Comms and leadership

Outside sources

Recent changes

4 changes to this chapter, newest first

The updates page has every change to the handbook, by date.

  1. AddedAlso: 1. What Trust & Safety is for, 5. Detection and prevention, 15. Working with Product, Legal, Comms and leadership

    Linked Steven's post on why treating Trust & Safety as an operations function is a mistake, and added it to chapter 8's view on where the team should report.

    Read the post: You can't moderate your way out of a systems problem →

  2. RevisedWith 18 other chapters

    Practical advice that varies by platform, such as cadences, sample sizes, targets and who owns what, is now set out as options with examples, so each team can choose what fits.

  3. RevisedWith 14 other chapters

    Added Steven's own calls from an interview: where Trust & Safety should report, what to automate first, the one number to track from day one, who makes the 2am call, and more. Practical choices that vary by platform are now laid out as options.

  4. DraftedWith 17 other chapters

    First full drafts of the other 18 chapters, built on Steven's posts, the handbook's principles and the Workbench's open content, with every legal and factual claim checked against its source. Stories from Steven's own work come next.