In one minute

Buy what's common, build what's specific to you, and run open source where you have engineers to keep it runningHash matching, transcription and extra review capacity are rarely worth building. Your policies, thresholds and the signals only your product has usually are.
Score vendors on evidence, not demosWeight eight criteria, treat reviewer wellness and security as minimums, and ask every bidder the same questions.
Pilot on your own past cases before you signA paid pilot with blind quality checks on your data, hard cases included, tells you more than any sales deck.
Put quality, response times, wellbeing, data handling and exit into the contractAnything left out becomes a favor you have to ask for later.
Govern vendors with the same numbers you use in-houseQuality agreement, overturn rates and cost per decision, side by side, in a meeting that happens on a schedule.
The mistake to avoidChoosing on price and a polished demo, then finding out about wellness, language gaps and real quality once switching has become expensive.

Why it matters

Few Trust & Safety teams build everything themselves. They bring in outside help: people to review content, models to score it, services to check ages or match known abuse images. That help shapes what your users experience and what your reviewers go through, and you stay accountable for it. Under the EU's Digital Services Act, a service's transparency report has to describe the training and assistance given to the people who moderate content, and give indicators of accuracy and possible error rates for the automated tools it uses (Article 15(1)(c) and (e)). If a vendor handles EU users' personal data on your behalf, the GDPR requires a contract that limits it to your documented instructions, allows audits, and returns or deletes the data when the work ends (Article 28). You can hand over the work. You can't hand over the responsibility.

The options have also changed. ROOST (Robust Open Online Safety Tools), a nonprofit launched in February 2025, now publishes free, open-source safety tools, including a review console and a rules engine. "Build or buy" has a third answer, and it needs the same scrutiny as the other two.

What good looks like

What you have, and what you can show, at each stage

Early

A founder or first safety hire covers trust and safety, usually with under a million users.

What you have

A few tools bought or adopted for clear needs (hash matching of known child sexual abuse material if you host images or video, a reporting flow, perhaps one classifier), a short list of decisions you'll never outsource, and a data agreement and exit clause in every contract.

What you can show

Why each tool was chosen, what it costs each month, and who owns each vendor relationship.

Growing

A dedicated safety team, millions of users, and new markets or features on the way.

What you have

A scored request for proposal for review staffing, a paid pilot on your own cases before any new vendor, contracts with quality floors, response times by severity and wellbeing standards, and a weekly vendor review.

What you can show

Quality agreement and cost per decision for each vendor, measured by your own team, and overturn rates by vendor and policy area.

At scale or regulated

Tens of millions of users, a heavily regulated sector, or extra duties as a very large platform under EU or UK law.

What you have

Several vendors across languages and time zones calibrated to one answer key, at least two sites or vendors able to take each critical queue, open-source or in-house parts where they fit, quarterly reviews against the scorecard, audited wellbeing and data access, and an exit plan you've tested.

What you can show

Vendor quality by language and market, how vendors performed in your last surge, wellbeing audit results, and the training, accuracy and error-rate evidence your transparency report needs.

How to do it

8 steps

  1. Step 01

    Decide what to build, buy or run from open source

    Start with the problem, not the product. For each need, ask three questions. Is it specific to your platform, or does every platform have it? Do you have the data, scale or people to do it well? Who will keep it running in two years?

    BuildBuyOpen source
    Fits whenThe problem is specific to your product: your policies, your behavior signals, your risk scoresThe problem is common, and a vendor has data, scale or reach you don'tThe problem is common, you want control of the code and data, and you have engineers to run it
    ExamplesRules that use your account history, your enforcement ladder, thresholds for your harmsReview staffing, transcription and translation, age checks, threat intelligenceA review console, a rules engine, hash matching, open safety models
    You pay inEngineering time, for as long as it runsFees, lock-in and less controlHosting, security reviews, upgrades and engineering time
    Watch forRebuilding something you could have boughtA tool your team can't see inside or measureCalling it free because there's no invoice

    Some things stay with your team whatever you buy. You own the policy, the thresholds, the final call on severe harm, the decision to report to law enforcement (chapter 12) and the record of what went wrong. A vendor can run a queue. It can't own your risk.

    The same goes for AI tools. My default: where a wrong decision can't be reversed, or a user's safety is at risk, a human owns the call. A vendor's model can prepare the case. A person closes it.

    An early team usually buys or adopts almost everything and builds almost nothing, because its scarce resource is attention. A growing team starts building the parts that are specific to its harms. At scale, the question becomes how many vendors you can afford to depend on, and what happens when one of them goes dark.

    Where you draw the line depends on team size. Some ways it tends to fall:

    • One or two people: run free and open-source tools where they fit, buy what you can't staff (review volume, languages, hash matching), and spend your own time on policy and the worst cases.
    • A growing team: buy volume and keep judgment. Vendors handle scale and coverage, while policy, severe escalations and the hardest calls stay in-house.
    • At scale: own the core, meaning your review tooling, your data and your risk signals, and buy the specialist pieces, such as age assurance or threat intelligence, where a vendor's reach beats yours.

    These are starting points, not rules. Revisit the line every year as the team and the tools change.

  2. Step 02

    Know the kinds of vendor

    Each kind of vendor solves a different problem, and each needs a different test.

    KindWhat they doWhat to test before you sign
    Review staffingPeople who review content at scale, from large outsourcing firms or specialist Trust & Safety providersAgreement with your experts on your policies, wellness protections, native-language coverage, how fast they can add capacity
    AI moderationClassifiers and language-model tools that score text, images, audio or video, some against a policy you writePrecision and recall on your own hard cases, by policy area and language, and how they tell you about model changes
    Transcription and translationTurning voice and other languages into text that reviewers and classifiers can readAccuracy on slang, dialect and your community's own terms, plus speed
    Age assuranceAge estimation, ID checks and parental consentError rates around the ages that matter (13, 16 and 18), completion rates and how long data is kept. See chapter 6
    Grooming and child safety detectionSignals across conversations, and hash matching of known abuse imagesResults on your past confirmed cases and false alarms, who sees the material, and how reporting works
    Threat intelligenceMonitoring the spaces where raids, fraud and attacks are planned, and signals about known bad actorsRelevance to your platform, how much warning they give, and how they collect information lawfully

    Each often has a different owner, depending on your team: operations for review staffing, detection engineering for classifiers, and product, legal and privacy together for age assurance. Name the owner before the request for proposal goes out.

  3. Step 03

    Run a request for proposal that asks for evidence

    A request for proposal (RFP) is the document that asks vendors to bid and answer your questions. Write it around the outcomes you need, and ask for evidence rather than assurances.

    The vendor scorecard scores each vendor from 1 to 5 on eight criteria. These are its default weights:

    CriterionWeightThe question it answersMinimum
    Decision quality20Will their reviewers make the same calls your experts would?
    Reviewer wellness15Will they protect the people who look at the worst content?3
    Language and market coverage15Can they understand your users in every language and market?
    Security and privacy15Can you trust them with your users' data?3
    Surge capacity10Can they scale up fast when something goes wrong?
    Cost10What will it really cost once everything is included?
    Reporting transparency10Will you be able to see how they're really performing?
    Tooling and integration5Will they work inside your tools and share data back?

    A score of 2 or lower on wellness or security rules a vendor out, whatever its total. Change the weights to fit your program. A regulated service might weight security higher. A service in many markets might weight language coverage higher. Agree the weights before you read any bids, so nobody tunes them toward a favorite.

    Ask every bidder the same questions. A few that separate strong vendors from polished ones:

    • Will you run a paid pilot on our data with blind quality checks?
    • What is the maximum daily exposure to graphic content per moderator?
    • Which of our languages do you cover with native speakers, not translation?
    • How fast can you add 30% capacity, and what happened in your last surge?
    • Can you share decision data back to us through an API?
    • Will you share raw quality samples, not just summaries?

    The full set is in the RFP question bank. For AI vendors, add questions about how you'll be told before a model or threshold changes, where your data is processed and kept, and whether it's used to train their models.

    Have two or three people score each vendor on their own, then compare. Where scores differ by more than a point, talk it through. That's usually where a demo impressed one person and the evidence didn't support it.

  4. Step 04

    Pilot on your own past cases

    A vendor's numbers come from someone else's content and someone else's policies. Before you sign, test them on yours.

    Build the test set from your own decisions. Include clear violations and clear non-violations, but weight it toward the cases that are hard for you. The Workbench's classifier eval builds a set from a written rule across ten kinds of case: clear violation, clearly allowed, borderline, counter-speech, context-dependent, adversarial, news and education, hyperbole, off topic and other languages.

    Label it with your experts first. Have them label the same cases and measure how often they agree with each other. That's the ceiling. No vendor will beat your experts' agreement with each other, so don't set a pass mark above it.

    Set pass marks before you see results. Agreement or precision per policy area, handle time, and how the vendor escalates cases it isn't sure about.

    For review vendors, run a paid pilot, give every bidder the same cases, and have your team score the results without knowing which vendor made which call.

    For AI vendors, look at precision and recall by kind of case, not one blended figure. A blended number hides failures in low-volume, high-severity areas. A single high-confidence score can describe a credible threat, hyperbole between friends or a survivor telling their own story. Check which of those the tool can tell apart.

    Sign a data agreement before the first case leaves your systems, and send only what the test needs. Keep child sexual abuse material out of every pilot set. Handling it is tightly regulated (chapter 12).

  5. Step 05

    Write the terms that matter into the contract

    The contract is where the scorecard becomes enforceable. Legal usually drafts it, but the T&S owner decides what goes in.

    TermWhat to write in
    QualityAn agreement floor per policy area, measured by your team on a blind random sample, and what happens when it's missed
    Response timesTargets by severity that match your own, and when the clock starts: at the report or detection, not when a reviewer opens the case
    WellbeingHard daily exposure limits per person, licensed counseling during and after employment, blurring and grayscale by default, attrition reporting and your right to audit. Hold vendors to the same standard as your own team (chapter 14)
    Data handlingProcessor terms where the GDPR applies, managed devices or clean rooms (secured workspaces where data can't be copied or removed), logged access, retention limits, no training on your data without your agreement, and a written process for illegal content
    SurgeHow much extra capacity, how fast, and at what price
    ReportingDecision-level data back to you, raw quality samples, quality misses escalated within an agreed time, and the information your transparency report needs
    Change noticeFor AI tools, notice before any model or threshold change, and the version on every result, so you can compare before and after
    ExitNotice period, help with the handover, return or deletion of your data, and confirmation that your policies, training material and labeled data stay yours

    Two terms need special care. First, wellbeing. Data labeling and AI training review are moderation work too, so labeling vendors get the same standards and escalation paths. Second, illegal content. In the US, the duty to report child sexual exploitation to NCMEC falls on the provider under 18 U.S.C. § 2258A, so write down exactly what the vendor does when a reviewer finds it, and how fast that reaches your team.

  6. Step 06

    Manage the budget across several vendors

    Price per decision is the number on the invoice, not what a decision costs. The real cost includes rework, appeals, your team's quality checks, vendor management time and tooling. A cheaper vendor whose decisions get overturned twice as often isn't cheaper.

    Track cost per decision for each vendor and queue, and always show it next to QA agreement. Cost on its own rewards rushing.

    Plan for the day a vendor goes dark. The Workbench tabletop The empty review floor starts with a storm closing the site that handles most of your review. Its lessons: agree your priority tiers before you need them, move your in-house team onto the most severe work, and remember that the vendor's reviewers are part of your safety system. How you treat them in a crisis shapes quality afterward.

    By stage:

    • Early: one vendor is fine. Avoid volume minimums you can't hit and paid ramp time you didn't plan for.
    • Growing: forecast volume by queue, and keep the most severe work, such as child safety and credible threats, with your own team or your strongest vendor.
    • At scale: make sure every critical queue can be handled by at least two sites or vendors, and that your answer key and training travel with the work.

    Chapter 19 covers making the case for the money.

  7. Step 07

    Govern vendors after the contract is signed

    Most vendor problems show up after signing, slowly. Catch them with the same numbers you use in-house, in meetings that already exist (chapter 11). One rhythm that works:

    • Weekly: quality agreement, response times, backlog and graphic exposure hours, per vendor and queue.
    • Monthly: overturn rates by vendor and policy area, and precision for each AI tool.
    • Quarterly: a business review that rescores the vendor against the same scorecard you used to choose it. If the scores have slipped, say which ones and what happens next.

    Calibrate every team and vendor against one shared answer key (chapter 10). Send policy changes to vendor reviewers the same day as your own, and give them the same escalation paths. Vendor staff who get guidance a week late will make last week's calls.

    Hold AI tools to the same rules as your own automation. A tool doesn't take on a new policy area until its overturn rate is at or below human review. Every model or threshold change gets a baseline before launch and a check after, and rolls back if the rate moves past an agreed limit.

  8. Step 08

    Use free and open-source tools where you can run them

    Open-source tools can give a small team capabilities that used to need a large budget. They also move the work from a vendor's team to yours.

    ROOST publishes free, open-source tools for safety teams:

    • Coop is a review console you host yourself: queues, routing, automatic enforcement rules, and appeals that go to a review queue. It includes hash matching through Meta's open-source Hasher-Matcher-Actioner and NCMEC CyberTipline reporting, and it blurs images and video by default, with grayscale and muted video as settings.
    • Osprey is a rules engine and investigation console for real-time events, originally built at Discord to fight spam, abuse, botting and scripting. It's infrastructure you host and maintain, so it suits a team with engineers.
    • The ROOST Model Community works on open safety models with partners, including OpenAI's open-weight gpt-oss-safeguard, which classifies content against a policy you write, and Roblox's Sentinel.
    • awesome-safety-tools is a directory of open-source tools for online safety, maintained by ROOST.

    Other free options exist. Microsoft offers PhotoDNA, for matching known child sexual abuse images, free to qualified organizations.

    Judge open source the same way you judge a vendor. Pilot it on your own cases. Score its security, including who reviews your deployment. Name the engineer who owns upgrades and the person on call when it breaks. "Free" means no license fee, not no cost.

    The Workbench connects to some of this. Running on Coop maps each metric's data to where it lives in Coop, and the Works with ROOST page turns a pre-mortem into starter Osprey rules and a Coop setup checklist. The Workbench is independent and not affiliated with ROOST.

Mistakes to avoid

And what to do instead

  1. 01

    Choosing on price and a demo

    Weight the criteria first, set minimums for wellness and security, and score evidence from a pilot.

  2. 02

    Letting the vendor measure its own quality

    Your team draws the sample and scores it blind.

  3. 03

    Testing an AI tool on easy cases

    Build the test set from your hard cases, and read results by kind of case, policy area and language.

  4. 04

    Treating reviewer wellness as the vendor's problem

    Write the same standards you hold your own team to into the contract, and audit them.

  5. 05

    Signing without an exit

    Agree the notice period, the handover, and the return or deletion of your data before you need them.

  6. 06

    Putting every critical queue with one vendor

    Keep the most severe work in-house or split across sites, and rehearse losing one.

  7. 07

    Calling open source free

    Budget for hosting, security and the engineer who keeps it running.

  8. 08

    Outsourcing the judgment

    Vendors can run queues and models. Your team owns the policy, the thresholds and the calls that can't be undone.

Start from this template

Copy it, fill it in, make it yours

Template

Vendor scorecard

Agree weights before you read bids. Score 1 to 5, and note the evidence for each score.

CriterionWeightMinimumVendor AVendor BEvidence
Decision quality20
Reviewer wellness153
Language and market coverage15
Security and privacy153
Surge capacity10
Cost10
Reporting transparency10
Tooling and integration5

Template

Pilot plan

Fill it in and get it signed off before any vendor sees a case.

ItemYour answer
Cases: how many, from which period, and how many of each kind
Who labels them, and how often your experts agree with each other
Pass marks per policy area, set in advance
Data agreement signed, and what data leaves your systems
Who scores the results, without knowing which vendor made each call
Decision date, and who makes the call

Template

Contract checklist

Before signing, confirm the contract covers: a quality floor measured by your team; response times by severity, with an agreed start for the clock; daily exposure limits, counseling during and after employment, and your right to audit; processor terms, access logging and retention limits; a written process for illegal content; surge capacity and price; decision data and raw quality samples back to you; notice before model or threshold changes; and an exit with handover help and return or deletion of your data.

Do it with

Free tools and metrics that go with this chapter

Workbench tool

Vendor scorecard

Choose a moderation vendor on evidence: weighted criteria, rubrics, minimums and RFP questions. Open content

Workbench tool

Classifier eval

Build a labeled set of hard cases from a rule, run a classifier against it and see where it fails. Open content

Workbench tool

Works with ROOST

Free, open-source safety tools, and where the Workbench connects to them.

Workbench tool

Incident tabletop

Rehearse The empty review floor, where the site that handles most of your review goes dark for a week.

Metric

QA agreement rate

How often a vendor's decisions match your experts' on a blind random sample.

Metric

Cost per decision

The fully loaded cost of each decision, by queue and vendor, read next to quality.

Reference

Running on Coop

Where each metric's data lives if your review console is ROOST's Coop, field by field, and the gaps you'll need to log yourself.

Further reading

Steven's posts on this topic, and sources worth the time

Outside sources

Recent changes

3 changes to this chapter, newest first

The updates page has every change to the handbook, by date.

  1. RevisedWith 18 other chapters

    Practical advice that varies by platform, such as cadences, sample sizes, targets and who owns what, is now set out as options with examples, so each team can choose what fits.

  2. RevisedWith 14 other chapters

    Added Steven's own calls from an interview: where Trust & Safety should report, what to automate first, the one number to track from day one, who makes the 2am call, and more. Practical choices that vary by platform are now laid out as options.

  3. DraftedWith 17 other chapters

    First full drafts of the other 18 chapters, built on Steven's posts, the handbook's principles and the Workbench's open content, with every legal and factual claim checked against its source. Stories from Steven's own work come next.