Scoring 75,000 Profiles a Session
A case study in applied fraud detection: how I built a creator fraud-scoring pipeline that screens 75,000+ profiles per session (~415 per minute) and turned it into the vetting standard behind an influencer media-buying agency as StreamRadar Pro.
Status: All figures are from real production runs. This article covers the method but withholds the exact scoring features, weights and data sources.
The constraint: reach is fake-able, budgets are real
Influencer marketing runs on a simple trade: a brand pays a creator for access to an audience. The catch is that every metric the industry prices deals on (followers, viewers, engagement) can be manufactured for a few dollars: bought followers, viewbots, engagement pods, recycled accounts.
When I ran an influencer media-buying agency, this was not an abstract risk. A single funded deal with a viewbotted channel doesn't just waste the fee; it poisons the campaign data used to plan every next deal. Manual vetting (scrolling a channel, eyeballing chat) does not scale past a handful of creators and is embarrassingly easy to fool.
So I built the vetting into software, and made passing it a precondition for every euro of spend.
The pipeline
The system has three stages, each built for hostile input:
- Collection at scale. A high-throughput scraper pulls public profile and activity data at 75,000+ profiles per session, ~415 per minute, from platforms whose pages are custom, inconsistently structured, and change without notice. A GraphQL-based parser normalises the mess into one typed schema.
- Validation and quarantine. Records that fail schema validation, look truncated, or contradict themselves are quarantined rather than silently dropped or silently trusted. The pipeline assumes the source is messy and the consumer is downstream of money.
- Scoring. Each profile receives a 0 to 100 fraud score built from feature families that are hard to fake together: follower-ratio consistency, engagement shape, and growth velocity. Any one signal can be gamed; the cost of gaming all of them at once, coherently, over time, is what the score actually prices.
Designed for triage, not verdicts
A score is only useful if you know what to do with it, so the output is designed for decisions:
- Clear passes proceed to deal negotiation.
- Clear fails are dropped before anyone spends time on them.
- The murky middle is flagged for human review with the specific signals that triggered suspicion, because a fraud tool that pretends to be infallible just moves the fraud to its blind spots.
Two operating principles carried over from the rest of my work:
- Adversarial decay. Fraud signals rot: yesterday's detection rule becomes today's checklist for fraudsters. Scoring is re-validated on fresh data rather than trusted as a solved problem.
- Explainability. Every flag is reviewable. A risk decision you cannot explain is a risk decision you cannot defend, to a brand, an auditor, or yourself.
From internal tool to product
The scorer started as the agency's internal standard: every creator passed the fraud check before a deal was signed, and results were reported to brands in verified deposits rather than impressions. That workflow (verify before spend, measure after) was the agency's whole proposition.
I built it out as StreamRadar Pro, internal tooling for the agency's own vetting rather than something sold on. The tool served SkinBet Agency through 2026, and its validation and scoring discipline carried into the data work behind 21Core.
What this demonstrates
- End-to-end fraud detection: data collection → feature engineering → scoring → human-in-the-loop triage, in production, against adversarial input.
- Data engineering under hostility: high-throughput scraping, normalisation of wildly inconsistent sources, validation and quarantine by default.
- Commercial outcome: the score decided where the agency's budget went, and low-evidence creators were gated rather than oversold.
Related: Deposits, not impressions: the affiliate desk this powered · Catching silent data corruption: the same integrity discipline, applied to pipelines