Works

Portaly AI scam-account blocking

Inserting a layer of AI judgement between the rules and the reviewers — and adding the path that turns a reviewer's decision back into data the system can use.

Product ManagerB2B Back-officeHuman–AI InteractionRisk Ops
00

Overview

Rapid user growth on the platform brought a flood of scam accounts, and the original two-stage flow — rule filtering, then manual review — could no longer carry the load: the rules could not catch obfuscated content, and everything they missed landed on a human. I led the redesign into a three-layer architecture — rules → AI judgement → human review → data feedback — and proposed the feedback mechanism itself. After launch, among directly-detected bans (excluding linked-account bans), 45.4% came through detection paths that did not exist before the redesign (28.9% from AI semantic judgement, 16.5% from the content restoration model), with a 0% appeal-reversal rate over the same period.

Key outcomes

  • Widened what the existing rules could see: added a content restoration model after the hard-rule stage, converting images and obfuscated content into scannable text and handing it back to the hard rules for a second pass.
  • Rebuilt review as a three-layer architecture: inserted AI semantic judgement between the existing hard rules and manual review, to handle suspicious accounts that rules struggle with and that require understanding the context of the content.
  • Built risk triage: routed cases to automatic ban, manual review or pass based on the verdict, raising coverage while keeping the risk of wrongful bans under control.
  • Verified what the new paths actually contributed: after launch, 45.4% of directly-detected bans came through the new detection paths — 28.9% from AI semantic judgement and 16.5% from the content restoration model — with an appeal-reversal rate of 0% over the same period.

Role & contribution

  • Problem research: audited the existing review flow, studied how comparable platforms handle this, and characterised the patterns of malicious accounts and the gaps in our system.
  • Solution design: delivered the system design for the three-layer architecture, the PRD, the interface flow and the UI.
  • Validation and iteration: confirmed feasibility and cost with engineering, tested with the reviewers themselves, and iterated to a version ready to build.
  • Impact analysis: reviewed the ban metrics after launch to understand how well it was performing.
Timeline
Dec 2025 – Mar 2026
Role
Product Manager
Team
1 PM + engineering + the review team
Tools
FigJam, Linear, LLM
01

Background & problem

As the platform's user base grew quickly, so did the number of scam accounts — fake customer service, accounts existing only to funnel traffic to a single external link, empty shell accounts. The existing review mechanism had two stages: a first layer of hard rules (malicious domains, keywords and so on), with everything not caught there falling through to manual review. At scale, that mechanism exposed three structural problems:

Coverage

Rules fail against obfuscated content — special characters, images or a reworded synonym are enough to slip past keyword matching.

Cost

Every grey-area account — uncertain but suspicious — fell through to a human, leaving the review team overloaded and the queue slow.

Scalability

New scam patterns had to be spotted by a human before a rule could be added by hand, so the system was permanently one step behind the attackers.

02

Defining the problem & goals

Problems to solve

  • How do you raise detection coverage for scam accounts without adding headcount?
  • How do you stop grey-area judgements from all falling to a human?
  • How do you let the system learn new scam patterns itself, rather than waiting passively for someone to notice them?

Project goals

  • Close the first layer's blind spot on special characters and images
  • Add a second layer of AI review to handle accounts that previously needed a human
  • Build a learning loop, accumulating manual review decisions as training signal
  • Build an early-warning mechanism, with the AI proactively surfacing suspicious patterns
→ As of May 2026, goals 1 and 2 are complete and live; goals 3 and 4 have a designed mechanism and a proposal but did not enter development within the project's timeframe. The design direction is described further down.
03

Research & iteration

Building a risk framework out of limited experience

External | The team had no AI review method or precedent it could simply adopt, so I built the first basis for judgement from external research and the platform's own data. On one side, I worked backwards from what comparable platforms publish. Reading transparency reports and other public material from link-in-bio platforms showed how violating content gets categorised, what methods detect it and how enforcement is tiered, which produced three rough conclusions:

  • Scams and spam are the single largest cause of suspension on platforms like ours, which confirmed where to start
  • Detection is never one model but a division of labour across text, links and images, with human oversight on top — which informed both the three-layer architecture and the content restoration model
  • Enforcement is tiered: remove the offending content first, suspend only in serious cases. And the rate at which appeals are upheld is not low, which shows wrongful bans are a real cost across the industry rather than a hypothetical risk

Internal | I gathered known malicious accounts from inside the platform and characterised the repeatable risk patterns visible in their account data, the semantics of their content and their behaviour.

Cross-referencing the two, I sorted those patterns into three groups by how you judge them:

  • Clear, stable signals that rules can intercept directly
  • Cases needing an understanding of context and meaning, which suit AI judgement
  • Ambiguous situations and edge cases, which still need a human

That classification became the basis for the three-layer review architecture.

Converge on an architecture first, then validate it fast internally

  1. I first drew the existing two-stage review flow as a diagram, clarifying the logic at each stage, how work was handed over and where AI could plausibly step in — then put forward the three-layer architecture and a first version of the flow.
  2. Rather than going straight to open-ended interviews, I chose to build a concrete draft so that cross-team discussion had something specific to react to: confirming technical feasibility, development cost and system constraints with engineering, and checking the actual decision logic, the real operating situations and how exceptions are handled with the review team.
  3. I kept adjusting the flow and the UI drafts on both sets of feedback, then asked actual reviewers to walk through it. Two rounds of iteration inside two weeks converged on a version engineering could build.
Flow diagram of the original review system and the three-layer architecture that replaced it (blurred for confidentiality)
Figure 1: the original review system and the three-layer architecture that replaced it (blurred for confidentiality; figure in Chinese)
04

The solution & the AI choices behind it

1. The solution: layered review with a feedback loop

The new flow inserts a layer of AI judgement between the rules and the humans, and adds a return path so that what a human decides feeds back into the system. Each layer handles the kind of judgement it is best at (see the diagram below):

  • Layer 1 | Rule filtering: handles clear-cut malice with fixed, enumerable signals — fast and explainable
  • Layer 2 | AI semantic judgement: handles the grey area that needs contextual understanding, running in two stages to control cost
  • Layer 3 | Manual review: handles low-confidence cases, and can override the first two layers in either direction
  • Feedback loop: every human override becomes a labelled data point, used to generate suggestions for new rules
De-identified architecture rules
Figure 2: the architecture rules, de-identified (figure in Chinese)

2. How the AI layer works

Before anything reaches layer 2, layer 1 gets a reinforcement of its own: a content restoration model converts special characters, images and foreign-language content into scannable text, so the existing hard rules can read what they previously could not.

Only accounts that genuinely need contextual understanding fall through to the AI layer, which then works on these principles:

  • Two stages, cheapest first. Look at the account's own profile page; only when that cannot settle the case does it go on to check what sits behind the external links. Scanning an outbound link costs far more per unit than scanning a profile, so it is designed to fire on a condition rather than by default.
  • Structured output. Every AI verdict has to come back in four parts:
OutputWhat it is for
Risk levelDecides whether the case is auto-actioned, sent to manual review, or passed
Violation categoryMaps to different enforcement actions, and lets us track how each type of scam rises or falls over time
Reasoning and cited evidencePoints to the specific content in the account that triggered the verdict
Account screenshotLets the reviewer preview the account's content, with a link through to the account itself

3. The implementation call: structure the existing standards as a prompt, rather than training a model

Working through cases with the review team, it became clear that the criteria for identifying a scam account could already be stated plainly. For example:

  • Is it impersonating official customer service?
  • Does it have nothing but a single external link?
  • Is it an empty shell account with no real content?
  • Is it using industry-specific patter to push users off the platform?

What the team lacked was the ability to apply those criteria at scale. A reviewer could explain exactly why an account needed banning; they just could not look at every account one by one. After evaluating the options, we chose for the first phase to structure the existing review criteria as a prompt rather than immediately investing in model training.

ApproachUpfront costSpeed of changeFits when…
Train our own classifierHigh — needs a large volume of labelled dataSlow — retraining required whenever the rules changeThe decision pattern is tacit and hard to state directly
Fine-tune an existing modelMedium — needs a reasonable amount of clean dataMediumStable labelled data has already accumulated in sufficient volume
Prompt an existing modelLowFast — edit the wording of the criteria directlyThe review criteria can already be stated explicitly

Scam tactics keep changing, so being able to adjust the criteria quickly and check the effect matters more than optimising a model once.

That choice does not rule out training a model later. The cases, the reasoning and the correct outcomes accumulating from human overrides gradually form usable labelled data; once there is enough of it and it is stable enough, the team can revisit fine-tuning or a dedicated classifier.

4. The reviewer's operating flow

In the interface, layer 3 comes down to movement in both directions between two lists:

Diagram of the reviewer's operating flow
Figure 3: the reviewer's operating flow (figure in Chinese)
  • Two lists, one set of controls: the blocked list and the safe list share the same date filter and keyword search, so reviewers never have to hold two mental models.
  • The sort order is the priority: the safe list is ordered by AI risk level, highest first, so a reviewer's time lands on the most suspicious accounts automatically — the most direct value the AI's output has in the interface. The blocked list leads with the most recent action, which makes it easy to check something immediately after the fact.
  • Overrides go both ways: banning and unbanning are really just an account moving between the two lists. The AI's verdict can be overturned in either direction — not only correcting a wrongful ban, but catching one that was missed.
  • Every decision is data: each override leaves a record, and those records are the source for monitoring wrongful bans and for adjusting the criteria later.
05

Risk & trade-offs

Crossing the AI's verdict against the truth gives four cases, each with its own risk:

AI says banAI says pass
Should be banned✓ Correct catch — high confidence means automatic action, and no reviewer time spent△ Missed — recoverable; a reviewer can ban it by hand from the safe list
Should not be banned✕ Wrongful ban — the costliest outcome: a real user is banned, trust is damaged and it is hard to win back✓ Correct pass — the overwhelming majority of accounts land here

When the AI is the one banning, those two errors do not cost the same. A missed ban carries risk, but there is still a chance to catch the account through a later mechanism. A wrongful ban hits a legitimate user immediately, and adds the cost of the appeal, the reversal and the repair to their trust.

So the automation strategy at launch was deliberately conservative, controlling the risk of wrongful bans rather than chasing the highest possible ban rate: only high-confidence malicious cases are banned automatically, every ambiguous case goes to a human, and even after a ban the appeal and manual-reversal path stays open as a way back.

Humans can also override the AI in either direction, and the system records the original verdict, the final outcome and the reason for the override. Those records serve both to monitor wrongful and missed bans, and as the data for tuning the prompt, the rules and the thresholds later.

06

Measuring impact

45.4%
of directly-detected bans came through paths that did not exist before the redesign
28.9%
share of bans contributed by AI semantic judgement
0%
appeal-reversal rate over the same period

How the metrics are defined

MetricCalculationWhat it tells us
New-path contributionBans via new paths ÷ directly-detected bans in the same periodHow much of the ban performance the new paths are responsible for
AI ban shareAI-decided bans ÷ directly-detected bans in the same periodHow much weight the AI actually carries in the decision chain
Scan hit rateBans decided ÷ scans at that stageCost efficiency, counted separately for the profile and outbound-link stages
Wrongful-ban rate (proxy)Bans reversed on appeal ÷ total bansWhether the conservative threshold needs adjusting, and whether wrongful bans are happening at scale

What actually happened

1 | Layer 1 | The content restoration model let the hard rules read obfuscated content: 0% → 10.9% → 16.5%

Zero before launch, 10.9% of directly-detected bans within two days of launch, and a steady 16.5% once the AI layer was running. Every account on this path had already survived the first keyword pass and was only caught once the restoration model had dealt with its special characters and foreign-language text — accounts the old architecture was structurally unlikely to catch. Not one hard rule was rewritten; the rules were simply given content they could read.

2 | Layer 2 | The AI layer took on 28.9% of directly-detected bans

It became a genuine second layer of risk decision-making rather than an advisory signal. The gap in cost efficiency between the two scan stages is stark:

StageShare of total scansHit rateShare of AI bans
Profile scan46%18.0%91.3%
Outbound-link scan54%1.5%8.7%

The overwhelming majority of malicious accounts can be decided at the first stage, from the profile alone (91.3% of AI bans), which validates the staged design — the model does not need to fetch detailed content for every account. Outbound-link scanning, by contrast, eats more than half the scan volume and produces under a tenth of the bans, at roughly 12 times the unit cost of a profile scan. The next step should be a stricter condition for entering it.

3 | The two together: 45.4% from new paths

Composition of directly-detected bansBaselineRestoration model liveRestoration model + AI
Caught after content restoration0%10.9%16.5%
AI semantic judgement0%0%28.9%
New paths combined0%10.9%45.4%

Once the AI was live, close to half of all direct detection came through paths that did not exist before the redesign. Accounts caught after content restoration only got there because the first hard-rule pass missed them; anything reaching AI semantic judgement had already survived two hard-rule passes.

4 | Risk control: 0% appeal-reversal rate over the same period

Every account actually banned automatically had a risk score in the highest band. Suspicious accounts in the middle ground are never auto-actioned and are kept for human judgement, while we watch the appeal-reversal rate and the accuracy of the model's bans in order to raise the threshold gradually.

5 | Human cost: reviewers reported handling time roughly halved

  • Interviewing the reviewers who use the system after launch, they said most of the accounts they used to open one by one had already been actioned or sorted by the AI, and that the same case volume took around half the time it used to.
  • This is the users' own estimate rather than a system measurement — the project never established a baseline for reviewer hours before launch, so there is nothing to cross-check it against.

What I proposed next: let the system catch up with the attackers on its own

The first two layers solved "the rules can't catch it" and "the reviewers are overloaded", but not the third structural problem — a new scam pattern still has to be spotted by a human first. Before I left I completed proposals for two mechanisms, neither of which entered development:

  • Feedback loop: the system already records the original verdict and the reason behind every human override, but at the time this was only used for monitoring. The proposal has the AI periodically analyse the two categories of error — should-have-been-banned and wrongly-banned — characterise what they have in common and produce suggested adjustments, so a misjudgement becomes a pattern to correct rather than a case to handle.
  • Early warning: an abnormal spike in registrations over a short window usually means one coordinated wave. The proposal has the AI, on detecting such an anomaly, scan that batch of accounts, characterise the shared tactic and summarise it for the review team to decide whether a new rule is warranted.

The principle is that the AI finds and characterises; the human keeps the decision — the asymmetry of anti-fraud risk makes fully automatic rule updates too dangerous, while "wait for someone to notice" leaves you permanently behind.

07

Reflection

1. Validating impact means counting cost, not just bans

  • Because the launch timeline was tight, the project never fully estimated daily scan volume, model call costs or the cost per ban before development. After launch we could really only assess it on the additional bans and the reviewer time saved, while monitoring what the AI was costing us.
  • Planning it again, I would try to establish a cost baseline before launch — daily scan cost, cost per effective ban, and time saved versus manual review — and add a cheap coarse risk filter so only accounts carrying some suspicious signal reach the model at all.
→ An AI project's impact cannot rest on proving it catches more. It also has to prove that catching them at this cost is worth it.

2. Prompt design presupposes that the team agrees on the criteria

  • Early on, engineering, the review team and I had different instincts about what makes an account suspicious. While the criterion is still "it looks off", there is no way to turn it into a consistent, testable prompt.
  • To converge, I used actually-banned accounts as shared cases, asked each role to judge them and explain their reasoning, then broke the disagreements down into concrete conditions — impersonation, content that does not match the link, shell accounts, intent to push users off-platform — and turned those into the basis for the model's judgement.
→ A prompt is not a piece of writing. It is the structuring of a product rule. Only once people can articulate the basis for a decision does the AI have any chance of applying it consistently.