Build 3 · Performance protection · Elabeh case study

Your dashboard can be green while TikTok is marking you red

TikTok grades every shop with a Performance Score, and that score decides what growth you're allowed: Campaigns, Flash Deals, TikTok-funded promotions, affiliate reach. Built for a 5-partner supplements brand selling on TikTok Shop, Shopify and Amazon.

Live measurement · window opened 22 Jul 2026 · results 19 Sep 2026

The bottleneck

The score that gates your growth recomputes daily. Most sellers check it weekly, if at all. Six days of possible lag, and the glance only answers where you stand. Not what's dragging it, how long you have, or what to do.

What it caught, week one

0.66%
TikTok's reading of our seller-fault return rate, failing a 0.33% bar. Our own dashboard read 0.34% for the same thing: green
31 days
of warning on a stock run-out, called the day the system went live and labelled a revenue risk, not a score risk
14/14
alert-path checks passed by test: dedupe, escalation, retry after a failed send. An alert system you can't trust is decoration
6 days
of detection lag in the old weekly check, gone. The score is read every morning, the drivers on every data pull

Before → after

Before
After
A human glanced at Seller Center weekly and hoped the week was quiet.
Nobody checks anything. The system messages a phone only when a decision is needed.
Our own derived metrics were trusted. One was wrong, in the comfortable direction.
TikTok's own figures are the source of truth, with our computation kept underneath as a drift check.
A flag meant ~30 minutes of digging to find which orders were behind it.
The advice card names the exact requests, TikTok's reason for each, and the official rule it's scored under.
A stockout would have been discovered as it happened.
Called a month out, with the gap to the next batch counted in days and the demand levers listed.

The catch that pays for the build

Our number said fine. Their number is the one that counts.

We compute seller-fault returns from our own order data. That figure read 0.34%, just inside TikTok's 0.33% bar. TikTok's own scoring of the same metric: 0.66%, failing, and the single biggest drag on a score that had drifted from 4.6 to 4.3 in three weeks. The gap had a boring cause, a filter that quietly excluded rejected refund requests, and rejecting a request does not remove it from TikTok's count.

The fix wasn't better arithmetic. It was making the platform's own payload the source of truth and demoting our computation to a cross-check that flags drift. Every derived metric in every dashboard has this failure mode: the platform grades you on its numbers, not yours, and the errors you don't notice are the ones in your favour.

Scored drivers table: every metric the shop is graded on, with TikTok's official value against its pass mark. Seller Fault Return and Refund Rate shows 0.65% against a 0.33% maximum, marked Breached, with our own cross-check figure of 0.61% underneath.
Every metric the score is built from, TikTok's figure against TikTok's bar. The breach reads 0.65% today against 0.66% at discovery; the cross-check underneath is our computation, kept only to expose drift. Captured 2026-07-24.

The assumption that inverted

The received wisdom is that fulfilment speed drives your score, so protecting it means obsessing over dispatch. For this shop that's wrong. It runs ~99% Fulfilled-by-TikTok, and TikTok's own rules exempt FBT orders from the delivery metrics. The real exposure is what's left on the seller: product-fault returns, negative reviews, response times, and compliance violations from what creators claim on your behalf.

That inversion changed what the system watches, and it reframed the coming stockout: an FBT stockout doesn't touch the score at all. It's a revenue problem, called a month early, with the gap to the next batch counted in days. A system that can't tell those two risks apart sends you chasing the wrong one.

Advice card: seller-fault returns and refunds at 0.65% against TikTok's 0.33% bar, with TikTok's own case count, reason breakdown, the two counted requests with dates and outcomes, the actions to take, and the official rule cited underneath.
What replaces the 30-minute dig: TikTok's count, the exact requests behind it, what to do, and the official rule it's scored under. Captured 2026-07-24.

The obvious objection

Why not just check Seller Center?

Seller Center is a scoreboard. It tells you your score when you go and look, and that's the whole service. It won't tell you the score started drifting on Tuesday, forecast a breach, count down a deadline, or connect a stock number to a revenue gap. Its rolling windows overwrite last month. And it will never tell you when its number and yours disagree, the exact failure that was costing us.

TikTok tells you your score. It doesn't tell you when to worry, how long you have, or what to do. That gap is the product.

There's a reason no tool category exists for this: for a UK shop there is no API access to the score. We probed the endpoint directly and got the region error to prove it. So the system reads it the only way available, daily, and verifies every driver against the shop's own order data.

Shop Performance Score and Account Health card: SPS 4.3 out of 5, Good, better than 84% of peers; AHR 212 with 0 violation points, low risk; a 30-day score trend showing 4.6 falling to 4.3; both readings marked fresh, read just now.
The daily reading: score, account health, and the 30-day trend TikTok itself asserts. Both values marked with their freshness, because a stale number shown as current is how dashboards lie. Captured 2026-07-24.

On a phone, unprompted

Telegram alert on a phone, shop name redacted. Urgent shop performance alert. Seller-fault returns and refunds at 0.65% against TikTok's 0.33% bar. Current 0.65%, bar max 0.33%. Affects the shop performance score. What to do, with TikTok's own count of 2 seller-fault cases against 309 delivered orders, the reason breakdown, and the two counted requests with dates.
The same breach, on a phone the moment it crossed the line. No dashboard opened, no scheduled check. This is what "near zero manual work" looks like in practice. Captured 2026-07-24.

What's being measured, in the open

The claim this has to earn

This page is published mid-measurement, deliberately. The window opened on 22 July 2026 and the results land here on 19 September. Two claims are on the line, and both have to survive:

1. The score holds or recovers. SPS stays at or above its 4.3 baseline, or the one failing metric comes back under TikTok's bar, with the daily readings as the record.

2. The manual work goes to near zero. No scheduled human checking. The alert ledger is the audit trail: every message sent, when it was acknowledged, what happened. A system that surfaces a score but still needs a human to look every morning is a dashboard, and the world has enough dashboards.

Partner reaction reserved for the measured result, 19 September 2026.

Co-founder

If you want to go further

See the instrument. I'll show you the live page from this business, breach and all, nothing staged.

Check your own exposure. Open Seller Center, note your score, then ask what it was last Tuesday. If you can't answer, you have the same gap we had. If the score has moved and you don't know why, that's the conversation worth having.