A brand runs two creator videos against each other. One gets a better cost per result, the winner is declared, and the next quarter of content is built on that conclusion.

In a market this size that conclusion is usually invented. Not wrong exactly, just unsupported, and a decision made on unsupported data feels exactly like a decision made on real data.

What a split test actually needs

A split test compares conversions, not views. To tell a real difference from a coincidence you need a meaningful number of conversions in each variant, and meaningful here means hundreds rather than dozens.

Most Portuguese campaigns do not produce that in a fortnight. They produce enough to see a difference, which is not the same as enough to trust one. Two videos that look twenty per cent apart on forty conversions each will frequently swap places if you run them another week.

That is not a reason to stop testing. It is a reason to stop deciding on it as though it were settled.

Three measures that work at low volume

The question that stops being asked

Every business has a question it answers over and over: does it fit, how long does it take, is it hard to install, do you deliver here.

If a video is doing its job, that question appears less often in messages, and that is measurable without any tool. Count the enquiries containing it for a month before and a month after. A drop is a real signal, because it does not depend on volume, it depends on the same question arriving less.

The reply rate on outbound

If the sales team sends a video in a first email, they will know within two weeks whether more people answer. That is a small sample and it is a comparison of the same list, the same sender and the same offer, which makes it far cleaner than an ad test.

Whether anyone reuses it

The most underrated signal in this business. Content that gets picked up by sales, by a distributor, by a retail partner or by an employee posting it themselves is content that solved a problem somebody had.

Nobody forwards a video they think is mediocre. Reuse is a judgement made by people with skin in the game, and it costs nothing to observe.

What you can measureWhether it works at low volumeWhat it tells you
Cost per result between two videosRarelyUsually noise, occasionally real
Watch time and completionYesWhether the opening works
The question disappearing from enquiriesYesWhether it answered something
Reply rate on the same email listYesWhether it earns attention
Internal reuse by sales or partnersYesWhether it is genuinely useful
Follower growthNoAlmost nothing about selling

The last row deserves saying plainly. Follower count is the easiest number to see and the least connected to whether anything was sold, and it is the one most often put in a report because it moves.

What not to compare

Do not compare a Portuguese campaign to the same campaign in a larger market. A country a fifth the size will always look worse in absolute numbers, and that comparison has ended more good local content than any creative decision.

Do not compare a video made for an ad account with one made for a product page. They had different jobs and only one of them was ever going to produce a cost per result.

And do not compare this month to the same month last year unless nothing else changed, which is never true. Compare the same content in the same place over a longer window, and accept that the window has to be longer here than the dashboards suggest.

The honest version of a report

For most brands in this market, a useful monthly report has four lines and no charts.

What we published. What we heard back, including from sales. What we would do again. What we are stopping. That is enough to steer by, and it is more truthful than a page of percentages calculated on samples too small to carry them.

The teams that improve fastest are not the ones with the best measurement. They are the ones who write down what they expected before publishing, and compare it afterwards to what happened.