When ROAS Lies: Building a Metrics System for a Luxury Brand That Sells Offline

|Diana Nekrasova

The context

The client is a European high-end fashion house in the ready-to-wear and couture segment, selling globally: its own atelier, cult pieces within its niche, an audience spanning Europe, the US, and the Gulf. A classic luxury fashion sales model: part of the orders go through the website, but the most valuable part of the business — bespoke couture — is closed offline, through a showroom appointment, personal correspondence, or a DM.

For performance marketing, that's a trap. Meta can only measure what happens inside its own funnel: a click, an add-to-cart, an on-site purchase. Here, most of the revenue comes from orders that never touch a cart at all — they're entered manually in the admin panel as draft orders, invisible to the pixel. ROAS can formally read zero while the business is growing and healthy. That's the point we were brought in at.

What we found going in

The first step was an audit of the client's historical ad accounts, covering two years. The picture was one we've seen across a lot of luxury brands that tried Meta Ads in-house or with a previous agency: a meaningful cumulative budget had been spent, and confirmed revenue in the system was zero.

The causes turned out to be structural, not creative. The pixel was misconfigured — campaigns had spent years sending traffic to an account literally labeled "do not use," and purchase events were firing without order value attached. Every ad in one of the accounts was missing primary text — they ran as visuals with no communication whatsoever. Ad set budgets sat at a fraction of what's needed for the algorithm to exit the learning phase. Essentially, two years of spend never produced a single signal usable for optimization.

The account structure itself was overcomplicated for the amount of data the business was actually generating: multiple overlapping ad sets, some audiences extremely narrow, and individual ad sets carrying long lists of interests with no clear hypothesis behind them. None of those choices is automatically wrong in isolation — the problem was the combination. A low-volume, high-ticket business was split into many small audience segments, even though there were nowhere near enough purchases for Meta to learn reliably from each of them. Instead of giving the algorithm more useful information, the structure was effectively dividing an already limited signal into even smaller pieces.

That was the real starting point — not "improve ROAS," but first build the infrastructure to measure anything at all.

Why standard ROAS doesn't work here

Once tracking was fixed, it became clear that a working pixel still doesn't solve the core problem. The brand doesn't sell on a "saw the ad → bought on the site" model — it sells on "saw the ad → got interested → came into the showroom or messaged the brand → bought offline." The pixel captures product views and sometimes an add-to-cart, but the purchase itself almost always happens outside its line of sight.

Reporting standard ROAS to the client under this model would be misleading, to both of us: well over 90% of revenue can show no formal connection to advertising, even when part of it is genuinely ad-driven. We needed a metrics language that measures intent quality, not transactions — one that works honestly whether the outcome is an online order or a showroom visit.

We tested the direct question too: will this audience buy online

Before fully moving away from the classic funnel, it was worth honestly testing the blunt hypothesis: would this audience buy online at all if campaigns were optimized directly for purchase, the way standard e-commerce works. The planned cost per purchase in this scenario was set at around $500 — and actual performance landed close to that plan.

The problem wasn't conversion cost itself, but scale. Meta's official benchmark for an ad set to exit the learning phase is roughly 50 optimization events over 7 days; the team also had its own, more modest internal minimum — about 20 purchases — used as a floor for a first directional read, not a final verdict, just an early signal on whether this direction was worth pursuing at all. Even hitting that internal minimum, given a segment AOV of $2.5–3k and a $500 target CPA, required a budget of at least $10k — several times what was allocated to the entire test campaign in a month. Reaching Meta's official 50-event threshold would have needed even more. Direct purchase optimization in this segment is mathematically possible, but it requires investment on a completely different order of magnitude than the current scale allowed. That math became one of the arguments for measuring earlier, more attainable intent signals instead of the purchase itself, given the budget actually available.

The metrics system, built from scratch

We designed a custom set of metrics on top of standard Meta and Shopify signals, each one built to answer a specific question a standard dashboard can't.

Creative quality is scored not by clicks but by a combination of full video views, post saves, and profile visits — because in luxury fashion, a save (not a like) is the strongest signal of "I plan to come back to this." A separate metric tracks warm-audience growth — people who engaged with the profile, watched most of a video, or visited the site over recent months — because at a modest ad budget, lookalike audiences simply don't reach usable size yet, and the first priority is building the retargeting pool. Another metric weighs add-to-carts against link clicks, showing how well creative and audience convert interest into purchase intent rather than just a click.

CPA wasn't usable as a working metric here — not because it "equals zero," but because too few purchases were observable in Meta for the number to mean anything: it's either uncomputable or technically blows out to values with no practical meaning. So instead of CPA, we introduced an internal directional metric — cost per intent signal. One caveat up front: this is not a replacement for revenue attribution, but an operating tool for decisions in conditions where purchase data is genuinely too thin. It also matters that a profile visit, a save, an add-to-cart, and an initiated checkout are events of very different intent strength — in this metric they are not summed as equal events. A profile visit reads as a weak intent signal, a save as medium, an add-to-cart as strong, and an initiated checkout as the strongest signal available short of the purchase itself. Each tier is tracked and weighted separately in the overall cost of signal — so the metric shows not an abstract "activity" score, but a rough distribution of the audience by depth of purchase readiness.

On top of that, we added rarer but important cuts: funnel velocity from product view to initiated checkout; branded search growth — not a standalone verdict on awareness, but a supporting demand signal that's only honestly read alongside PR activity, organic, and paid together; and the share of impressions landing in target geographies, because on a limited budget it's critical not to "leak" impressions into countries that generate engagement but never purchases.

We also defined a separate reporting language for the client. Instead of a formal ROAS, the monthly report started using lines like "the ads reached N unique people in the target cities this month," "this many people saved posts from the collection," "the warm retargeting audience grew to this size." For a brand that literally can't see most of its sales inside the ad account, this wasn't just reporting — it was a way to actually trust the channel.

How hypotheses were tested

This metrics system underpinned the launch structure itself. Instead of pushing budget straight into a broad audience, testing was split into phases with hard checkpoints: the first wave of creatives gets a small budget and is quickly filtered, survivors get a second, larger wave, and only one final "winner" scales onto the core budget. The switch from test mode to scale mode is triggered by data, not a calendar — as soon as cost per add-to-cart and its share of clicks cross a set threshold.

There was no automation behind this — campaigns were reviewed manually, every day, by the team. But discipline was anchored on three priority metrics that outweighed everything else: cost per click no higher than $1, cost per add-to-cart around $30 or below, and ATC Rate (add-to-carts as a share of clicks) — the primary KPI the team looked at first. As long as those held within range, a campaign was considered working even if everything else looked weak — and conversely, if the key metrics broke range, the campaign was paused or reworked, regardless of how good it looked on reach or engagement.

The bar wasn't uniform, either by geography or by audience type. For the US, Europe, and the UK, the team worked to those cost benchmarks; MENA ran on a different scale from the start — CPM, cost per click, and cost per add-to-cart all run structurally lower there — so "is this campaign working" was judged against its own, lower Gulf-region thresholds rather than one global standard. ATC Rate was treated the same way depending on audience temperature: around 0.5% was normal for cold traffic unfamiliar with the brand, while retargeting to an already-warm, brand-aware audience was expected to run several times higher — 2–3%.

Every hypothesis in the media plan — six were declared at launch, ranging from the performance of founder-led content to a 9:16 format versus the classic one — was tracked separately, with its own budget, expected outcome, and confirmation status. That turned the media plan from a static document into a live tracker: some hypotheses were confirmed and got more budget, others were judged inconclusive and killed in the first month, before the team burned through the full test budget on them.

Early results

Already in the first month of launch, overall business return relative to ad spend (blended, without splitting online vs. offline sales) came in well above plan — growth landed at more than one and a half times the target — even though the ad system itself hadn't recorded a single formal purchase. Engagement metrics — reach, impressions, cost per click — came in several times better than plan, and the combined value of items added to cart exceeded planned revenue several times over. That's not revenue, and it's not a guarantee a purchase follows — but it's a clear sign that the creative and offer were pulling the audience into meaningfully deeper product interaction than a simple click.

In the second month, budget started shifting toward the Gulf region based on data rather than intuition: one country was taking more than a quarter of the budget with almost no recorded checkouts — and based on the custom metrics, it was moved in time from a conversion role to an awareness role, freeing budget for formats and geographies with a confirmed intent signal. Retargeting to people who'd already shown interest was driving more than half of all add-to-cart value on a budget share several times smaller — and that finding is what got its budget increased in the next cycle.

The metrics system also helped honestly separate the effect of paid media from an organic spike: a celebrity placement of the brand at a major social event mid-month brought in several times more new followers than the entire month's ad budget — and rather than crediting that growth to the ad campaigns, the report showed the client exactly where the PR effect ended and paid traffic began.

What this gave the client

The main outcome of this project isn't a specific ROAS number — it's a working decision-making infrastructure for a business where Meta's standard tools structurally can't give an honest answer. The audit closed a two-year gap in the data and stopped budget that had been spent blind. The custom metrics system gave the client a language to discuss ad performance without treating online conversions as the only measure of success. And the discipline of hypothesis testing with hard checkpoints made it possible to shut down what wasn't working quickly, and put budget behind what the data actually confirmed, without delay.

For brands with an offline or hybrid sales model — showrooms, bespoke orders, DM-driven sales — this looks like a more honest, and more importantly, a more workable approach than trying to force standard ROAS onto a business that's simply built differently.


Want a measurement system for your own offline sales?

If most of your highest-value orders close outside the pixel's reach — a showroom, a DM, a phone call — standard ROAS will keep lying to you no matter how well the pixel is configured. Get a free marketing audit from SciGrowth and we'll show you exactly what your ad account can and can't measure, before you spend another dollar scaling on a number that might be noise.


FAQ

How do you measure ad performance when most sales happen offline?
You stop treating on-site purchase as the only valid outcome and build a custom set of intent signals instead — saves, profile visits, add-to-carts, initiated checkouts — each weighted by how strong a signal of purchase intent it actually is. None of these replace revenue attribution; they give you an honest, directional read when the purchase itself happens off-platform.
Why not just fix the pixel and report standard ROAS once tracking works?
Because a working pixel only fixes what happens inside Meta's own funnel — clicks, add-to-carts, on-site purchases. It still can't see a sale that closes in a showroom or over DM. For a brand where most revenue is structurally invisible to the pixel, standard ROAS stays misleading even with perfect tracking.
What is "cost per intent signal," and how is it different from CPA?
CPA measures cost per purchase — and becomes meaningless when too few purchases are observable to compute a stable number. Cost per intent signal instead tracks the cost of earning weighted engagement (a profile visit, a save, an add-to-cart, an initiated checkout), each counted at a different strength. It's not a replacement for revenue attribution — it's an operating tool for decisions when purchase data is genuinely too thin to use.