Case study
62% of online revenue was attributed to no source at all: a 90-day Shopify order reconciliation
6 real orders worth €1,905.00 against 3 transactions recorded in Google Analytics worth €720.00. The loss turned out to be chronic rather than recent, and we advised against buying the most expensive service we sell.
A technical outdoor apparel brand on Shopify, spending a few hundred euros a month on advertising, almost entirely on Instagram.
The constraint: they had just switched agencies, suspected tracking was broken, and wanted to know whether to rebuild it from scratch. We had read-only access, no way to touch the theme, and an ad spend too low for any infrastructure to pay for itself. The engagement had to end with a recommendation defensible on the numbers, not with a quote.
The result: 6 real orders worth €1,905.00 against 3 transactions recorded in Google Analytics worth €720.00. 62% of the period’s online revenue was attributed to no source at all, and the loss proved to be chronic across twelve months rather than recent.
Below is how we got there, what we ruled out along the way, and why we advised against buying the most expensive service we sell.
What this case covers
- how to tell whether a drop in conversions is real or a measurement problem
- how to establish which of the three platforms is telling the truth, instead of averaging them
- how to distinguish a recent break from a defect that was always there
- how to decide whether a tracking infrastructure pays for itself, before selling it
Contents
- The client and the starting point
- Order reconciliation against the back-office ledger
- Assessment of the Meta infrastructure
- GA4 analysis and source-data hygiene
- Economic assessment
- Method note on the numbers
- What didn’t work
- What we would do now
- Conclusion
The client and the starting point
A brand founded in late 2024, with a first line of merino wool active apparel and a second line in development. It sells on Shopify. Acquisition runs almost entirely through Instagram content, with a podcast and an editorial series pulling traffic.
The relationship with the previous agency had ended over weak conversions and suspicions about tracking. Since late April they had been running Meta in-house. No Google Ads spend, no server-side setup built by us, consent handled by a third-party platform.
The starting position, in numbers and with the source of each:
| Metric | Value | Source |
|---|---|---|
| Real online revenue, 90 days | €1,905.00 | Shopify order ledger |
| Transactions recorded in GA4, same period | €720.00 | GA4, 8 May / 4 Aug window |
| Meta spend, 28 days | €262.10 | Events Manager |
| Meta spend, 90 days | €1,485.30 | Ads Manager |
| Purchase events on Meta, 28 days | 0 | Events Manager |
| Total orders, 12 months | 92 | Shopify order ledger |
Order reconciliation: the comparison that decides how everything else reads
IN SHORT
What we did: put the Shopify order ledger next to the GA4 transactions, line by line, before opening any advertising platform.
Why: without that comparison, every subsequent analysis starts from a baseline that may already be 60% wrong.
What we found: 3 orders out of 6 with no user journey recorded, and the first- and last-visit fields empty at platform level.
The number for this section: €1,185.00 unattributed out of €1,905.00 in revenue, i.e. 62%.
The first decision was about order of operations, not method. We asked for the full order-ledger export before looking at Meta, GA4 or Google Ads.
The reason is that the three platforms only check each other if a fourth reference point exists that depends on none of them. The order ledger is that reference, because it is accounting rather than measurement.
Over the 90-day window the store had six real orders originating from the online store, worth €1,905.00. Another four orders were in-person sales, correctly absent from the web reports and out of scope.
GA4 showed three, worth €720.00.
| Order | Date | Total | Present in GA4 |
|---|---|---|---|
| #1085 | 10 May | €63.00 | yes |
| #1089 | 1 Jun | €360.00 | yes |
| #1090 | 3 Jul | €297.00 | yes |
| #1091 | 10 Jul | €280.00 | no |
| #1092 | 24 Jul | €675.00 | no |
| #1093 | 25 Jul | €230.00 | no |
Every tracked order stopped at 3 July. The immediate reading was a recent break, and that would have been the convenient conclusion.
We extended the window to twelve months to find the exact date of the failure. Out of 92 total orders, excluding 48 test or import orders and 4 in-person sales, roughly 40 real orders remained. Of those 40, 22 had the user journey populated, i.e. 55%.
The 50% in the audit window was in line with the store’s own average. The break date did not exist, because the loss is chronic and spans the entire life of the store.
The rate does degrade over time, though: 85% in December 2025, 50% in January 2026, 20% in February 2026. From March onwards monthly volumes are too low to be interpretable.
On the untracked orders, the first- and last-visit fields recorded by Shopify are null. And the platform’s own internal analytics, which depends on neither Google nor Meta, attributes only 3 orders out of 6 to a session.
It is the platform that fails to associate the order with a browsing session (the link between who bought and the visit they came from). With that link missing upstream, every downstream system is left without raw material. It is the only explanation compatible with a simultaneous loss across three independent systems.
What worked
Asking for the order ledger before the platforms changed the verdict of the whole engagement. With only the three GA4 transactions we would have written an audit on a baseline reduced to 38% of reality, and every percentage calculated on top of it would have been formally correct and substantially false.
Assessment of the Meta infrastructure: healthy, and already more effective than they thought
IN SHORT
What we did: verified deduplication between the browser channel and the server channel event by event, and measured the server’s net contribution.
Why: before proposing a server-side infrastructure you need to know whether one is already running and how much it delivers.
What we found: both channels active from the native apps, deduplication above threshold, no Custom Pixel configured.
The number for this section: 23.9% of the signal arrives only through the server channel, without anyone having set it up.
Events received over 28 days: 7,352 PageView, 5,705 ViewContent, 7 Search, 5 InitiateCheckout, 2 AddToCart, 0 Purchase.
The zero on Purchase is a consequence of the previous section, not a flaw in the integration. With a baseline loss of around 45-50% at session level and three real orders in the window, the probability of observing zero by chance is around 12%. There is no need to posit a specific failure.
The prefix on the event identifiers indicates the platform’s native server-side implementation. We checked the panel: no Custom Pixel configured. Both channels come from the native apps.
Split: browser 5,649 against server 7,422. Across 13,071 events received, matchable pairs number 5,648 and unique events roughly 7,423, which means the headline figure in Events Manager is 1.76 times real traffic. That is expected behaviour, not an error, but you need to know it before reading any number on that screen.
Deduplication rate 78% (the share of events the system recognises as already received from the other channel, and therefore does not count twice), against an operational reference threshold of 75%. Identifier present on all events and shared across both channels, checked event by event.
The server channel’s net contribution is 1,774 events that exist only there, i.e. 23.9% of the overall signal and 31.7% on ViewContent alone. In detail: PageView 4,025 against 3,327, plus 21%. ViewContent 3,389 against 2,316, plus 46%.
The value is amplified by the composition of the traffic: 64% Safari and 27% Instagram in-app browser on Android, together over 90%. Those are the two environments where the browser channel loses the most.
Event match quality at 5.9 on PageView and ViewContent. IP address and user agent present at 100%, internal user identifier at 100%, ad click identifier at 68.6% and 73.0%. Email, phone, name and city absent, which is a property of those events rather than a defect: on a product view that data does not exist yet.
One compliance note that has to be stated. The native consent restriction is active for Italy only. On Italian traffic that 23.9% is technical recovery within the population that consented. On foreign traffic, which is the larger share, it may include non-consenting users, and that needs checking in the consent platform’s panel.
What worked
Measuring the server channel before proposing it avoided selling something the client already had running. And it produced the figure that later overturned the economic recommendation.
GA4 analysis and source-data hygiene
IN SHORT
What we did: inspected the payload of the e-commerce events and cross-checked line totals against transaction totals.
Why: dirty data at source produces clean, wrong reports, which is the worst possible condition.
What we found: discount not propagated to the line price, test orders in production, a single market configured.
The number for this section: 1 order out of 92 carries tracking parameters on the ads.
Across all three recorded transactions the sum of line prices is €800.00 against a transaction value of €720.00. A ratio of exactly 0.9 on all three.
We established the cause in the order ledger: a 10% discount code applied to all three. The platform allocates the discount to the lines correctly, with amounts visible order by order, but that allocation is not propagated to the price field sent to GA4. The compare-at price field is empty across all variants, which rules out the alternative explanation.
One of the two codes used does not appear among the active ones in the panel. Origin to be verified with the client.
48 test or import orders sit in production, concentrated on two days (26 November 2025 and 22 April 2026), with repeated amounts of 1, 18 and 45 euros, no address, and in part with first access from a password-protected page. They enter every native report and distort average order value, customer counts, and any future lifetime-value calculation.
The heaviest finding concerns marketing, not tracking. Across 92 orders in twelve months, only one carries UTM parameters (the labels that say which campaign a visit came from). The ads are boosted posts with no parameters. The consequence is that no report, however well built, can establish whether ad spend produces orders. You fix it by tagging campaigns, not by migrating the stack.
Finally, catalogue and markets: a single market configured and a generic default country code on the catalogue. Product titles contain line-break tags, which the Merchant Center channel does not accept.
What worked
Reading the payload instead of inferring it from totals. A 0.9 ratio identical across three transactions is the kind of regularity an average never shows.
Economic assessment: when the infrastructure does not pay for itself
IN SHORT
What we did: compared the recurring cost of a full server-side infrastructure with the store’s actual ad spend.
Why: the recommendation had to be defensible even when it took the most profitable work off our own table.
What we found: spend is an order of magnitude below the break-even threshold.
The number for this section: €262.10 of spend over 28 days, against a reference threshold of €1,000-1,500 per month.
Meta spend over 28 days is €262.10, which generated 99,505 impressions and 4,226 clicks, at a cost per thousand impressions of €2.63 and a click rate of 4.25%. Over 90 days spend rises to €1,485.30. Google Ads is connected but inactive, with zero spend.
The reasoning is arithmetic. A full server-side infrastructure has a setup cost and a recurring monthly cost. For it to pay for itself, the additional signal has to translate into campaign optimisation, and for that to happen you need a spend volume that generates enough conversions for the algorithm to work with.
With these numbers that condition is absent. On top of that, as measured in the previous section, the native server channel is already active and delivers 23.9% of the signal at no cost.
The final recommendation was structured on three levels, ordered by impact-to-cost ratio.
First, fix the session loss at platform level. It is the only intervention that acts on the cause identified in the first section, and the only one that unblocks everything else.
Second, tag the campaigns. Zero cost, immediate effect on the very possibility of measuring return on spend.
Third, no full infrastructure, with the spend threshold for reassessing it written plainly in the document.
What worked
Putting the economic assessment inside the audit instead of inside a separate quote. The client received the criterion for turning down our service, and that made the two recommendations we did make credible.
Method note on the numbers
Declared source of truth: the Shopify order ledger. Not GA4, not Events Manager. The ledger is accounting, the platforms are measurement, and when they diverge accounting wins.
| Published number | Source | Window | Model |
|---|---|---|---|
| 6 orders, €1,905.00 | Shopify order ledger | 90 days, 8 May / 4 Aug | Online Store orders, point of sale excluded |
| 3 transactions, €720.00 | GA4 | 90 days, same window | purchase event, last non-direct click |
| 55% journeys populated | Shopify order ledger | 12 months, 1 Sep 2025 / 6 Aug 2026 | on ~40 real orders, excluding 48 test and 4 point of sale |
| Deduplication 78% | Events Manager | 28 days, 8 Jul / 4 Aug | matchable pairs over events received |
| 23.9% server contribution | Events Manager | 28 days, same window | server-only events over unique events |
| Match quality 5.9 | Events Manager | 28 days, same window | average across PageView and ViewContent |
| €262.10 of spend | Events Manager | 28 days, same window | actual spend, not budget |
The gap between platform data and back-office data is 62% over the 90-day window and roughly 45% on the twelve-month average. We have not averaged or rounded them, because they are two measurements over two different windows and must be read separately.
What we are not able to measure, and with what degree of certainty.
The cause of the missing association between session and order is established at system level, meaning it is the platform and not GA4 or the pixel, but it is estimable with medium confidence at mechanism level. The two compatible explanations are consent refusal, given that the restriction is active for Italy only and every order in the period was Italian, and session fragility in the Instagram in-app browser, which since January has been the dominant share of traffic. Separating them requires an active test on session-cookie behaviour on refusal, which needs write access and was out of scope.
The effect of the loss on campaigns cannot be quantified with the available data, because out of 92 orders only one has tracking parameters. That is a limitation of the source data rather than the analysis, and it can be fixed in a week by tagging campaigns.
What didn’t work
We raised a blocking issue on a hypothesis that later proved false. In the first draft we flagged as a blocker the suspicion that the add-to-cart event was not firing with the variant selector, inferring it from the anomalous ratio between product views and add-to-carts. An active test on the store falsified the hypothesis, because the event fires correctly. The conclusion had been deduced from numbers instead of verified in the interface. What we did instead: we introduced an explicit hierarchy between structural findings read from configuration, active findings verified by test, and accounting reconciliations, and every verdict now states which level it belongs to.
We lost time looking at the platforms before the order ledger. The first version of the analysis was organised by channel, with the accounting reconciliation at the end as a final check. In that order, the six-real-orders figure arrived too late to steer the work, and several intermediate conclusions had to be rewritten. What we did instead: reconciliation against the back office became section one of every audit we run, ahead of any platform.
We read the browser-versus-server split on the wrong baseline. In an intermediate draft the two channels came out as 4,187 against 3,054, figures taken from a view that did not cover the same scope. Corrected to 5,649 against 7,422 after re-checking. What we did instead: every pair of numbers put into relation inside an audit now carries the window and the scope written next to it, and if they don’t match we publish a single number.
What we would do now
An active test on session-cookie behaviour when consent is refused. It is the only way to separate the two remaining explanations for the loss. Requires write access and half a day.
Full campaign tagging and a 30-day measurement. With parameters in place, in a month’s time there will for the first time be a basis for saying whether spend produces orders. It is the best impact-to-cost intervention in the whole document.
Clean-up of the 48 test orders. As long as they stay in production, every average-order-value and lifetime-value calculation starts out corrupted, including the one needed to decide whether the full infrastructure makes sense.
Reassessment of the threshold in six months. If monthly spend exceeds €1,000 the economics change, and the infrastructure comes back on the table with different numbers.
Compliance check on foreign traffic. The native consent restriction is active for Italy only. On the rest of the traffic, what the consent platform does needs checking.
Conclusion
62% of the period’s online revenue, that is €1,185.00 out of €1,905.00, was attributed to no source at all. Not because of a recent failure, but because of a defect in the association between order and session that has been with the store since it opened.
What this engagement leaves the client, picking the four opening points back up.
How to tell whether a drop is real or measurement: you put the order ledger next to the reports before looking at any platform. If the two lists don’t match, the problem is measurement until proven otherwise.
How to establish which platform tells the truth: you don’t average them, you compare all of them against a source that depends on none of them.
How to distinguish a recent break from a chronic defect: you widen the window until you find the date of the failure. If the date doesn’t exist, the failure never started, because it was always there.
How to decide whether an infrastructure pays for itself: you compare the recurring cost with actual ad spend, and you write down the threshold above which to reassess.
When this work makes sense, and when it doesn’t
It makes sense if you have more than one source telling you how much you sold and at least once you have wondered which one to believe, and if someone in the company has to answer for those numbers in front of a partner or an investor.
It doesn’t make sense if you are looking for a guarantee that two platforms’ numbers will match, because they never will and part of the gap is structural. And it doesn’t make sense if you want the work to start before an audit, because on an undiagnosed system the scope cannot be estimated and a badly estimated scope costs both of us.
The first step is a reconciliation between your order ledger and your platforms over 90 days. Ten working days, one document, and the freedom to do nothing else.