# Xoot — full content Full text of every blog post on https://www.xoot.uk, provided for LLM crawlers that prefer plain text over JavaScript-rendered HTML. # Fix the product before you fix the bidding URL: https://www.xoot.uk/blog/fashion-returns-tug-of-war-fix-basics-before-ad-spend Published: 2026-07-21 Category: Fashion e-commerce **Fashion returns are eating ad budgets alive. The fashionable answer is to get clever with bidding and audiences. The right answer is usually more boring than that.** Fashion e-commerce has a returns problem it would rather not look at directly. In 2020, US shoppers sent back an estimated $428 billion of goods, around 10.6% of everything they bought, with clothing leading the way. For online fashion specifically, return rates hover around 25%. One order in four goes back. That's not a logistics footnote. Every returned dress was paid for twice: once by the ad budget that won the sale, and again by the reverse logistics that unwound it. McKinsey puts the average cost of processing a single return at about two thirds of the item's price once you count shipping, handling and restocking, and that's before the stock misses its full-price window and gets marked down. Little wonder 83% of retailers describe returns as a serious threat to profitability. Faced with numbers like that, the tempting move for a marketing team is to optimise around the problem. Tag high-return products in the feed and bid down. Exclude serial returners from prospecting. Feed refund data back into the platforms. All of these are good ideas, and I'll come back to them, because done properly they work. But there's a catch that gets skipped in most of the how-to content: if the returns are being caused by operational problems, such as inconsistent sizing, misleading photography or patchy quality control, then adjusting your marketing is a bandage on a leaky pipe. The water keeps coming. You've just stopped watching that bit of the floor. ## When marketing and operations pull in opposite directions Here's the trap in miniature. A particular dress has a 40% return rate. Marketing spots it, quite reasonably concludes the thing is unprofitable to advertise, and cuts spend. Sales fall, returns fall with them, blended ROAS ticks up, and everyone moves on. Nobody asked why 40% of buyers were sending it back. Maybe the sizing was mis-marked. Maybe the studio lighting made a dull maroon look bright red. Whatever the cause, it hasn't gone anywhere. Customers still arriving through organic search or email buy the dress, get the same unpleasant surprise and return it, except now there's less revenue around to absorb the cost. Marketing saved some return processing and threw away the legitimate sales that would have stuck if the product page had told the truth. Two departments cancelled each other out. This happens because, in most businesses, returns don't belong to anyone. McKinsey found that 58% of retailers admit no single team owns the problem, even though preventing returns and recovering value from them cuts across merchandising, e-commerce, marketing and finance. Marketing optimises the metrics it can see, which are gross ones. The product team never learns how much a sizing inconsistency is costing in wasted ad spend, because that cost lands in someone else's report. Most retailers can't even break returns down by root cause at product level, and as McKinsey rather drily notes, causes you can't see are causes you can't fix. So before touching the bidding, plug the bucket. ## Why things actually come back The reasons customers return clothes are well documented and mostly mundane. In one survey of apparel retailers, size and fit accounted for 53% of returns, colour or appearance not matching expectations for 16%, damage or defects 10%, price disputes 9%, the feel of the material 7%, and slow delivery 5%. Other studies put fit-related returns anywhere between 53% and 70%. Nearly all of it reduces to the same thing: the product that arrived wasn't the product the customer thought they'd ordered. Sizing is the big one, and it runs deeper than "brands vary". One brand's Medium is another's Small; fine, everyone knows that. Less forgivable is variation within a single brand, where two "Medium" shirts from the same retailer fit differently because they came off different cuts or out of different factories. A BBC investigation found that H&M trousers in black and beige, nominally the same size, fitted very differently, probably because darker fabrics go through extra treatment that changes their stretch. Layer vanity sizing on top (a 38-inch chest is a Small here and a Medium there) and customers respond rationally: they order two sizes and send one back. Bracketing isn't customer misbehaviour. It's a workaround for an industry that can't standardise itself. Photography is next. If the shade of the dress on the site doesn't match the one in the parcel, whether through studio lighting, screen calibration or a dye batch that drifted, that's a return. A newer version of the same problem is AI-generated model imagery. Dressing a virtual model costs a fraction of a photoshoot, but generators still struggle with fabric texture and drape, so a stiff blouse can look fluid on the page and disappoint in person. Plenty of retailers also still sell garments off a single flat shot on a hanger, which tells the customer almost nothing about length, movement or fit. Feel and quality round it out. A customer can't touch the fabric, so they infer weight, softness and stretch from photos and adjectives, and when a jumper that looked thick and cosy turns up thin and scratchy, back it goes. Multi-factory production quietly makes this worse: the "same" shirt cut from different fabric lots can fit and feel like two different products, and the customer experiences it as a quality lottery. Charging them a return fee when the product was at fault is a particularly efficient way to lose them for good. None of this is exotic, which is rather the point. The majority of fashion returns trace back to fixable gaps in sizing consistency, content accuracy and quality control, not to fickle customers. ## What fixing it at the source looks like The fixes are unglamorous. Standardise your size specifications and hold factories to them. Publish actual garment measurements, not just S/M/L, so people can compare against something they own. Photograph products on real bodies, check colour against the physical item under neutral light, and if you use AI imagery, review it for honesty and keep at least one real photograph in the set. Write fabric composition, weight and stretch into the description. If a style runs short or a colour runs snug, say so on the page; better to lose the sale than win the sale and the return. Video helps more than almost anything, because it shows true colour in moving light and how the fabric actually behaves. One study found shoppers 64% more likely to buy after watching a product video, so this isn't even a trade-off between conversion and returns. It improves both. The more interesting development is retailers pointing machine learning at the problem, and the results are worth a pause. H&M, whose online returns were dominated by fit issues, built AI-powered virtual fitting rooms. Customers enter their measurements or scan themselves, and a 3D model shows how a garment will sit on their body, accounting for cut, fit type and fabric stretch. Early rollouts showed meaningful reductions in returns, with lower reverse-logistics costs and a sustainability story thrown in. Better still, the try-on data flows backwards into the business: if a particular cut fits badly across thousands of avatars, the design team hears about it before the returns arrive. Playful Promises, a UK lingerie brand, went a similar route with Prime AI's fit prediction after ordinary size charts proved useless for bras, which is about the least forgiving fit category there is. The system learns from real purchase and return data, factors in each garment's cut, material and even colour (dye processes can make a black lace bra fit more snugly than the same style in nude), and recommends a size for that specific item rather than assuming your usual size travels. Returns fell 27%, conversion rose 18%, and sizing queries to customer service dropped. The model keeps improving as more outcomes feed in, which is exactly the shape you want: a system that learns from past returns to prevent future ones. Zalando runs a lighter-touch version, mining customer feedback at scale to put "size flags" on product pages ("Runs small, consider sizing up"), reportedly cutting size-related returns by around 10%. Notice the shape of these wins. Nobody reduced returns by suppressing demand. They reduced returns by helping people buy the right thing, and sales rose as a side effect. Industry surveys suggest 85% of apparel retailers are now using or planning virtual fitting tools, and 80% of those running a size recommender report higher conversion. Compare that with the alternative timeline where H&M just quietly bids down on customers who return bras. Returns fall because sales fall. The sizing stays inconsistent, the customers stay annoyed, and the "optimisation" is a slow leak dressed up as efficiency. ## Get the data before you get clever You can't fix what you can't diagnose, and most return data is under-used at best. Reason codes at the returns portal are the start: too small, too large, looks different from the photo, faulty, changed my mind, arrived too late. On their own they're ambiguous. "Too small" might mean the garment runs small, or that the customer guessed wrong. The trick is triangulation: if 80% of a dress's returns say "too small" and a decent chunk of those customers reordered the next size up and kept it, that's a sizing problem, not a customer problem, and the size guide needs changing. Free-text comments are worth collecting too ("returning because the colour is much greener than on the site"), and worth mining, along with reviews and social mentions. Language models have made it cheap to comb unstructured feedback for themes at a scale no merchandising team could read manually. Then slice return rates by attribute. By factory, by supplier, by fabric, by collection, by photography style. You'll find patterns you didn't expect, like every product shot on a mannequin rather than a model returning above average, and each pattern is an instruction. The other half of the diagnostic job is joining returns to marketing data, and this is where most companies quietly fail: the e-commerce team tracks return reasons, marketing tracks conversions, and nobody has merged the tables. Until they're merged, your ROAS is a work of fiction. A campaign that spent £1,000 to drive £5,000 looks like a 5x return; if 30% of that revenue came back, it actually drove £3,500 of kept revenue, and once you subtract the cost of processing those returns it may not have broken even. Two thirds of retailers say they have a strategy for improving the economics of returns, but without unit-level data joined across systems they can't see what anything actually costs. The order ID is the key that links a click to a sale to a return to a reason; if your data model can't make that join, that's project number one. Track it over time as well. If the fixes are working, the reason codes should move: fit returns fall after you launch a size recommender, "not as described" falls after you add video. That feedback loop tells you you're fixing the right things, and it tells you when the preventable returns are mostly gone, which is the moment the remaining returns data becomes safe to hand to your ad platforms. ## Now you can bring in the ad platforms Once the operational floor is solid, returns data stops being a symptom you're papering over and becomes a genuinely useful signal. A few ways to use it. **Custom labels in Google Shopping.** Product feeds allow five custom labels, and one of the better uses going is a returns tier: label products ReturnRate_High, Medium or Low from trailing return data, refreshed automatically by a feed rule or script. Then structure campaigns so the tiers carry different targets. High-return products are less profitable per conversion, so they need to earn more per sale: put them behind a higher ROAS target, or in a tightly budgeted campaign of their own, and let low-return products run freer. Concretely: dress A returns at 5%, dress B at 40%. Left alone, Smart Bidding spends happily on both, because it can't see that B's conversions keep un-converting. Tier them, demand say 30% more ROAS from the high-return group, and budget migrates towards revenue that stays. Feed specialists have been arguing for years that custom labels should carry operational data like margin and returns rather than just "summer sale", and one UK agency reports that labelling high-return items and bidding accordingly directly improved marketing ROI. Two cautions. This is a compensator for categories with inherently high returns (occasionwear that gets worn once and sent back), not life support for products you should be fixing. And don't leave a product in the naughty tier after the fix has landed. **Value-based bidding on net revenue.** More powerful, and more work: tell the platforms what a sale was actually worth after returns. Google supports conversion adjustments, so when an order is refunded you retract the conversion or restate its value, keyed on the order ID, and Smart Bidding gradually learns which contexts produce revenue that sticks. Meta's value optimisation works on the same principle: feed accurate values through the pixel and Conversions API and delivery chases predicted value rather than raw purchase counts. You never tell the algorithm "people who search 'red dress' bracket like mad"; you feed it net values, and it works out on its own that specific-product searches keep their orders while broad-query browsers don't. The catch is plumbing. Closed-loop refund feeding needs your returns system talking reliably to your ad accounts, and few brands have built it, which is exactly why it's an edge for the ones that do. **Audiences.** If you can identify serial returners in your customer data, and I mean the buy-everything-return-everything pattern rather than your best customers, who return plenty because they buy plenty, exclude them from prospecting via Customer Match or a custom audience. Point lookalikes at the opposite group: multi-order customers with low return rates, so the platform hunts for people who resemble keepers. And treat a return as a marketing trigger rather than a dead end. Someone who just sent back ill-fitting shoes doesn't want a generic "come back soon" ad, but a similar style with a sizing nudge and free exchange delivery might turn the return into an exchange instead of a lost customer. The common thread is that you're teaching the machines what a good sale looks like. Do that on top of a broken product experience and you've automated the doom loop from earlier. Do it after the fixes and every pound of spend starts flowing towards revenue that survives the returns window. ## Where this leaves you The order of operations is the whole argument. Fix sizing, photography, product content and quality control, using machine learning where it genuinely helps, because that removes returns while growing sales. Build the data spine that joins ad spend to orders to returns to reasons, because that makes the problem visible. Then, and only then, get clever with labels, adjusted values and audiences, because at that point you're fine-tuning a machine that works rather than compensating for one that doesn't. H&M's chief executive put the goal simply: whatever customers buy, they should want to keep. That's a sentence that aligns operations and marketing better than any org chart. When it's true, the next marketing pound goes towards kept revenue rather than an initial sale, and that, more than any bidding trick, is what makes the ad spend count. --- # Building Off the Algorithm: 1,344 Anti-Algorithm Playlists, Start to Finish URL: https://www.xoot.uk/blog/building-fringe-fm-anti-algorithm-playlists Published: 2026-06-10 Category: Fringe FM *The product: [Fringe FM](/fringe-fm) — playlists sorted by mood, activity and era, live at [fringefm.net](https://fringefm.net).* ## Where this came from This whole project owes its existence to [**Sonosaurus**](/sonos-controller) — a previous build of mine that lets me request playlists by voice through Siri. The setup: a bot running on my home network listens for a Shortcut trigger from my phone or watch, takes a natural-language prompt ("something cinematic and Brazilian for cooking dinner", "the kind of thing you'd hear in a Berlin record shop at 4pm"), and three things happen at once. Claude generates a tracklist. The music starts playing on Sonos. The same tracklist hits the Spotify API to create a real, persistent Spotify playlist with a deliberately silly name — something like "Wrestling With Gravy" rather than "Chill Vibes 2026" — and the playlist link plus tracklist gets pushed to a Telegram bot so I can save it, share it, or come back to it later. End-to-end, voice-to-music-with-shareable-link in about twenty seconds. Sonosaurus was a hobby tool — a faster way to get music going than scrolling through my own library or arguing with Spotify's algorithm. But two things kept nagging at me after I'd been using it for a few months. First: the playlists Sonosaurus produced were genuinely better than what I was getting from any of Spotify's editorial or algorithmic suggestions. The deep-cut, taste-driven, anti-default-pick prompt I'd written for it was doing real curation work. Better than what Spotify served me, better than what my friends sent me, occasionally better than what I'd have picked for myself. Second: Sonosaurus was a one-person tool. Every playlist it generated landed in my Telegram and on my Sonos — but the curation engine doing the work was, by any reasonable measure, capable of serving a lot more than one household. Friends would ask me to "do a Sonosaurus" for their dinner party. The Spotify links I'd share would get follows from people who weren't me. There was clearly demand for this kind of curation beyond what a voice-triggered home setup could serve. The idea for this project was the obvious next step: take Sonosaurus's exact technical pattern — generate a tracklist with Claude, create a real Spotify playlist via the API, give it a memorable name — but instead of running it on-demand for me, run it ahead of time across every plausible mood/activity/era combination, save every output, and put the whole library on a [public website](https://fringefm.net). The single-user voice-controlled hobby becomes a 1,344-playlist public resource. The Sonosaurus DNA is throughout this project. The prompt that selects the tracks is a direct descendant of the Sonosaurus system prompt, just with the discipline cranked up because public-facing playlists need to be more consistent than throwaway dinner-party background music. The playlist naming voice — the funny, slightly absurd, deliberately-not-corporate names — came directly from what Sonosaurus already did via Telegram. Even the Spotify-playlist-creation code is a near-port of the same logic, just batched across 1,344 cells instead of called one at a time. If Sonosaurus is the single-user, voice-triggered version of this idea, **Fringe FM is the same idea scaled across every listening context I could think of, pre-generated, and made public.** ## The brief I wanted to build a website that solved one specific problem: people who are tired of Spotify's recommendations and want to find music outside the algorithmic loop. Not another playlist app, not another "AI DJ", not a social music network. Something dumber and more useful — pre-curated playlists with deep, taste-driven track selection, organised by mood, activity, and era, ranked to be discoverable through search engines. The mental model was record-collector tier curation at scale. The kind of selections you'd hear at a good independent record shop or read about in the back pages of Wire magazine. Light in the Attic, Music From Memory, Numero Group, Habibi Funk, kankyō ongaku — that whole reissue-driven, deep-cut sensibility, applied across every mood and listening context. The constraint that defined the architecture: I'm one person, and I needed to ship the whole library. So everything had to be automatable, but the *judgement* of what made a good playlist couldn't be. The bet was that I could put enough taste into a system prompt that a language model would consistently apply it across 1,344 different contexts. ## The taxonomy The library is built on three dimensions multiplied together. **25 moods**: melancholic, contemplative, euphoric, restless, defiant, nostalgic, tender, dreamy, anxious, sombre, hopeful, weary, joyful, lonely, focused, energised, bittersweet, yearning, meditative, playful, romantic, introspective, ethereal, brooding, uplifting. **20 activities**: coffee, reading, driving, cooking, late-night, working, dinner-party, hangover, rainy-day, walking, studying, cleaning, travelling, early-morning, before-bed, gardening, running, solo-evening, road-trip, afternoon-slump. **3 eras**: vintage (reissue-tier, pre-1990 lean), modern (post-2010 underground), timeless (cross-decade). 25 × 20 × 3 = 1,500 raw combinations. After filtering incompatible pairings (you don't need a euphoric playlist for bedtime, or a sombre one for running), the final library is 1,344 cells. Each cell becomes one playlist of 15 tracks. The URL structure is `/playlists/{mood}-{activity}-{era}` — for example `/playlists/melancholic-coffee-vintage` or `/playlists/euphoric-driving-modern`. Hub pages roll up by single dimension: `/melancholic` lists all 36 melancholic playlists across activities and eras, `/coffee` lists all 70 coffee playlists. Natural internal linking, long-tail search coverage, no duplicate content. ## How the tracks were chosen This is the part where the actual work of curation happens — and the part where most "AI playlist" projects fail. Out of the box, large language models will give you Bonobo, Tycho, Khruangbin, Mac DeMarco, Nils Frahm, and Bon Iver in every playlist. They've been trained on every Spotify-Discover-Weekly-adjacent thing ever written, and they reach for the safe pick by default. Getting Claude to produce taste-driven, deep-cut selections required a heavily engineered system prompt. The key components were these. **Numeric listener targets.** Most picks needed to be artists with under 100,000 monthly Spotify listeners. The mix was specifically: 65% under-100K, 25% mid-tier, 10% accessibility anchors. Forcing the model to think about actual listener counts shifted the centre of gravity dramatically. **The First Track Rule.** Three tiers of forbidden openers, in order of severity. Tier 1: AI defaults (Norah Jones, Sade, Massive Attack — the algorithm's comfort food). Tier 2: "indie respectable" — Khruangbin, Mac DeMarco, Big Thief, Japanese Breakfast — the artists that signal taste without actually demonstrating it. Tier 3: ambient defaults — Bonobo, Four Tet, Floating Points, Nils Frahm. These artists could appear later in a playlist as anchors but never as the opening track. The opener sets the brand promise: this is going to be different. **Era and geography mandates.** Every playlist had to include at least 4 pre-2010 picks, at least 2 pre-1990 picks, 3-4 non-Anglo selections, and span at least 5 distinct subgenres. This forced diversity that the model wouldn't generate on its own. **Explicit reissue label naming.** The prompt named specific labels — Light in the Attic, Music From Memory, Numero Group, Habibi Funk, Awesome Tapes from Africa, Soundway, Analog Africa, Glossy Mistakes — to give the model a vocabulary for the aesthetic. Once you tell it "think Music From Memory tier", you get Suzanne Ciani, Pauline Anna Strom, Hiroshi Yoshimura, Pep Llopis, Suso Saiz, Roberto Musci. Without that vocabulary, you get Tycho. **Curator angles.** A list of ten "lenses" — second-person British humour, lost female composers, post-punk diaspora, kosmische, MPB deep cuts, chanson modernity, etc. — two random ones injected into each prompt to shift the focus and prevent the library feeling homogeneous. ## On temperature Worth a small detour because it matters more than people think. When you call a language model, you set a parameter called **temperature**, usually between **0.0 and 2.0**. It controls how much randomness the model uses when picking its next word at each step. At 0.0, the model always picks the single most probable word — same input produces same output, every time, completely deterministic. At 2.0, the model samples wildly from less probable options, which produces creative but often incoherent output. Most real-world applications sit between 0.3 (technical writing, factual answers) and 1.0 (creative writing, brainstorming). For this project I started at **1.0** — the upper end of "creative but coherent". The output was genuinely interesting but a bit too wild. I'd get playlists where the model would commit to an obscurity so deep it had invented an artist that didn't exist, or include three tracks from genuinely-existing-but-only-on-Bandcamp artists that Spotify had never heard of. The kind of failure mode where you can see the model trying so hard to be adventurous that it's stopped being useful. Dialled down to **0.9** and the difference was immediate. Still genuinely surprising picks — Pep Llopis, Sibylle Baier, Mort Garson's gardening album, kosmische deep cuts most people have never heard. But now mostly real tracks that actually existed on streaming services. The sweet spot for music curation in this style turned out to be a single decimal step below "maximum reasonable creativity." Worth knowing if you're doing similar work: temperature is not a binary on/off for creativity. Small movements have large effects. ## On the Batch API The thing that made this project economically and practically feasible was Anthropic's **Batch API**. A regular API call is synchronous: you send a request, the model thinks, the response streams back, you wait. For 1,344 sequential requests at maybe 8-15 seconds each, that's hours of wall-clock time before you can do anything else. Anthropic's Batch API inverts this: you upload all 1,344 requests as a single batch file, Anthropic processes them in parallel across their fleet, and you collect the results when they're done. **Crucially, you also get a 50% discount on tokens for using it.** The trade-off is supposed to be that batch jobs can take up to 24 hours to complete. In reality, my 1,344 playlists came back in roughly **10 seconds**. Anthropic's infrastructure is genuinely good at parallel inference, and a batch of this size barely registers. The implications are significant. For one-time bulk generation tasks like this, the cost of Opus drops to about the same as Sonnet's regular pricing — meaning you can afford to use the best model for tasks you'd normally compromise on. Combined with prompt caching (the system prompt is identical across all 1,344 calls so it gets cached after the first one, dropping per-call costs further), the total Opus spend for generating the entire library was around $15. The same library generated via individual synchronous calls would have cost roughly $30 and taken six to ten hours of wall-clock time. Batch turned it into "submit it, make tea, come back to a complete library." This is the right tool for any "I need to do N variations of the same prompt" problem. Bulk content generation, dataset creation, evaluation runs, A/B testing different prompts at scale. The cost and time savings stack and they're substantial. ## The human involvement I want to be honest about this because it matters: the model picked the tracks, but I shaped how it picked them, and I made every architectural decision about what kind of library this should be. Specifically, the human work was deciding the brand positioning (anti-algorithm, record-collector tier, not another mood playlist app); designing the taxonomy and compatibility rules; writing and iterating the system prompt — probably 30+ revisions across the course of the project, including the three-tier forbidden opener rule, the listener-count targets, the reissue label vocabulary, the era/geography mandates, and the temperature tuning; running test playlists, listening to them, flagging when the model was reaching for the safe pick, adjusting the prompt; auditing the final library for duplicates, opener leaks, structural integrity; renaming all 1,344 playlists using a separate Claude pass with the Sonosaurus-style humour brief, then re-auditing for duplicates and bland names; and making every call about what to do when things broke. What the model did: produce 1,344 specific tracklists that followed those rules. With taste. Without me needing to write each one. The track selection logic was working within a week. The brand voice and the discipline to keep the prompt tight took another month of iteration. ## The pipeline The mechanical steps from "taxonomy file" to "1,344 published playlists" went like this. **Generate taxonomy** — a Python script produces a JSON file listing all 1,344 mood/activity/era combinations. **Generate playlists** — Anthropic Batch API processes all 1,344 in parallel, each cell gets one Claude call with the full system prompt and its specific vibe metadata. Output: 1,344 JSON files, one per playlist, each containing 15 artist/title pairs. **Audit the library** — duplicate detection (no two playlists with identical tracklists), opener-leak detection (catch any forbidden artists as openers), structural integrity check. **Rename playlists** — separate Claude pass to give each playlist a distinctive, funny, British-leaning name in the Sonosaurus style. Initial Haiku run produced bland results; Sonnet was meaningfully better at the humour. **Resolve track URIs** — for each track in each playlist, search Spotify and Tidal to find the actual streaming URI. This is where most of the time and pain went. **Create real playlists** — for each platform, OAuth as the brand account, POST a new playlist, populate it with the resolved URIs, save the playlist ID back into the JSON. **Generate the static site in Lovable** — by this point each playlist JSON contains everything the site needs: name, mood, activity, era, vibe metadata, the 15 tracks, the Spotify URI, the Tidal URI, the Spotify playlist URL, the Tidal playlist URL. I handed the JSON files to Lovable to produce [fringefm.net](https://fringefm.net) and described what I wanted: a page per playlist with embedded players for both services, hub pages by mood/activity/era, schema.org markup for SEO, and a homepage that surfaces the brand. Lovable generated the whole thing. The fact that the whole pipeline outputs structured data was the entire point — any frontend tool that can consume JSON can render this library. Steps 1-4 took about a week. Step 5 on Spotify took two weeks. Step 5 on Tidal took twelve minutes. More on that below. ## On letting tools do their jobs A point worth making explicitly: the JSON-first pipeline meant the static site was a separate problem that could be solved with a separate tool. I didn't need to learn a framework, write a templating engine, or hand-code 1,344 HTML pages. Lovable took the structured data and produced the site. Same JSONs would have worked fine in Next.js, Astro, Eleventy, or anything else that reads files and renders templates. The discipline of keeping the data layer clean and platform-agnostic meant the frontend choice was reversible and low-stakes. If Lovable doesn't suit later, I swap it for something else, same JSONs, same site shape. This is the underrated benefit of structuring AI-generated work this way. The model produces *data*, not deliverables. Every downstream step — naming, verification, playlist creation, site generation — operates on the same JSONs and adds to them. No regenerating, no re-prompting, no compounding errors. The library exists as data first and as anything else second. ## What went wrong, and how Claude and I worked through it A note before this section: Claude (Anthropic's assistant) wrote essentially all the code in this project. It also made most of the mistakes I'm about to describe. That's not a complaint — it's an accurate description of how AI-assisted development works in 2026. Claude is genuinely excellent at code, and also confidently wrong sometimes. Knowing where to apply scepticism is the human's job. **The opener leak problem.** Early playlist generations kept defaulting to Norah Jones, Khruangbin, or Bonobo as opening tracks despite the prompt saying "be adventurous". When I flagged that the openers were too safe, Claude proposed the three-tier forbidden list. Once specific artists were named explicitly, Claude genuinely avoided them. The lesson: a vague instruction like "avoid algorithm-tier picks" isn't enough; you have to name the specific traps. **The duplicate playlist problem.** Claude's first playlist-naming pass used Haiku for cost reasons and produced 189 duplicate names across 1,344 playlists ("Late Night Lounge" appeared seven times). The audit script Claude had also written caught this. We re-ran with Sonnet for the same task — meaningfully better at humour and variety — plus a global collision check. Worth knowing: humour and wordplay are tasks where model size genuinely matters. Haiku is fine for routine work; for anything depending on creative voice, Sonnet or Opus is the right call. **The Spotify rate-limiting saga.** The first version of the Spotify script that Claude wrote had 5 parallel workers and no rate limiting. It got the Fringe FM app soft-banned for 10.5 hours within minutes. Each subsequent iteration was Claude rewriting its own previous attempt as new ban data came in. First attempt: 5 workers, no rate limit — soft-banned for 10.5 hours after ~60 tracks. Second attempt: Claude added a token bucket, dropped to 2 workers at 1 req/sec — soft-banned for 7 minutes. Third attempt: 1 worker at 0.5 req/sec, capped retry-after, abort on long bans — got through 446 tracks before another ban. Fourth attempt: Claude pivoted from Client Credentials to user OAuth. User-authenticated traffic has dramatically more headroom because Spotify treats it as "this user is using their own data" rather than "this app is scraping us". Worth noting that Sonosaurus had been hitting the same Spotify Create Playlist endpoint dozens of times a week for months without ever being rate-limited — but Sonosaurus only creates one playlist per call, with hours between calls. The pattern that triggers Spotify's abuse detection is volume and burstiness, not the endpoint itself. The same code running at "background music for a Wednesday evening" cadence is invisible to Spotify; running at "build an entire library in one go" cadence is exactly what their abuse layer is designed to catch. Even after the OAuth pivot, Spotify imposes a daily quota on new apps. Early sessions capped out around 65 playlists per 24-hour window before being soft-banned for another 24 hours. The Spotify side ran as a Windows scheduled task firing every 25 hours, chipping away at the library overnight. Interestingly, the per-day throughput climbed over time. By the second week, the same script was creating around 150 playlists per session before being throttled — more than double the early-run ceiling. Two things were going on, and both helped. The first was a persistent local cache of resolved track URIs. The same track shows up in multiple playlists across the library — Caetano Veloso's "Cucurrucucú Paloma" might fit melancholic-cooking-vintage, melancholic-rainy-day-timeless, and tender-coffee-vintage. Once Claude had searched Spotify for that artist and title and stored the URI, every subsequent playlist that needed it skipped the API call entirely. As the cache filled up, each playlist required fewer real Spotify requests to assemble. By the second week the cache was hitting on roughly 30-40% of tracks, which meant the daily quota stretched proportionally further before being throttled. The second was probably trust. I can't prove this mechanism, but the most likely explanation for the *remaining* speedup is that Spotify's abuse detection treats apps with sustained, predictable, well-behaved traffic patterns more leniently over time. New apps making bursty requests at unpredictable times get throttled aggressively; an app that does roughly the same thing every 25 hours from the same account looks less suspicious. The infrastructure is essentially learning that this app isn't a scraper. Useful to know if you're building anything similar: the early days are the hardest. Cache aggressively. Stay polite and consistent. Both effects compound. **The Spotify endpoint removal that wasn't (and where Claude got it wrong).** Spotify's February 2026 developer API migration removed `POST /users/{user_id}/playlists` — the historical endpoint for creating playlists. Claude read the changelog, saw "REMOVED" against playlist creation, and confidently told me dev mode apps could no longer create playlists at all. It walked me through a substantial architectural pivot: drop the "real Spotify playlists" model entirely, switch to per-track embeds, restructure the website to render 15 individual track players per page with a custom JS queue, the whole thing. I was halfway down that path when I cross-checked with ChatGPT, which immediately pointed out that `POST /me/playlists` is the documented replacement endpoint — same purpose, different URL shape. The "REMOVED" line in the changelog only applied to the old `users/{user_id}` form. The new `me/playlists` form was right there in the same docs. When I fed ChatGPT's response back to Claude, it acknowledged the mistake immediately and produced the correct script in about ten minutes. Cost of the detour: maybe three hours of work building toward an architecture I didn't need. The lesson isn't "Claude is bad" — Claude was extremely good at every other stage of this build, including parts where it had to debug its own previous mistakes from clues in HTTP responses. The lesson is that even excellent assistants can confidently misread a single line of documentation, and the fix is to triangulate across multiple sources before committing to a major architectural decision. Trust but verify, especially for anything that would require throwing away work. **The Tidal endpoint discovery process.** Tidal's developer API documentation isn't crawlable — their reference site is a JavaScript-rendered Swagger UI that returns an empty shell to scrapers. Claude built diagnostic scripts that probed various URL shapes and Accept-header combinations, with verbose request/response logging. Every probe returned 404 with an empty body. After several rounds of guessing, I gave up and just opened the docs in my browser, copied the visible endpoint list, and pasted it back to Claude. Claude spotted the issue immediately: the endpoint is `GET /v2/searchResults/{query}` — camelCase, not lowercase. Tidal's API gateway is case-sensitive on the path, so every previous attempt at `/searchresults/` was being rejected at the routing layer before reaching the search service. Single-letter casing bug, hours of misdiagnosis on Claude's part because it couldn't see the docs directly. **The JSON-edit-with-regex disaster.** When migrating a stale flag out of the playlist JSONs, Claude suggested a one-line PowerShell regex to strip the field. I ran it. The regex removed the field but left trailing commas behind, breaking the JSON in every file. 1,344 files damaged in one command. Claude then wrote a repair script using `json.loads`/`json.dumps` that fixed everything in one pass. The actual lesson Claude articulated afterward: never use regex to edit structured data, even for trivial changes. Use the data format's actual parser. Worth noting that Claude knew this rule going in — it just didn't apply it to a "small" task. AI assistants, like humans, can violate their own best practices when the task feels too quick to warrant care. ## Spotify vs Tidal: night and day The most striking part of this whole project is the difference between the two platforms' developer experiences. **Spotify's developer API for new apps is hostile.** The rate limits aren't published, the abuse detection is sensitive to specific request patterns (consecutive failed searches trigger soft-bans even at conservative rates), and the daily quota for unverified apps starts somewhere around 800-1,000 API calls per 24 hours. Working through two weeks of scheduled restarts with gradually-increasing throughput is the only path forward for a small developer building a public-facing music site. The extended access tier that would unblock this requires 250,000 monthly active users — a chicken-and-egg requirement that effectively blocks any new app from being viable at launch. **Tidal's developer API for new apps is generous.** Once the camelCase bug was fixed, the full library of 1,344 playlists was created in twelve minutes. No daily quota, no soft-bans, no scheduled restart pattern needed. The 1 req/sec rate limit Claude built into the script as a precaution turned out to be vastly more conservative than necessary. Tidal's user-authenticated API tolerated continuous traffic without issue. The hit rate difference is interesting too. Spotify resolved roughly 84% of tracks; Tidal resolved closer to 78%. The catalogs overlap but aren't identical — particularly in the deep reissue territory where Music From Memory titles, Awesome Tapes from Africa releases, and certain Brazilian, Japanese, and library music catalogues are present on one service but not the other. Some of the deepest cuts in the library are on Tidal but not Spotify; others vice versa. **The implication for "anti-algorithm" branding.** Tidal's audience already self-selects for caring about music over convenience. Their userbase is smaller than Spotify's by an order of magnitude, but it skews heavily toward audiophiles, DJs, music journalists, and the exact "bored with the algorithm" demographic this project targets. The brand probably fits better there. Spotify still gets the project's larger audience reach, but Tidal is where the actual readers are. ## What I'd do differently Three things, in order of how much they would have saved. **Start with Tidal, not Spotify.** I assumed Spotify-first because Spotify is bigger. But Spotify's API is hostile to new apps and ate weeks of iteration. Tidal was twelve minutes. If I'd started with Tidal I'd have had a working product two weeks earlier and could have used that as a demo to apply for Spotify extended access. **Build the diagnostic-first pattern from the start.** By the time we got to Tidal, Claude had landed on a useful pattern: build a small diagnostic script that does one operation end-to-end and prints every request/response in full detail. Before scaling to 1,344 of anything, run the diagnostic with one. This caught the camelCase bug immediately. Earlier in the project — through all the Spotify rate-limiting iterations — Claude was writing production-shaped scripts directly and learning their flaws by running them against the real API. Going diagnostic-first from day one would have saved most of those iterations. Worth establishing this as a default for any new API integration with an AI assistant. **Treat API documentation as untrustworthy until proven otherwise.** Both Spotify and Tidal had documentation that didn't match runtime behaviour in important ways. Spotify's changelog implied playlist creation was removed entirely (it wasn't — just renamed). Tidal's npm wrappers referenced endpoints that returned 404 in production (they were lowercase in the wrapper code but actually camelCase). The only reliable knowledge came from running real HTTP requests and reading what came back. ## Costs For anyone considering similar work: Anthropic API (all generation, naming, regeneration) was under $25 total. Spotify and Tidal developer accounts are free. Lovable (static site) was covered by the free tier. Cloudflare Pages hosting is free at this scale. Domain: ~£12/year. My time was a few weeks of calendar time, but rarely a deliberate sit-down session. Mostly five minutes here, twenty minutes there — whenever I had a spare moment between life and the children. The natural rhythm of the project turned out to be: prompt Claude, run something, hit a wall, hit Claude's session limit, go play with the kids until the quota reset, come back and pick up where I'd left off. The session limits were arguably load-bearing — they forced the breaks that would have been hard to take otherwise. The economics of this kind of project are wildly different from a year or two ago. The model spend for generating 1,344 carefully-curated playlists was roughly the cost of a takeaway. The infrastructure is free. The work that mattered was the taste decisions, the prompt engineering, and the patience to keep debugging APIs that don't cooperate. That's the part you can't shortcut. --- # The Bidding Ladder, and the Awkward Truth About Climbing It Too Fast URL: https://www.xoot.uk/blog/bidding-ladder-climbing-too-fast Published: 2026-04-15 Category: Google Ads *Why the smartest lifetime-profit model in the world won't save you from rubbish data, impatient boards, and the tyranny of moving too many things at once.* It's 9.47am on a Tuesday in February, and in the glass-walled meeting room of a converted Manchester mill — exposed brick, Edison bulbs, the obligatory motivational neon — a performance agency is preparing to do something brave on behalf of its biggest e-commerce client. They've been running the account on ROAS bidding for the best part of three years, and have quietly spent the better part of £100k building a proprietary customer lifetime value model on the side. The dashboards are immaculate. The data scientist has built something that would make an actuary whistle. This morning, finally, they are going to flip the switch from ROAS straight to lifetime profit. By the end of March, the campaign is paused, the model is shelved, and the account is up for pitch. This is a story you've heard before, possibly more than you'd care to admit. It is, in my experience, almost always the same story — and it is almost never about the model. ## The ladder, briefly In a sufficiently mature Google Ads account, bidding signals tend to ascend a now-familiar ladder, each rung representing a more sophisticated picture of value: **CPC**: are people clicking? **CPA**: are they converting? **ROAS**: how much revenue per pound spent? **POAS**: how much *profit* per pound spent (margins, shipping, payment fees, the lot)? **Lifetime revenue**: what do these customers spend over a year, two years, five? **Lifetime profit, with predictive returns**: what do they actually leave behind once you net out the trainers they sent back, the discount code they used, and the cost of acquiring the next cohort? Every rung up this ladder is, in principle, a better proxy for the thing you actually care about, which is whether your business will still be here in 2030. CPC is a signal about idle curiosity. Lifetime profit is a signal about whether you have a business at all. So, naturally, you should sprint to the top. Right? ## The data is the weakest link Here is the problem, stated plainly: a bidding strategy is only as good as the data feeding it. And the higher you climb, the more data points the model needs, and the more brittle the chain becomes. Server-side tagging, Google Tag Gateway routing requests through your CDN, Enhanced Conversions sending hashed first-party data straight to Google — these aren't fashionable acronyms to drop into a quarterly review. They are the difference between a model that knows what your customers actually did, and a model that is hallucinating in the polite, statistical sense of the word. Every cookie blocked by Safari's ITP, every consent banner click that wasn't wired to update Consent Mode properly, every Klarna postMessage you didn't capture — all of that becomes a small, plausible lie that the model dutifully repeats back to Google's algorithms with the confidence of a man explaining wine. Garbage in, garbage out is the cliché. The reality is rather more uncomfortable: garbage in, *plausible-looking* garbage out. The dashboard still works. The numbers still tick over. Smart Bidding still optimises confidently towards a target that bears an increasingly distant resemblance to your actual P&L. And here's the bit that quietly haunts every senior practitioner: a sophisticated model on poor data is **worse** than a simple model on good data. The simple model degrades visibly. The sophisticated one degrades invisibly. You don't notice you've been steering with a broken compass until you look up and realise you've sailed into the Strait of Hormuz. So before anyone so much as whispers "predictive lifetime profit" in a planning meeting, the boring questions are the load-bearing ones: Is your tagging genuinely server-side, or are you still relying on a browser GTM container that half your visitors politely decline? Is Google Tag Gateway routing your tag requests through your own CDN so they look first-party to browsers and ad-blockers, or are your hits dying in Safari and the ad-blocker market on the way out? Are Enhanced Conversions actually wired up and verified, sending hashed first-party data to Google so the algorithm has something to match on when the cookie isn't there, or did somebody tick the box once and never look at it again? Is your consent state actually synced — Cookiebot or CookieInformation or whoever you've chosen — to both Consent Mode v2 and every downstream tool, or do you have a fun little discrepancy nobody has audited since launch? Is your conversion value built on **true profit** per order — shipping, payment fees, discount codes, the lot — or just product margin? Do your refunds, returns, and cancellations flow back into your conversion data, or are you optimising towards revenue your warehouse is currently repacking? If any of those answers makes you wince, you don't have a bidding strategy problem. You have a measurement problem wearing a bidding strategy as a disguise. ## Why baby steps win Assume, generously, that your data is in order. Server-side tagging is humming, Google Tag Gateway is pushing your hits through a clean first-party endpoint, your POAS includes everything down to payment processor fees, and your lifetime model has been validated against twelve months of cohort data. You are ready. This is the moment a lot of accounts blow themselves up. The temptation, having done all that work, is to flip from ROAS bidding to lifetime-profit bidding in one confident motion. The trouble is that Google's bidding algorithms are themselves models, learning from your conversion signals over a multi-week ramp-up, and they don't enjoy surprises. Change the conversion definition — even if your new definition is genuinely better — and you have effectively asked the algorithm to relearn your account from scratch, with new value distributions, new variance, and a new relationship between clicks and reported outcomes. Combine that with the changes happening in your own business — seasonality, a product launch, that influencer who unexpectedly went viral, a competitor's price cut — and you've now got at least three confounders moving simultaneously. When the campaign underperforms in week four, nobody will be able to tell you whether it's the model, the algorithm's relearning curve, or the fact that it rained for a month. What internal stakeholders see, meanwhile, is a graph going down. Patience is a finite resource, especially after Q1 results. The accounts that successfully reach the top of the ladder almost all do it the same boring way: one rung at a time, with overlap, and with a control. Move from ROAS to POAS first. Run that for a month. Validate that the algorithm has settled and that profit-weighted demand actually shapes the way you think it should. Then layer in lifetime revenue. Then introduce predictive returns into your value calculation. **Then**, eventually, the full lifetime-profit-with-returns picture. This is unglamorous. It will not earn you a speaking slot at a conference. It will, however, keep your campaigns running, your client's leadership calm, and the account off the next pitch list long enough to enjoy the sophisticated model you've built. ## The optimisation strategy, on a Post-it note Strip the hierarchy back to the practical advice and it fits on something rather smaller than a slide deck. **Audit before you ascend.** Every rung up the bidding ladder is a tax on data quality. Server-side tagging, Google Tag Gateway, Enhanced Conversions, proper consent plumbing, and a returns feed aren't optional once you move beyond ROAS — they're prerequisites. **Validate the model offline first.** Before you let Google bid on lifetime profit, prove that your lifetime profit numbers reconcile to your finance team's view of the world. If they don't, the algorithm will be optimising to a fiction. **Move one rung at a time.** Each transition is its own learning period for both your account and Google's algorithms. Stacking changes guarantees you won't know what worked. **Run with overlap.** Keep a control campaign on the previous bidding strategy while you test the new one. It's the only way to tell skill from weather. **Sell the journey internally.** Stakeholders don't lose faith in lifetime profit because the maths is wrong. They lose faith because they were promised a three-week dip while the algorithm relearned, and lived through six. Pre-commit to the realistic timeline *and* the dip, up front. ## The grown-up version of the story Lifetime profit with predictive returns is, genuinely, a wonderful place to bid from. It aligns Google's considerable machine learning muscle with the only metric that ultimately pays your bonus. It is, in a real sense, the destination. But the destination is rarely the interesting part of the journey. What separates the accounts that get there from the accounts that don't is almost never the cleverness of the final model. It's the unglamorous, patient, slightly tedious work of keeping the data clean and the changes small. Climb the ladder. Just don't take three rungs at a time. Ladders are unforgiving like that — and so, eventually, are boards. --- # Meta's Privacy Sandbox pixel: a technical deep dive URL: https://www.xoot.uk/blog/meta-privacy-sandbox-pixel-technical-deep-dive Published: 2026-03-01 Category: Meta Pixel Meta quietly embedded Chrome's Attribution Reporting API (ARA) into its standard Facebook Pixel (`fbevents.js`), routing conversion signals through a dedicated endpoint at `facebook.com/privacy_sandbox/pixel/register/trigger/`. This endpoint represents Meta's implementation of the **trigger registration** step of Chrome's ARA — the privacy-preserving mechanism for recording conversions without third-party cookies. The endpoint runs in parallel alongside the traditional `/tr/` pixel, serving as a complementary attribution path for browsers that restrict cookie-based tracking. No official Meta developer documentation exists for this endpoint. In a significant twist, Google officially **retired the Attribution Reporting API** in October 2025, making this implementation effectively transitional — though Meta's broader privacy-preserving attribution work through the W3C now forms the backbone of the successor standard. ## How trigger registration works under the hood The `/privacy_sandbox/pixel/register/trigger/` endpoint implements one half of Chrome's ARA two-step protocol. The full flow operates as follows: when a user views or clicks a Meta ad on a publisher's site, Meta registers an **attribution source** with the browser via the `Attribution-Reporting-Register-Source` HTTP response header, storing encrypted metadata (campaign ID, destination site, expiry) in the browser's private cache. Later, when that user visits an advertiser's site and performs a conversion — a purchase, form submission, or other tracked event — Meta's pixel fires a request to the `/privacy_sandbox/pixel/register/trigger/` endpoint. This request uses the `fetch()` API with the `attributionReporting: {eventSourceEligible: false, triggerEligible: true}` option, causing Chrome to attach an `Attribution-Reporting-Eligible: trigger` request header. Meta's server then responds with the `Attribution-Reporting-Register-Trigger` response header containing JSON-encoded trigger configuration — event trigger data, aggregatable trigger data, priority values, deduplication keys, and filter specifications. The browser processes this header internally, **matches the trigger against stored sources** based on destination site and reporting origin, and schedules delayed attribution reports to be sent back to Meta. Event-level reports arrive with **2–7 day delays** and contain limited, noised data; aggregatable reports are encrypted and must be processed through a Trusted Execution Environment before Meta receives differentially private summary statistics. The critical difference from the traditional `/tr/` endpoint is **who controls the attribution logic**. With `/tr/`, Meta receives full event data in real time and performs server-side matching using third-party cookies (`_fbp`, `_fbc`) and user identity graphs. With the ARA trigger endpoint, the **browser itself** performs attribution matching — Meta never receives cross-site identifiers, only delayed, noised, or aggregated reports. ## Parameters decoded: cd[], rqm, expv2, and the rest The request parameters sent to `/privacy_sandbox/pixel/register/trigger/` blend established Facebook Pixel conventions with Privacy Sandbox–specific data. Here is what each parameter group carries: **`cd[]` (Custom Data)** is the best-documented parameter, inherited directly from the traditional pixel. It carries event-specific conversion data — `cd[value]`, `cd[currency]`, `cd[content_type]`, `cd[content_ids]`, and similar fields. When an advertiser calls `fbq('track', 'Purchase', {value: 115.00, currency: 'USD'})`, the pixel serializes those properties into `cd[]` query parameters. In the Privacy Sandbox context, this data informs Meta's server how to construct the ARA trigger response header — determining trigger_data values, aggregatable key structures, and contribution budgets. **`rqm=FGET` (Request Method: Fetch GET)** indicates the HTTP transport mechanism. Traditional pixel calls use `rqm=GET` for `` tag requests and `rqm=POST` for XMLHttpRequest calls. The **FGET** value means the request was made via the browser's `fetch()` API using the GET method. This is not arbitrary — ARA trigger registration requires `fetch()` with specific `attributionReporting` options to signal trigger eligibility to Chrome. The `FGET` designation lets Meta's server distinguish Privacy Sandbox–eligible requests from standard pixel fires. **`expv2` (likely Experiment Version 2)** is undocumented publicly but almost certainly encodes Meta's internal **A/B testing and feature flag assignments**. Meta operates one of the largest experimentation platforms in the industry, and the `exp` prefix plus `v2` suffix strongly suggest this carries experiment bucket identifiers that determine which ARA configuration variant — trigger data encoding, aggregation key structure, privacy budget allocation — the server should use when constructing the response header. **`pmd[]` and `ap[]`** are also undocumented. Based on naming patterns and ARA protocol requirements, `pmd[]` most likely stands for **"Pixel Metadata"** — configuration data about the pixel instance itself (version, initialization state, consent signals) distinct from the event data in `cd[]`. The `ap[]` group likely carries **"Attribution Parameters"** — data specifically needed for ARA trigger construction, such as aggregatable values, key pieces, or priority signals that map to the `aggregatable_trigger_data` and `aggregatable_values` fields in the ARA response header. These interpretations remain informed inference; Meta has not publicly documented these parameters. ## Meta's broader strategy bypassed Google's Privacy Sandbox The most striking finding is that Meta was **never a listed tester** of Google's Privacy Sandbox APIs. While companies like Criteo, RTB House, NextRoll, and Index Exchange actively tested ARA through Google's origin trials, Meta pursued an entirely different path. Starting in February 2022, Meta engineer Ben Savage and Mozilla engineer Martin Thomson co-developed **Interoperable Private Attribution (IPA)** — a competing proposal submitted to the W3C's Private Advertising Technology Community Group. IPA differs fundamentally from Google's ARA. Where ARA relies on Trusted Execution Environments and is Chrome-specific, IPA uses **Multi-Party Computation (MPC)** with independent "helper parties" and was designed to work across browsers and devices. Crucially, IPA leverages **encrypted match keys** tied to user logins — a design that exploits Meta's core competitive advantage of billions of logged-in users across Facebook, Instagram, and WhatsApp. This enables cross-device attribution that ARA could never achieve. Despite this strategic divergence, Meta's `fbevents.js` code does fire requests to Privacy Sandbox endpoints in production — both the `/privacy_sandbox/pixel/register/trigger/` ARA endpoint and a separate `/privacy_sandbox/topics/registration/` Topics API endpoint. The earliest documented observation of the trigger endpoint came in **January 2025**, when a Gravity Forms community member noticed these requests firing during form submissions. This suggests Meta hedged its bets, implementing Chrome's APIs at the code level while advocating for its own standard at the policy level. ## The ICO report reveals Meta's five-stage PPA architecture The most detailed official description of Meta's privacy-preserving attribution system comes from the UK Information Commissioner's Office, which published a **Regulatory Sandbox Final Report** in September 2025 after collaborating with Meta from June 2024 to April 2025. Meta's Privacy Preserving Attribution (PPA) system operates through five stages: source registration stores ad impression data on the user's device; trigger registration records conversion events; on-device extraction creates encrypted "histogram contributions"; MPC computation splits encrypted data between two independent helper parties who aggregate it with differential privacy noise; and finally a differentially private histogram is returned to the advertiser. The ICO concluded that UK privacy regulations (PECR) **do apply** to PPA because information is stored on and accessed from users' devices, meaning user consent would be required. However, the report acknowledged that PPA "reduces data protection risks in comparison with existing industry practices" — a qualified endorsement of the approach. ## Google retired ARA, vindicating Meta's W3C bet On **October 17, 2025**, Google officially retired the Attribution Reporting API along with Topics, Protected Audience, Private Aggregation, and most other Privacy Sandbox technologies, citing "low levels of adoption." Google simultaneously announced it would contribute learnings to the **W3C's Privacy-Preserving Attribution: Level 1 specification** — the interoperable standard being developed through the Private Advertising Technology Working Group, where Meta's IPA proposal serves as a foundational input alongside Apple's Private Ads Measurement. This retirement renders Meta's `/privacy_sandbox/pixel/register/trigger/` endpoint effectively transitional. As Chrome phases out ARA support (specific removal timeline not yet published as of March 2026), these calls will cease functioning. However, Meta's strategic positioning has proven prescient: the W3C successor standard incorporates IPA's MPC architecture, cross-device match keys, and server-side privacy enforcement — all Meta innovations. ## Conclusion Meta's Privacy Sandbox pixel implementation represents a pragmatic hedge rather than a strategic commitment. The company built ARA trigger registration into its production pixel code while simultaneously developing and advocating for a fundamentally different architecture through W3C standardization. The `/privacy_sandbox/pixel/register/trigger/` endpoint is genuine Chrome ARA infrastructure — it fires `fetch()` requests with attribution reporting options, carries conversion data through familiar `cd[]` parameters, and elicits proper `Attribution-Reporting-Register-Trigger` response headers. But the undocumented status of key parameters like `pmd[]` and `ap[]`, the absence of any Meta developer documentation, and Meta's conspicuous absence from Google's official tester lists all signal that this was a parallel implementation rather than a primary investment. With ARA now retired and the W3C standard ascending, Meta's real attribution future lies not in this endpoint but in the MPC-based, cross-device PPA system it has been building through the standards process — a system designed around its greatest asset: billions of logged-in users. ---