How I built a scraping-free data pipeline for a marketplace that refuses datacentre IPs
If a marketplace does not want to be scraped, you will not scrape it — not reliably, not at a cost that supports a $13/month SaaS tier. The architecture has to invert: the server never reaches out for data, the data reaches in. Here is how NoonSpy ships that for the Noon marketplace, why the obvious alternatives failed, and the specific SQLite, TLS fingerprint and Stripe incidents that production taught me.
The marketplace
Noon is the second marketplace of the Gulf — tens of millions of SKUs across the UAE and Saudi Arabia, growing. The sellers on it have no Helium 10, no Jungle Scout, no research layer at all. They guess at demand, guess at pricing, and find out whether they were right after the inventory has already landed in a Noon warehouse.
I spent 2024 building NoonSpy as the layer I wanted to exist myself. The hardest part was not product. The hardest part was that every conventional way of getting data out of Noon dies on contact with the edge.
What dies, and why
Noon is fronted by Akamai. The architecture I had to abandon in the first week looks like this, and if you have ever tried to point a Node or Python scraper at a protected marketplace you have written this exact file:
Hetzner, Vultr, DigitalOcean, Linode — every datacentre IP I tried returned 403 before the response body opened. Rotating residential proxies worked, but at $0.80 per GB and a 15% failure rate that cascaded into whole category refresh jobs timing out at 2am, the economics did not survive the first month. Headless-browser farms (Playwright on $5 Spot instances) worked for about an hour and then burned, because Akamai's fingerprinting notices the TLS stack of your Chromium regardless of what the user-agent string says.
The lesson is not that scraping is hard — scraping has always been hard — it is that a marketplace spending Akamai money on edge protection is not going to be out-engineered by a Dockerised Playwright, no matter how good your proxy pool is. The arms race is not winnable by a single developer with a $12/month VPS.
The inversion
So the architecture had to flip. Instead of the server reaching out for data, data reaches in. A thin collector runs on a residential ISP connection (one of my own machines at home), reads the Noon pages a seller already visits in normal browsing, and pushes authenticated snapshots to a single endpoint on the platform:
The server never scrapes. It only receives, normalises, dedupes, and reasons. The hardest architectural decision in the product is a non-decision: there is nothing in the server for Akamai to notice, because the server has no scraper to break.
A side-effect that turned out to carry the whole product: trending data is literally what my users browse. The sellers producing the signal are the ones consuming the insight. The feed got more honest the more it was used — I did not have to simulate a global crawler, because my users became one together.
Three things that broke in production
1. SQLite 'database is locked' at peak
I run NoonSpy on SQLite in WAL mode, on one VPS. Postgres was the obvious choice and the wrong one for a single-operator app at this scale — SQLite handles ~2M snapshot rows and daily writes from the collector plus live API reads for well under a dollar a month of hardware, and backups are a scp from cron.
Then in Jun 2026, during a traffic peak, 'database is locked' started cascading into 500s. The collector was writing, the API was serving reads, and the AI service was writing cached verdicts — all competing for the same file. Default SQLite in Python opens with a 5-second busy_timeout and a non-WAL journal, which is fine for one writer and deadly for three.
The fix was four lines of settings and one architectural separation:
Zero downtime since. The lesson I took from this is not that SQLite is weak — it is that SQLite is production-grade if you know its specific edges. Most of the people who recommend 'move to Postgres' are solving a problem a settings change would have handled.
2. The Akamai fingerprint regression
In Aug 2026, Noon's edge upgraded its TLS fingerprinting. The collector started getting silently rate-limited — not blocked, just slowed. This is the worst kind of failure, because dashboards look fine; it took me two days of trending data loss before I noticed that category refresh counts had quietly dropped by 40%.
The fix was in two parts. First, TLS fingerprint rotation via a custom httpx transport, so the collector presents a plausible range of real-browser TLS stacks instead of Python's always-identifiable one. Second, a canary check that runs every 6 hours and compares the collector's observed category size against the previous day's — a drop larger than 25% now pages me.
The real lesson is that silent degradation is worse than loud failure. 500s page you; a 40% quiet drop in a background job does not. Build the canary that watches your data, not the alert that watches your dashboard.
3. The Stripe statement descriptor
Early Pro subscriptions showed up on customers' bank statements as BLISSFULTREND.COM — an old, dormant brand on the same Stripe account that I had forgotten about. Three refund requests arrived in a single week from users who did not recognise the charge.
The fix is one setting per charge:
In a multi-brand Stripe account, the account-level descriptor is only the default — explicit beats implicit every time. Write the test that fails CI if the descriptor is ever anything else, and the whole class of error is dead.
The economics that falls out of all this
Here is what shipping the inversion made possible financially:
- Infrastructure cost under $40/month total — $12 VPS + Claude API usage + Stripe fees
- Pricing starts at AED 49/month (≈$13) for Pro, AED 149/month for Business
- Positive unit economics from the first paying user, because scraping cost is zero
- No proxy bills, no browser farm, no Playwright burn-and-restart loops
A scraping-based NoonSpy would have needed ~$300/month in proxies and infrastructure before the first user paid, and would still have been fighting Akamai's upgrades every quarter. The inversion is not a performance optimisation — it is what makes the product possible at a price its market can bear.
What I would do differently if I started again
I would ship the canary before I shipped the collector. Two days of silently missing data in Aug 2026 was the sharpest reminder that observability is a feature, not an afterthought, and it is cheaper to build during the first week than during the first incident.
I would also give the collector a public reporting page from day one — a status endpoint that shows what categories have been seen in the last hour, which SKUs are missing, which merchants are undersampled. Everything that is implicit in a solo-run system becomes explicit the day you add a second operator or want a user to trust the data; shipping it upfront buys honesty for free.
Beyond that, the choices held. SQLite was right. Flask was right. Chrome MV3 was right. Claude for Arabic reasoning was right. The inversion was right. The edges of each one bit back exactly once, which is how I now know they will not bite again.
Where to find the rest
The full NoonSpy case study — nine chapters including the hard decisions, the honest limitations, and the stack rationale — lives on the portfolio at abdullah-haroon.com/work/noonspy/. The product itself is at noonspy.com. If you are building against a marketplace that would prefer you did not, DMs on LinkedIn and Twitter are open — I am always interested in how other operators are solving the same shape of problem.
FAQ
Questions this answers
Why not use the Noon API?
There is no public Noon API that gives seller-side marketplace data. Noon's partner API exists for logistics integrations, not for research. The research layer has to be built from the outside, which is what forces the architectural decisions in this post.
Isn't a residential collector just scraping with extra steps?
It is scraping only in the sense that it reads HTML. The important difference is that it reads pages a human is already visiting as part of their normal browsing, from the human's own machine, on the human's own residential ISP — the server never initiates a request to the marketplace. Legally and operationally this is a very different posture to a data-centre scraper, and marketplaces treat it very differently in practice.
Would this architecture work for Amazon or Shopify?
The inversion works anywhere the user is already in the marketplace. For Amazon the equivalent is already a well-trodden pattern — Keepa and SellerSprite both ship Chrome extensions that produce their data. The novelty in NoonSpy is applying it to a marketplace where no one else has, not inventing the pattern.
What stops a competitor from copying the Chrome extension approach?
Nothing — and that is fine. The moat is not the collector, it is the normalised history nobody else is accumulating. A competitor starting today has months of catch-up before their Verdict can look back as far as mine can, and the Verdict is only useful insofar as it can look back.
Written by Abdullah Haroon
Independent creative developer, photographer and SaaS founder based in Abu Dhabi, UAE. Available for freelance and contract work across the UAE, Saudi Arabia, the UK, the US and worldwide.