TL;DR
- Build a Zalando scraper only for public pages you are permitted to collect, and check the site's current terms and robots policy before running it.
- Start with one product URL and extract embedded structured data before adding browser automation.
- Use Playwright only when the required fields appear after JavaScript execution; plain HTTP is faster when the initial response already contains them.
- Store stable product identifiers, canonical URLs, currency, availability, retrieval time, and a content hash so updates are idempotent.
- Treat consent pages, locale redirects, missing variants, and access-denied responses as explicit terminal states.
What is a Zalando scraper and why would you need one?
A Zalando scraper is a controlled program that reads permitted public product pages and converts selected fields into a consistent record. Nstdata Crawl can serve as a managed acquisition option when a team does not want to operate browsers and routing itself, while Python libraries provide more direct control. Appropriate uses include internal QA, approved catalog research, and monitoring products your organization is authorized to observe.
The goal is not to copy the entire storefront. A reliable job starts from an approved URL inventory and a narrow schema. The ecommerce product-data guide explains why source records and accepted business records should remain separate.
What do you need before you start?
You need Python 3.11 or newer, an isolated environment, httpx, beautifulsoup4, and Playwright for pages that truly require a browser. You also need written permission or a documented lawful basis, a small test URL set, a locale and currency policy, and a destination schema.
python -m venv .venv source .venv/bin/activate python -m pip install httpx beautifulsoup4 playwright python -m playwright install chromium
Define these fields before writing selectors: product_id, canonical_url, name, brand, currency, price_text, availability, color, size_options, retrieved_at, and source_hash. Do not collect customer, account, or checkout data. Read the current Zalando terms and the site's robots file for the exact host and locale you plan to access.
How does a Zalando product page actually work?
A product page can combine server-rendered HTML, embedded JSON, and client-side requests. The visible DOM is not always the best source. Inspect the initial response and <script type="application/ld+json"> blocks first because structured product data is usually more stable than presentation classes. If required data appears only after rendering, use a browser and wait for a meaningful product element rather than sleeping for an arbitrary duration.
Connect to the Right ProxyChoose the location and session mode that fit your workflow, then connect through Nstdata. Set Up Proxy |
Sticky
Client Nstdata
🇺🇸US
🇩🇪DE
🇸🇬SG
|
Detailed Tutorial
Method 1: Extract JSON-LD from a permitted product page
Step 1: Fetch one URL with a bounded client
Use a descriptive user agent, a timeout, redirect limits, and a low request rate. The target URL must come from your approved inventory.
import os import httpx url = os.environ["TARGET_URL"] with httpx.Client(timeout=20, follow_redirects=True) as client: response = client.get(url, headers={"User-Agent": "CatalogQA/1.0 contact@example.com"}) response.raise_for_status() html = response.text
Step 2: Parse product JSON-LD
import json from bs4 import BeautifulSoup soup = BeautifulSoup(html, "html.parser") products = [] for node in soup.select('script[type="application/ld+json"]'): try: value = json.loads(node.string or "") except json.JSONDecodeError: continue values = value if isinstance(value, list) else [value] products.extend(x for x in values if isinstance(x, dict) and x.get("@type") == "Product") if not products: raise RuntimeError("No Product JSON-LD found; inspect the response before changing selectors")
Step 3: Normalize without inventing values
Map only present fields. Preserve the original price string and currency, and use None for unavailable data. Never infer a size or stock state from an absent button.
Method 2: Render the page with Playwright
Step 1: Launch a bounded browser context
import os from playwright.sync_api import sync_playwright with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page(locale="en-GB", viewport={"width": 1440, "height": 1000}) page.goto(os.environ["TARGET_URL"], wait_until="domcontentloaded", timeout=30000) page.locator('script[type="application/ld+json"]').first.wait_for(timeout=10000) rendered_html = page.content() browser.close()
Step 2: Detect the wrong page state
Check the final URL, page title, expected product identifier, locale, and presence of product evidence. Reject consent-only, login, challenge, and generic error pages. The official Playwright Python documentation is the source for current installation and waiting APIs.
Step 3: Save a review artifact
Save a screenshot or HTML hash for failures, but apply retention limits. Review artifacts should help diagnose a selector or rendering change without becoming an uncontrolled copy of the page.
Method 3: Use a managed crawl API for acquisition
Step 1: Keep acquisition separate from parsing
Use Nstdata Crawl when managed rendering, task state, and artifact delivery are more useful than owning browser workers. Submit only approved public URLs, request the smallest set of formats, and inspect body-level success rather than assuming a successful HTTP response means the product loaded.
Step 2: Validate the returned artifact
Confirm canonical URL, product evidence, locale, currency, and content completeness before parsing. Check the current Crawl pricing model when estimating cost per accepted product rather than cost per submitted URL.
Step 3: Feed the same normalization layer
The HTTP, Playwright, and managed methods should all produce the same internal source record. Provider-specific fields belong in the acquisition adapter, not in downstream catalog tables.
Why is the Zalando scraper returning empty or inconsistent data?
Empty data usually means the response is a different page state, the field moved into embedded JSON, the locale redirected, or the content requires rendering. Compare response.url, status, title, body length, and a saved hash. If Playwright finds the field but HTTP does not, inspect authorized network requests before adding more browser interactions.
Inconsistent prices often come from locale, currency, promotion, membership, or variant context. Store those dimensions with every observation. Nstdata's price-monitoring pipeline guide shows why normalized offers need provenance and acceptance rules.
How do you make the scraper safe and maintainable?
Use a bounded queue, low concurrency, retry only transient failures, and stop on persistent denial. Version the parser, keep fixtures for each supported page state, and alert on missing-field rate rather than silently writing nulls. The Robots Exclusion Protocol describes standardized robots rules, but compliance also requires terms, privacy, copyright, and internal review.
Conclusion
Start with one permitted product page and JSON-LD, add Playwright only when evidence shows the required field is client-rendered, and keep acquisition separate from normalization. Validate locale, canonical URL, and product identity before accepting a record. For a broader catalog job, use bounded URL discovery and the web-scraping best-practices checklist. If routing across multiple approved regions becomes a separate operational concern, consider Nstdata Proxy Manager without coupling proxy logic to the product schema.
Experience Nstdata — Start Your Free Trial Today
FAQ
Q: Is it legal to scrape Zalando?
Legality depends on the target, method, jurisdiction, terms, and data use. Collect only public or otherwise authorized pages and obtain legal guidance for the specific project.
Q: Should a Zalando scraper use Beautiful Soup or Playwright?
Use Beautiful Soup when the initial response or embedded JSON contains the required data; use Playwright only when JavaScript rendering is necessary. This reduces cost and failure surface.
Q: Why does the scraper see a different language or price?
Locale redirects, currency settings, geography, promotions, and variant context can change the page. Store those dimensions and reject records that do not match the requested market.
Q: How often should product pages be refreshed?
Refresh frequency should follow business need, page volatility, permission, and site load. Use change history to slow stable products and prioritize volatile ones.
Q: How do you prevent duplicate products?
Use a stable product identifier plus market and variant dimensions, canonicalize URLs conservatively, and make writes idempotent with content hashes.





