SYSTEM all green source hay.dk queue 3,492 pages p99 latency 214ms dataflirt.com · scraper/hay-dk
RUN . 14 active pipelines . hay.dk live

Hay.Dk data,
at warehouse scale.

We extract designer metadata, material specifications, variant pricing, and stock availability from hay.dk. Delivered as clean JSON, CSV, or Parquet to S3 or BigQuery on your cadence.

Products extracted
4,192 /run
Variant updates
18,304 /24h
Image assets
32,911 /run
Active pipelines
14
Uptime
99.94%
Data Dictionary

Every field we extract from hay.dk

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Product Listings objects from hay.dk. All fields typed and schema-versioned.

product_idskutitledesignercollectioncategorybase_pricecurrencydescriptioncare_instructionsdimensionsweight
product_listings
● 200 OK
"product_id": "HAY-M-1002",
"sku": "100293",
"title": "Mags Sofa",
"designer": "HAY Studio",
"category": "Furniture > Sofas",
"base_price": 18499.0,
"currency": "DKK",
"dimensions": "W268.5 x D95.5 x H67 cm"
# product_idskutitledesignercollectioncategory
1
2
3

Complete list of extractable fields for Variant Data objects from hay.dk. All fields typed and schema-versioned.

variant_idparent_skucolourmaterialfinishpricein_stockimage_urlsean
variant_data
● 200 OK
"variant_id": "VAR-8392",
"parent_sku": "100293",
"colour": "Steelcut Trio 133",
"material": "Kvadrat Fabric",
"price": 21499.0,
"in_stock": true,
"ean": "5710441238492"
# variant_idparent_skucolourmaterialfinishprice
1
2
3

Complete list of extractable fields for Designers objects from hay.dk. All fields typed and schema-versioned.

designer_idnamebioprofile_image_urlcollection_countproduct_skusorigin_countryactive_years
designers
● 200 OK
"designer_id": "DES-042",
"name": "Ronan & Erwan Bouroullec",
"origin_country": "France",
"collection_count": 4,
"product_skus": "['10492', '10493', '10494']",
"active_years": "2004-Present"
# designer_idnamebioprofile_image_urlcollection_countproduct_skus
1
2
3

Complete list of extractable fields for Materials objects from hay.dk. All fields typed and schema-versioned.

material_idnametypefinishsustainability_certcare_guidedurability_ratingimage_url
materials
● 200 OK
"material_id": "MAT-091",
"name": "FSC Certified Oak",
"type": "Wood",
"finish": "Matt Lacquered",
"sustainability_cert": "FSC Mix",
"care_guide": "Wipe with damp cloth"
# material_idnametypefinishsustainability_certcare_guide
1
2
3

Complete list of extractable fields for Categories objects from hay.dk. All fields typed and schema-versioned.

category_idnameparent_categoryurlproduct_countdescriptionhero_imagemetadata_tags
categories
● 200 OK
"category_id": "CAT-012",
"name": "Lounge Chairs",
"parent_category": "Furniture",
"product_count": 42,
"url": "https://hay.dk/category/furniture/lounge-chairs",
"metadata_tags": "['seating', 'living room', 'lounge']"
# category_idnameparent_categoryurlproduct_countdescription
1
2
3

Capabilities

Extract the complete Hay.dk catalogue

Our infrastructure navigates the complex front end of modern eCommerce platforms. We capture exact material finishes, dynamic pricing, and high resolution assets without missing a single variant.

Variant Mapping

Map every fabric, colour, and finish combination back to its parent product with accurate SKU relationships.

Dimension Parsing

Extract and normalise height, width, and depth metrics into structured numeric fields for spatial analysis.

Designer Metadata

Capture biographical data, collection associations, and origin details for every featured designer.

High-Res Imagery

Extract direct URLs for uncompressed product images, lifestyle shots, and specific variant textures.

Stock Tracking

Monitor inventory status across different regional storefronts and warehouse locations.

Localised Pricing

Capture region specific pricing, tax inclusions, and currency conversions across European markets.

3D Asset Links

Identify and extract URLs for AR models and CAD files embedded within product pages.

Material Specifications

Extract sustainability certifications, care instructions, and exact composite percentages for fabrics.

Scheduled Runs

Configure continuous pipelines at weekly or daily cadences to capture new collection drops immediately.

// engagement pipeline

From catalogue URL to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target categories, designer profiles, or specific regional storefronts. We design the schema together.

Pipeline Build
d 2–4

We configure Scrapy and Playwright crawlers to handle dynamic variant loading and geo-location routing.

Validation & QA
d 4–6

Schema validation, null-rate checks, and variant integrity tests before full production launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket or Snowflake stage on an agreed schedule.

Under the hood

Navigating modern eCommerce front ends

Design brands use complex JavaScript frameworks to render variants and regional pricing. Here is how we extract clean data.

pipeline-monitor · hay.dk · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Dynamic rendering
Playwright execution for SPA content

Furniture variants load dynamically based on user interaction. We run full Playwright browser sessions to trigger JavaScript events, ensuring every fabric and frame combination is captured.

Asset pipelines
High resolution image extraction

We bypass compressed thumbnails to locate the source URLs of high resolution assets and 3D models, delivering clean links ready for your internal DAM systems.

Geo-routing
Regional proxy targeting

Pricing and availability change based on the visitor location. Our residential proxy pools target specific European regions to capture accurate local market data.

Schema stability
Resilient DOM selectors

We utilise multiple fallback chains for field extraction, combining CSS selectors with structured JSON-LD data to survive minor site updates.

Change detection
Delta exports

We maintain a hash index of last seen values. Subsequent runs only push modifications to pricing or stock, reducing downstream processing load.

Applications

Who uses furniture catalogue data

Teams across industries use hay.dk data to build competitive products and smarter operations.

01
Competitor Intelligence

Retailers track pricing, material choices, and collection launches to benchmark their own product lines.

02
Interior Design Aggregation

Marketplaces ingest structured dimension and material data to build comprehensive search filters for design professionals.

03
Trend Analysis

Analysts track the introduction of new materials, colours, and sustainability certifications across seasons.

04
AI Model Training

Machine learning teams use structured metadata paired with high resolution imagery to train spatial recognition and style matching models.

05
Supply Chain Visibility

Procurement teams monitor stock levels across regional stores to predict manufacturing constraints.

06
Retail Arbitrage

Distributors track cross border pricing disparities to optimise their purchasing strategy.

Why DataFlirt

"Hay.dk holds the definitive digital record for contemporary Danish design, but extracting structured variant data requires navigating complex front end architecture."

Extracting furniture catalogues involves more than parsing text. We capture high resolution imagery, exact material specifications, designer metadata, and dynamic variant pricing across multiple regional storefronts. DataFlirt manages the complete extraction lifecycle so your team can focus on analysis.

Technical Spec

Hay.dk scraper technical specifications

Everything supported by our hay.dk scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for dynamic variant loading
Supported
Variant expansion
Iterates through all colour and material combinations per product
Supported
High-res image URLs
Extracts uncompressed source images rather than thumbnails
Supported
3D model extraction
Captures links to AR/CAD files embedded in the page source
Supported
Multi-region pricing
Captures EUR, DKK, GBP, and USD via proxy localisation
Supported
Change detection
Emits only changed records since the previous pipeline run
Supported
B2B Trade Pricing
Trade discounts require authenticated dealer accounts
Partial
Order History
Customer specific order data sits behind a login wall
Partial
Infrastructure

Infrastructure powering the pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy and Playwright

Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering and variant interaction flows.

Residential Proxy Pools

We maintain proxy pools across European regions to ensure accurate capture of localised pricing and availability.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Nested structures ideal for complex variant hierarchies
CSV
Flat file format for simple catalogue ingestion
XLS
Excel compatible format for manual review
Parquet
Columnar format optimised for analytical queries
AWS S3
Direct delivery to your cloud storage buckets
Webhook
HTTP POST delivery for real time catalogue updates
API
REST endpoints to query your extracted datasets
PostgreSQL
Direct database insertion with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About hay.dk scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Hay.dk legal?

Scraping publicly available product information is generally permissible. DataFlirt extracts only public catalogue data, pricing, and specifications. We do not bypass authentication walls or extract personal user data. Clients should review target terms of service and consult legal counsel.

How do you capture all product variants?

We utilise headless browsers to interact with the page exactly as a user would, clicking through every available fabric, colour, and finish option to capture the specific SKU, price, and image associated with that combination.

Can you extract 3D models and AR files?

Yes. If the product page embeds links to .glb, .gltf, or .usdz files for augmented reality viewing, our pipeline identifies and extracts those source URLs for your use.

How frequently can the pipeline run?

For a complete catalogue of this size, daily or weekly runs are standard. We can configure specific categories to run at higher frequencies if you are tracking limited edition drops or flash sales.

Do you provide B2B trade pricing?

No. Trade pricing on Hay.dk requires an authenticated dealer login. We only extract the publicly visible retail pricing available to standard consumers.

What format is the dimension data delivered in?

We parse raw text strings like 'W268.5 x D95.5 x H67 cm' into discrete numeric fields for width, depth, and height, standardising the unit of measurement across the dataset.

$ dataflirt scope --new-project --source=hay.dk ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a complete catalogue export or continuous tracking of new collections and pricing changes. Tell us your requirements.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in furniture

Services

Data Extraction for Every Industry

View All Services →