SYSTEM all green source santacole.com queue 3,492 pages p99 latency 218ms dataflirt.com · scraper/santacole-com
RUN - 12 active pipelines - santacole.com live

Santa & Cole data,
at warehouse scale.

We extract high-end lighting catalogues, designer portfolios, material specifications, and photometric files from santacole.com. Delivered as clean JSON, CSV, or Parquet.

Products extracted
4,192 /run
Designers tracked
114
CAD/IES files mapped
8,291
Active pipelines
12
Uptime
99.98%
Data Dictionary

Every field we extract from santacole.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Lighting Products objects from santacole.com. All fields typed and schema-versioned.

skunamedesignercategorysub_categorydescriptionpricecurrencymaterialsdimensionslight_sourcedimmableip_ratingweightimages
lighting_products
● 200 OK
"sku": "SC-LMP-042",
"name": "Cesta",
"designer": "Miguel Mila",
"price": 850.0,
"currency": "EUR",
"ip_rating": "IP20",
"dimmable": true
# skunamedesignercategorysub_categorydescription
1
2
3

Complete list of extractable fields for Furniture & Accessories objects from santacole.com. All fields typed and schema-versioned.

skunamedesignercategorymaterialsfinishesdimensionsweightcare_instructionspricecurrencyimagesassembly_required
furniture_& accessories
● 200 OK
"sku": "SC-FURN-011",
"name": "Cadaques",
"designer": "Federico Correa",
"category": "Sofas",
"materials": "['Wood', 'Fabric']",
"price": 3200.0,
"currency": "EUR"
# skunamedesignercategorymaterialsfinishes
1
2
3

Complete list of extractable fields for Designer Profiles objects from santacole.com. All fields typed and schema-versioned.

designer_idnamebiographybirth_yearnationalityawardsproducts_designedprofile_image_urlstudio_url
designer_profiles
● 200 OK
"designer_id": "D-MM-01",
"name": "Miguel Mila",
"birth_year": 1931,
"nationality": "Spanish",
"awards": "['National Design Award']",
"products_designed": "['Cesta', 'TMM']"
# designer_idnamebiographybirth_yearnationalityawards
1
2
3

Complete list of extractable fields for Technical Files objects from santacole.com. All fields typed and schema-versioned.

skuproduct_namefile_typefile_urlfile_sizelanguageformatlast_updated
technical_files
● 200 OK
"sku": "SC-LMP-042",
"product_name": "Cesta",
"file_type": "Photometric",
"file_url": "https://santacole.com/files/cesta.ies",
"format": "IES",
"language": "EN"
# skuproduct_namefile_typefile_urlfile_sizelanguage
1
2
3

Complete list of extractable fields for Variant Finishes objects from santacole.com. All fields typed and schema-versioned.

parent_skuvariant_skufinish_namematerialcolour_heximage_urlprice_modifierstock_status
variant_finishes
● 200 OK
"parent_sku": "SC-LMP-042",
"variant_sku": "SC-LMP-042-CHE",
"finish_name": "Cherry Wood",
"material": "Wood",
"colour_hex": "#5C4033",
"stock_status": "in_stock"
# parent_skuvariant_skufinish_namematerialcolour_heximage_url
1
2
3

Capabilities

Extract design specifications with engineering precision

Santa & Cole relies on visually heavy, JavaScript-rendered product pages with complex variant selectors. Our infrastructure parses the underlying state to deliver structured technical data.

Product Catalogues

Extract full product metadata including SKU, name, description, category, dimensions, and weight.

Designer Mapping

Link products to designer profiles, extracting biographical data, awards, and historical portfolios.

Technical Specifications

Capture light source details, dimming capabilities, IP ratings, and power requirements for lighting fixtures.

Photometric Files

Locate and map IES and LDT file URLs for architectural lighting simulation workflows.

3D Model Links

Extract URLs for CAD models, SketchUp files, and BIM objects associated with each product.

Material Finishes

Map parent-child variant relationships for different wood, metal, and fabric finishes.

Pricing Signals

Capture retail pricing, currency, and tax inclusions across different geographic regions.

Region Localisation

Extract localised catalogue variations for European, North American, and Asian markets.

Scheduled Syncs

Run weekly or monthly pipelines to track new product launches and discontinued items.

// engagement pipeline

From designer catalogue to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Specify the categories, designer profiles, or asset types required from the Santa & Cole catalogue.

Pipeline Build
d 2–4

We configure Playwright crawlers to handle bespoke frontend routing and variant state hydration.

Validation & QA
d 4–6

Schema validation ensures technical specifications and asset URLs map correctly to parent SKUs.

Delivery
ongoing

Structured JSON or Parquet files pushed to your S3 bucket or data warehouse on schedule.

Under the hood

Navigating boutique eCommerce architecture

High-end design sites prioritise visual experience over standard DOM structures. Here is how we extract structured data from bespoke frontends.

pipeline-monitor · santacole.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
JS Rendering
Full Playwright execution for visual frontends

Boutique design websites rely heavily on client-side rendering. We run full browser sessions to execute JavaScript, await animation frames, and extract the hydrated application state.

Variant State Extraction
Mapping complex material finishes

Product variations (like wood type or fabric colour) often exist only as JavaScript objects rather than separate HTML pages. We intercept XHR requests and parse window objects to map all variants.

Asset Mapping
Aggregating technical files

Architectural workflows require photometric data and 3D models. Our pipeline identifies, validates, and maps IES, LDT, and DWG file URLs directly to the corresponding product SKU.

Region Localisation
Handling geo-routed catalogues

Santa & Cole displays different pricing and availability based on IP location. We use region-specific residential proxies to extract accurate data for your target market.

Change Detection
Tracking catalogue updates

We maintain hash indexes of product states to detect new finishes, discontinued items, or pricing updates, delivering clean diffs rather than redundant full exports.

Applications

Who uses Santa & Cole data

Teams across industries use santacole.com data to build competitive products and smarter operations.

01
Architecture & BIM Platforms

Aggregate technical specifications and photometric files for inclusion in architectural design software.

02
Competitor Pricing

High-end furniture manufacturers monitor retail pricing strategies across different European markets.

03
Interior Design Aggregators

Populate digital catalogues with accurate dimensions, materials, and high-resolution imagery.

04
Market Research

Analyse trends in materials, designer collaborations, and product lifecycle within the luxury lighting sector.

05
Supplier Audits

Distributors verify their listed specifications and pricing against the official manufacturer catalogue.

06
AI Training Data

Train computer vision models on high-quality furniture imagery and corresponding descriptive metadata.

Why DataFlirt

"Design brands embed critical technical data inside bespoke, visual-first interfaces. Extracting it requires infrastructure that parses state, not just HTML."

Boutique manufacturers like Santa & Cole do not offer public APIs for their catalogues. We build managed extraction pipelines that handle JavaScript rendering, variant state mapping, and asset aggregation so your engineering team receives clean, warehouse-ready schemas.

Technical Spec

Santa & Cole scraper - technical capabilities

Everything supported by our santacole.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions to handle client-side routing and animations
Supported
Variant mapping
Extract all material and finish combinations per product
Supported
IES/LDT file extraction
Map photometric data URLs for lighting simulation
Supported
Multi-region pricing
Capture EUR, USD, and GBP pricing via residential proxies
Supported
High-res image aggregation
Extract uncompressed image URLs from the CDN
Supported
Change detection
Identify new product launches and discontinued SKUs
Supported
Designer portfolio mapping
Link products to historical designer profiles
Supported
Assembly instruction PDFs
Extract URLs for technical manuals and care guides
Supported
B2B trade pricing
Requires an approved trade account login to access wholesale rates
Partial
Wholesale inventory levels
Real-time stock data is gated behind the internal ERP system
Partial
Infrastructure

Infrastructure powering the extraction

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Handles complex DOM traversal and state extraction on visually heavy, JavaScript-dependent product pages.

Asset Aggregation Pipeline

Validates and maps thousands of technical files, ensuring CAD and IES links resolve correctly before delivery.

Cloud-Native Orchestration

Containerised workloads managed by Kubernetes and Airflow ensure reliable execution and SLA adherence.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Nested structure ideal for complex variant mapping
CSV
Flat files for pricing and basic catalogue analysis
XLS
Excel format for procurement and merchandising teams
Parquet
Columnar format for efficient data warehouse querying
AWS S3
Direct delivery to your cloud storage infrastructure
Webhook
HTTP POST notifications for catalogue updates
API
REST endpoints to query extracted product data
PostgreSQL
Direct database inserts for immediate application use
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About santacole.com scraping, legality, and pipeline operations.

Ask us directly →
Can you extract photometric files for all lighting products?

Yes. We locate and map URLs for IES and LDT files wherever Santa & Cole provides them on the product page or technical specification tabs.

How do you handle the different material finishes?

We parse the frontend state to map every parent-child variant relationship, capturing the specific SKU, finish name, and image URL for each material option.

Do you scrape pricing for different countries?

Yes. By routing requests through region-specific residential proxies, we extract localised pricing and currency data for your target markets.

Can I get the 3D CAD models?

We extract the direct download URLs for 3D models (DWG, SketchUp, etc.) and associate them with the correct product SKU in the final dataset.

How often can the catalogue be refreshed?

For boutique catalogues like Santa & Cole, we typically recommend weekly or monthly runs to capture new releases and pricing adjustments without unnecessary compute overhead.

Do you bypass login walls for trade pricing?

No. We only extract publicly available retail data. B2B trade pricing and wholesale inventory levels require authenticated access and are not supported.

$ dataflirt scope --new-project --source=santacole.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue export or continuous pricing syncs across regions, we build and operate the pipeline. Tell us your requirements.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in lighting

Services

Data Extraction for Every Industry

View All Services →