SYSTEM all green source patagonia.com queue 12,841 pages p99 latency 184ms dataflirt.com · scraper/patagonia-com
RUN · 14 active pipelines · patagonia.com live

Patagonia data,
at warehouse scale.

We extract apparel listings, Worn Wear used inventory, material compositions, Footprint Chronicles data, and pricing signals from Patagonia. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products extracted
14,291 /run
Worn Wear items
8,422 /day
Stock updates
42,104 /24h
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from patagonia.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Main Apparel objects from patagonia.com. All fields typed and schema-versioned.

skutitlecategorysub_categorypricecurrencydiscount_pctavailable_coloursavailable_sizesfit_typeweight_gramsdescriptionbullet_pointsimage_urlsurl
main_apparel
● 200 OK
"sku": "84212",
"title": "Men's Nano Puff® Jacket",
"category": "Mens",
"sub_category": "Jackets & Vests",
"price": 239.0,
"currency": "USD",
"available_colours": "['Black', 'Forge Grey', 'Nouveau Green']",
"weight_grams": 337
# skutitlecategorysub_categorypricecurrency
1
2
3

Complete list of extractable fields for Worn Wear objects from patagonia.com. All fields typed and schema-versioned.

worn_wear_idoriginal_skutitlecondition_gradecondition_descriptionpriceoriginal_priceyear_madesizecolourflaws_notedin_stock
worn_wear
● 200 OK
"worn_wear_id": "WW-84212-M-BLK",
"original_sku": "84212",
"condition_grade": "Excellent",
"price": 119.0,
"original_price": 239.0,
"year_made": 2021,
"flaws_noted": "None",
"in_stock": true
# worn_wear_idoriginal_skutitlecondition_gradecondition_descriptionprice
1
2
3

Complete list of extractable fields for Materials & ESG objects from patagonia.com. All fields typed and schema-versioned.

skurecycled_pctfair_trade_certifiedbluesign_approvedfabric_compositioninsulation_typeorigin_countryfactory_namefactory_locationenvironmental_impact_notes
materials_& esg
● 200 OK
"sku": "84212",
"recycled_pct": 100,
"fair_trade_certified": true,
"bluesign_approved": true,
"fabric_composition": "1.4-oz 20-denier 100% recycled polyester ripstop",
"insulation_type": "60-g PrimaLoft Gold Insulation Eco",
"origin_country": "Vietnam",
"factory_name": "Pungkook Corporation"
# skurecycled_pctfair_trade_certifiedbluesign_approvedfabric_compositioninsulation_type
1
2
3

Complete list of extractable fields for Reviews & Ratings objects from patagonia.com. All fields typed and schema-versioned.

review_idskureviewer_nicknameratingfit_slider_valuereview_titlereview_bodyprimary_uselocationdate_postedhelpful_votes
reviews_& ratings
● 200 OK
"review_id": "REV-992144",
"sku": "84212",
"rating": 5,
"fit_slider_value": "True to size",
"review_title": "Perfect for layering",
"primary_use": "Everyday Wear",
"date_posted": "2023-11-12",
"helpful_votes": 14
# review_idskureviewer_nicknameratingfit_slider_valuereview_title
1
2
3

Complete list of extractable fields for Category & SERP objects from patagonia.com. All fields typed and schema-versioned.

category_idkeywordpositionskutitlepricebadge_textcolour_countproduct_urlscraped_at
category_& serp
● 200 OK
"category_id": "mens-jackets-vests",
"position": 1,
"sku": "84212",
"title": "Men's Nano Puff® Jacket",
"price": 239.0,
"badge_text": "Bestseller",
"colour_count": 8,
"scraped_at": "2023-11-14T08:12:00Z"
# category_idkeywordpositionskutitleprice
1
2
3

Capabilities

Extract the complete Patagonia catalogue

Our scraper handles the complexities of Patagonia's Salesforce Commerce Cloud backend, dynamic Worn Wear inventory, and nested material composition data.

Full Catalogue Extraction

Title, price, descriptions, fit matrices, and weight specifications extracted across all primary categories and sub-categories.

Worn Wear Tracking

Monitor used inventory drops, condition grades, and secondary market pricing across the Worn Wear platform.

Material & ESG Data

Extract recycled material percentages, Fair Trade certifications, and Footprint Chronicles factory origin data per SKU.

Colourway & Size Matrices

Map every available size and colour combination, including out-of-stock variants and seasonal colour additions.

Review & Fit Analysis

Capture text reviews, star ratings, and the critical customer fit slider (runs small/large) for product analysis.

Web Specials & Pricing

Track discount percentages, Web Specials inventory, and historical price changes across the catalogue.

Regional Storefronts

Support for patagonia.com, eu.patagonia.com, and regional variants for global pricing comparisons.

Asset Extraction

High-resolution product imagery and technical diagram URLs captured and delivered alongside metadata.

Automated Diffs

Receive only changed records on subsequent runs. Track new product launches and discontinued items automatically.

// engagement pipeline

From category URL to structured data

Brief in. Clean data out.

Define Scope
d 0

Specify target categories, Worn Wear segments, or specific data points like material composition and factory origins.

Pipeline Build
d 2–4

We configure crawlers to handle Demandware pagination, dynamic inventory endpoints, and rate limits.

Validation & QA
d 4–6

Schema validation, null-rate checks on nested material tags, and price-outlier detection before production.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Handling Patagonia's architecture

Extracting structured data from modern headless commerce platforms requires specific infrastructure. Here is how we maintain pipeline stability.

pipeline-monitor · patagonia.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Commerce Cloud
Salesforce Demandware API interception

Patagonia uses Salesforce Commerce Cloud. Instead of parsing complex DOM trees for inventory, we intercept the underlying JSON API responses for accurate, real-time stock and pricing data across all colourways.

Worn Wear
Dynamic inventory monitoring

Worn Wear inventory is highly dynamic with single-SKU items dropping frequently. We utilise high-frequency polling on specific category endpoints to capture items before they sell out.

Rate limiting
Distributed request timing

Aggressive crawling triggers Akamai edge blocks. We distribute requests across our residential proxy network and pace page loads to mimic organic browsing patterns.

Data nesting
Complex schema flattening

Material compositions and factory origin data are deeply nested within product pages. Our extraction logic normalises these fields into flat, queryable columns for your data warehouse.

Regional pricing
Geo-targeted session management

To capture accurate EU vs US pricing, we maintain isolated browser sessions with region-specific cookies and exit nodes, preventing currency redirect loops.

Applications

Who uses Patagonia data

Teams across industries use patagonia.com data to build competitive products and smarter operations.

01
ESG Benchmarking

Apparel brands analyse Patagonia's material compositions, recycled percentages, and factory disclosures to benchmark their own sustainability initiatives.

02
Circular Economy Analysis

Retail strategists track Worn Wear pricing models and inventory velocity to understand the economics of brand-owned resale platforms.

03
Competitor Pricing

Outdoor brands monitor Web Specials and seasonal discount depth to optimise their own promotional calendars.

04
Assortment Planning

Merchandisers track colourway availability and size-run depth to identify trending styles in the outdoor apparel sector.

05
Supply Chain Intelligence

Analysts map the Footprint Chronicles data to understand global textile sourcing and manufacturing dependencies.

06
Product Development

Design teams mine customer reviews and fit-slider data to identify common design flaws or sizing issues in technical outerwear.

Why DataFlirt

"Patagonia's catalogue is the industry benchmark for sustainable material sourcing and circular economy pricing — but it requires a pipeline to analyse at scale."

Most teams underestimate the investment required: reliable Patagonia scraping requires handling Demandware endpoints, Worn Wear's dynamic inventory, and nested size-colour matrices. DataFlirt absorbs that complexity so your engineers focus on analysis.

Technical Spec

Patagonia scraper — technical capabilities

Everything supported by our patagonia.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

API interception
Direct extraction from Demandware JSON endpoints for accurate variant data
Supported
Worn Wear extraction
Capture of unique used-item IDs, condition grades, and original price comparisons
Supported
Material tag parsing
Extraction of specific certification tags (Bluesign, Fair Trade, Recycled)
Supported
Regional pricing
Extraction across US, EU, and UK storefronts using localised proxies
Supported
Review pagination
Full capture of historical reviews and fit-slider metrics
Supported
Change detection
Hash-based diffing to identify new product drops or price changes
Supported
Pro Program pricing
Extraction of discounted pricing tiers requiring professional credential verification
Partial
Customer order history
Extraction of personal purchase history or Ironclad Guarantee return status
Partial
Infrastructure

Infrastructure powering the pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheusSnowflakeBigQuery
Scrapy + Playwright Stack

Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows. Combined via scrapy-playwright middleware.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.

Cloud-Native Orchestration

Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested — schema versioned per run
CSV
Flat file with typed columns — Excel/Sheets compatible
XLS
Excel format for business teams and analysts
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery — compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoint for querying extracted catalogue data
BigQuery
Streamed directly into your dataset with schema auto-detect
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About patagonia.com scraping, legality, and pipeline operations.

Ask us directly →
Can you extract data from the Worn Wear platform?

Yes. We track the Worn Wear subdomain, extracting unique item IDs, condition grades, pricing, and specific noted flaws. This requires distinct pipeline logic from the main apparel catalogue due to the single-SKU nature of used inventory.

Do you capture environmental and material data?

Yes. We extract the material composition text, recycled percentages, and binary flags for certifications like Fair Trade and Bluesign, as well as factory origin data from the Footprint Chronicles.

How frequently can you update pricing and stock?

For the main catalogue, we typically run daily diffs. For high-velocity segments like Worn Wear or Web Specials, we can configure hourly polling pipelines to capture inventory before it sells out.

Do you support regional Patagonia sites?

Yes. We can extract from patagonia.com, eu.patagonia.com, and other regional variants. We use geo-located proxies to ensure accurate local pricing and prevent forced currency redirects.

Can you track out-of-stock items?

Yes. We map the entire size and colourway matrix for each product, recording null or false availability flags for variants that are currently out of stock.

What format is the data delivered in?

We deliver structured JSON, CSV, or Parquet files directly to your AWS S3 bucket, Google Cloud Storage, or data warehouse (BigQuery/Snowflake) on a schedule.

$ dataflirt scope --new-project --source=patagonia.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue extract or continuous monitoring of Worn Wear inventory — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in fashion and apparel

Services

Data Extraction for Every Industry

View All Services →