SYSTEM all green source on-running.com queue 4,192 pages p99 latency 214ms dataflirt.com · scraper/on-running-com
RUN · 14 active pipelines · on-running.com live

On-Running data,
at warehouse scale.

We extract product listings, pricing signals, sizing inventory, and technical specifications from on-running.com. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products extracted
3,481 /day
Inventory checks
89.2K /24h
Review records
42K /run
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from on-running.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Footwear Models objects from on-running.com. All fields typed and schema-versioned.

product_idmodel_namecategorypricecurrencycolourway_idcolour_nameavailable_sizesweight_gheel_to_toe_drop_mmcushioning_leveldescriptionimage_urlsurl
footwear_models
● 200 OK
"product_id": "M-Cloudmonster-2",
"model_name": "Cloudmonster 2",
"category": "Road Running",
"price": 170.0,
"currency": "USD",
"colour_name": "Undyed-White | Flame",
"weight_g": 295,
"heel_to_toe_drop_mm": 6
# product_idmodel_namecategorypricecurrencycolourway_id
1
2
3

Complete list of extractable fields for Apparel & Accessories objects from on-running.com. All fields typed and schema-versioned.

product_idproduct_namegenderfit_typematerialspricecurrencycolour_nameavailable_sizescare_instructionsweather_conditionfeaturesurl
apparel_& accessories
● 200 OK
"product_id": "W-Weather-Jacket",
"product_name": "Weather Jacket",
"gender": "Women",
"fit_type": "Athletic",
"price": 240.0,
"currency": "USD",
"colour_name": "Black",
"weather_condition": "Wind, Rain"
# product_idproduct_namegenderfit_typematerialsprice
1
2
3

Complete list of extractable fields for Inventory & Pricing objects from on-running.com. All fields typed and schema-versioned.

skuproduct_idregion_codepricelist_pricecurrencysizein_stockstock_levellow_stock_warningscraped_at
inventory_& pricing
● 200 OK
"sku": "3ME10120108-US10",
"region_code": "US",
"price": 170.0,
"currency": "USD",
"size": "US M 10",
"in_stock": true,
"low_stock_warning": false,
"scraped_at": "2026-05-12T10:15:00Z"
# skuproduct_idregion_codepricelist_pricecurrency
1
2
3

Complete list of extractable fields for Technical Specs objects from on-running.com. All fields typed and schema-versioned.

product_idmodel_namecushioningrunning_profilespeedboard_typeupper_materialmidsole_techoutsole_patternsustainability_pctrecycled_content
technical_specs
● 200 OK
"model_name": "Cloudsurfer",
"cushioning": "Plush",
"running_profile": "Daily Training",
"speedboard_type": "None (CloudTec Phase)",
"midsole_tech": "Helion superfoam",
"sustainability_pct": 30,
"recycled_content": "100% recycled polyester upper"
# product_idmodel_namecushioningrunning_profilespeedboard_typeupper_material
1
2
3

Complete list of extractable fields for Reviews & Ratings objects from on-running.com. All fields typed and schema-versioned.

review_idproduct_idratingreview_titlereview_textreviewer_namereview_dateverified_buyerfit_feedbackcomfort_feedbackquality_feedback
reviews_& ratings
● 200 OK
"review_id": "REV-982341",
"product_id": "M-Cloudmonster-2",
"rating": 5,
"review_title": "Maximum cushioning",
"verified_buyer": true,
"fit_feedback": "True to size",
"comfort_feedback": "Excellent",
"review_date": "2026-04-20"
# review_idproduct_idratingreview_titlereview_textreviewer_name
1
2
3

Capabilities

Extracting the complete On-Running catalogue

On-Running uses a modern headless commerce architecture. We intercept internal state and API responses to extract clean, structured data for every shoe, apparel item, and regional storefront.

Full Catalogue Extraction

Extract every shoe model, apparel item, and accessory across the entire on-running.com domain.

Variant & Colourway Mapping

Capture parent-child relationships for every colourway and size combination per product model.

Real-Time Inventory Tracking

Monitor stock availability and low-stock warnings at the exact SKU and size level.

Technical Specifications

Extract deep product metadata including heel-to-toe drop, weight, CloudTec configuration, and Speedboard details.

Multi-Region Pricing

Capture geo-localised pricing and availability across US, EU, UK, and APAC regional storefronts.

Sustainability & Material Data

Track recycled content percentages and material composition metrics listed on product detail pages.

Review & Rating Aggregation

Compile customer feedback, star ratings, and specific fit/comfort metrics from verified buyers.

Headless Commerce Support

Intercept Next.js hydration state and internal GraphQL queries to capture data before it reaches the DOM.

Scheduled + Streaming Modes

Run one-off bulk exports or configure continuous pipelines at hourly or daily cadences.

// engagement pipeline

From product URL to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Select target regions, product categories, and update frequencies. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy and Playwright to handle Next.js state interception, proxy rotation, and regional localisation.

Validation & QA
d 4–6

Schema validation, null-rate checks, and variant mapping verification before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our pipeline handles headless commerce

Modern single-page applications hide data in complex state objects rather than HTML. Here is how we extract structured data from on-running.com.

pipeline-monitor · on-running.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
State interception
Extracting Next.js hydration data

Instead of parsing complex DOM structures, our Playwright scripts intercept the internal JSON state used to hydrate the frontend. This provides cleaner, more comprehensive data including hidden inventory metrics.

Geo-localisation
Regional proxy targeting

Pricing and availability change based on the user location. We route requests through residential proxies in specific target countries to capture accurate local market data.

Variant handling
Dynamic sizing matrix extraction

Shoe sizes and colourways are loaded dynamically. We map every possible SKU combination, ensuring complete coverage of the product matrix.

Change detection
Only re-scrape what changes

We maintain a hash index of last-seen values per field. Subsequent runs only push diffs, reducing downstream processing load for inventory and price updates.

Monitoring & alerting
24/7 pipeline health

Every run emits structured logs to our observability stack. We alert on null-rate spikes, schema drift, and coverage drops, responding before you notice.

Applications

Who uses On-Running data — and how

Teams across industries use on-running.com data to build competitive products and smarter operations.

01
Competitor Price Monitoring

Footwear brands track On-Running pricing strategies across different global markets to adjust their own positioning.

02
Assortment & Gap Analysis

Retailers analyse colourway availability and model lifecycles to optimise their own buying and merchandising strategies.

03
Inventory & Restock Tracking

Track stockouts and restock velocities at the size level to estimate demand and production volumes.

04
Technical Spec Benchmarking

Product teams extract weight, heel drop, and material data to benchmark against competing running shoe models.

05
Sentiment Analysis

Marketing teams aggregate review data to understand customer sentiment regarding fit, durability, and comfort.

06
AI Training Data

Machine learning teams use structured product descriptions and technical specs to train retail recommendation engines.

Why DataFlirt

"On-Running's headless architecture hides deep technical specifications and regional inventory data behind complex API calls — we structure it into queryable tables."

Extracting data from modern single-page applications requires intercepting internal state and mimicking localised user sessions. DataFlirt manages the residential proxies, API interception, and daily schema maintenance so your engineers focus on analysis, not infrastructure.

Technical Spec

On-Running scraper — technical capabilities

Everything supported by our on-running.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Next.js state extraction
Direct extraction of application hydration state for clean JSON data
Supported
Residential proxy rotation
ISP-grade residential IPs for accurate regional pricing capture
Supported
Multi-region geo-targeting
Support for US, UK, EU, and APAC storefront localisation
Supported
Variant mapping
Parent to child SKU relationships for all colours and sizes
Supported
Change detection (diffs)
Hash-based diff: only emit records with changed fields since last run
Supported
Webhook delivery
HTTP POST per record or batch for inventory alerting
Supported
GraphQL interception
Capture underlying API requests powering the frontend experience
Supported
Customer purchase history
Requires authenticated user sessions and violates privacy policies
Partial
Cyclon subscription details
Account-specific subscription management and billing data
Partial
Infrastructure

Infrastructure powering the pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheusBigQuery
API & State Interception

Playwright handles JavaScript execution and intercepts internal API calls and state objects, bypassing fragile DOM parsing.

Geo-Distributed Proxy Pools

Residential ISP proxies routed by target country ensure accurate pricing and inventory data per region.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested — schema versioned per run
CSV
Flat file with typed columns — Excel/Sheets compatible
XLS
Formatted spreadsheet for immediate business analysis
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery — compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoint to query your extracted dataset
BigQuery
Streamed directly into your dataset with schema auto-detect
PostgreSQL
Upsert into your existing schema with conflict resolution
Snowflake
Stage + COPY INTO workflow — incremental or full-replace
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About on-running.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping on-running.com legal?

Scraping publicly available product, pricing, and review data is generally permissible. DataFlirt targets only public, non-authenticated catalogue data. We do not extract personal data or circumvent authentication walls. Clients should review terms of service and consult legal counsel for specific use cases.

How do you handle the Next.js frontend?

We use Playwright to execute the JavaScript and intercept the internal JSON hydration state and GraphQL queries. This yields structured data directly from the application state rather than relying on fragile CSS selectors.

Can you extract regional pricing?

Yes. We route requests through residential proxies located in the target region (e.g., US, UK, Germany) to capture localised pricing, currency, and availability.

How fresh is the inventory data?

We can configure pipelines to run at hourly cadences for specific high-priority SKUs, capturing near real-time stock availability and low-stock warnings.

Do you extract technical specs like heel drop and weight?

Yes. We extract all structured metadata provided on the product detail pages, including weight, heel-to-toe drop, cushioning level, and material composition.

What is the minimum viable engagement?

Our minimum engagement typically covers a weekly full-catalogue extraction for a single region. Contact us with your specific volume and frequency requirements for a scoped quote.

Can I request a sample dataset?

Yes. We provide a sample run covering a subset of products across footwear and apparel during the scoping process, allowing you to validate the schema before committing.

$ dataflirt scope --new-project --source=on-running.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a daily inventory feed or a comprehensive technical specification catalogue — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in fashion and apparel

Services

Data Extraction for Every Industry

View All Services →