SYSTEM all green source cromwell.co.uk queue 12,904 pages p99 latency 218ms dataflirt.com · scraper/cromwell-co.uk
RUN · 31 active pipelines · cromwell.co.uk live

MRO catalogue data,
at warehouse scale.

We extract industrial tool specifications, stock availability, pricing tiers, and technical datasheets from Cromwell. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products extracted
482K /run
Price updates
1.2M /week
Stock checks
85K /day
Active pipelines
31
Uptime
99.94%
Data Dictionary

Every field we extract from cromwell.co.uk

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Product Specifications objects from cromwell.co.uk. All fields typed and schema-versioned.

skutitlebrandmanufacturer_part_numbercategorysub_categorydescriptionfeaturesimage_urlsdatasheet_url
product_specifications
● 200 OK
"sku": "KEN5827980K",
"title": "10mm Combination Spanner Chrome Vanadium",
"brand": "Kennedy",
"manufacturer_part_number": "5827980K",
"category": "Hand Tools",
"sub_category": "Spanners",
"description": "Manufactured from chrome vanadium steel with a highly polished mirror finish.",
"datasheet_url": "https://www.cromwell.co.uk/datasheets/KEN5827980K.pdf"
# skutitlebrandmanufacturer_part_numbercategorysub_category
1
2
3

Complete list of extractable fields for Pricing & Stock objects from cromwell.co.uk. All fields typed and schema-versioned.

skubase_price_exc_vatbase_price_inc_vatcurrencybulk_tier_1_qtybulk_tier_1_pricebulk_tier_2_qtybulk_tier_2_pricein_stockstock_levellead_time
pricing_& stock
● 200 OK
"sku": "KEN5827980K",
"base_price_exc_vat": 4.15,
"base_price_inc_vat": 4.98,
"currency": "GBP",
"bulk_tier_1_qty": 10,
"bulk_tier_1_price": 3.85,
"in_stock": true,
"stock_level": 142,
"lead_time": "Next Day Delivery"
# skubase_price_exc_vatbase_price_inc_vatcurrencybulk_tier_1_qtybulk_tier_1_price
1
2
3

Complete list of extractable fields for Technical Attributes objects from cromwell.co.uk. All fields typed and schema-versioned.

skumaterialfinishoverall_lengthsizemetric_imperialstandardweighthardness
technical_attributes
● 200 OK
"sku": "KEN5827980K",
"material": "Chrome Vanadium Steel",
"finish": "Mirror Polished",
"overall_length": "140mm",
"size": "10mm",
"metric_imperial": "Metric",
"standard": "DIN 3113",
"weight": "45g"
# skumaterialfinishoverall_lengthsizemetric_imperial
1
2
3

Complete list of extractable fields for Category Taxonomy objects from cromwell.co.uk. All fields typed and schema-versioned.

category_idnameparent_categorylevelurlproduct_counttop_brandsdescription
category_taxonomy
● 200 OK
"category_id": "cat_1024",
"name": "Combination Spanners",
"parent_category": "Spanners",
"level": 3,
"url": "/shop/hand-tools/combination-spanners/f/402",
"product_count": 845,
"top_brands": "['Kennedy', 'Facom', 'Stahlwille']"
# category_idnameparent_categorylevelurlproduct_count
1
2
3

Complete list of extractable fields for Search Results objects from cromwell.co.uk. All fields typed and schema-versioned.

keywordpositionskutitleprice_exc_vatbrandratingin_stockurl
search_results
● 200 OK
"keyword": "10mm spanner",
"position": 1,
"sku": "KEN5827980K",
"title": "10mm Combination Spanner Chrome Vanadium",
"price_exc_vat": 4.15,
"brand": "Kennedy",
"in_stock": true,
"url": "/shop/hand-tools/combination-spanners/10mm-combination-spanner/p/KEN5827980K"
# keywordpositionskutitleprice_exc_vatbrand
1
2
3

Capabilities

Everything you need from Cromwell: nothing you don't

Our Cromwell scraper extracts complex technical tables, tiered pricing structures, and stock availability across the entire MRO catalogue. Built with proxy rotation and JavaScript rendering to handle dynamic page elements.

Technical Specification Parsing

Extract complex HTML tables into structured key-value pairs. Normalise attributes like length, diameter, and material across different brands.

Tiered Pricing Extraction

Capture base prices, VAT calculations, and bulk discount tiers for large volume procurement analysis.

Stock Level Tracking

Monitor inventory availability, specific stock counts, and estimated lead times for out-of-stock items.

Datasheet Links

Extract direct URLs for Material Safety Data Sheets (MSDS) and technical PDF documents associated with PPE and chemicals.

Brand Hierarchy Mapping

Map private label brands like Kennedy and Sherwood alongside premium manufacturers like Facom and 3M.

Search Rank Scraping

Track product visibility and organic positioning for high-volume MRO keywords.

Category Traversal

Crawl entire category trees to build a complete map of the industrial supply taxonomy.

Incremental Updates

Run daily or weekly diffs. Only receive records where prices, stock levels, or specifications have changed.

Warehouse Delivery

Push structured data directly into your analytical environment via S3, BigQuery, or Snowflake.

// engagement pipeline

From product list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide category URLs, brand lists, or specific SKU sets. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, proxy rotation, and session management for cromwell.co.uk.

Validation & QA
d 4–6

Schema validation, null-rate checks, and technical attribute normalisation before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our Cromwell pipeline handles the hard parts

Industrial catalogues feature deep category trees and complex specification tables. Here is how we ensure data quality.

pipeline-monitor · cromwell.co.uk · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Anti-bot layer
UK residential proxy rotation

B2B distributors use edge protection to block automated traffic. Our crawlers route requests through UK-based residential IPs with realistic browser fingerprints to maintain high success rates without triggering rate limits.

Table normalisation
Structured key-value extraction

MRO products have dozens of technical attributes presented in variable HTML tables. Our pipeline parses these tables into clean JSON objects, standardising field names so 'Overall Length' and 'O/A Length' map to the same column.

JavaScript rendering
Playwright for dynamic pricing

Bulk discount tiers and real-time stock levels often load via asynchronous JavaScript requests. We use Playwright to execute page scripts and capture the exact data presented to a human buyer.

Change detection
Hash-based diffing

For catalogues exceeding 400,000 SKUs, full daily exports create unnecessary storage costs. We hash field values and only deliver records where price, stock, or specifications have changed since the previous run.

Monitoring & alerting
Pipeline health tracking

Every run emits structured logs. We alert on null-rate spikes, missing price fields, and layout changes. Our engineers update selectors before your downstream processes fail.

Applications

Who uses Cromwell data: and how

Teams across industries use cromwell.co.uk data to build competitive products and smarter operations.

01
Competitor Price Benchmarking

Industrial distributors track Cromwell pricing on exact-match SKUs to adjust their own margins and promotional strategies.

02
MRO Procurement Optimisation

Large manufacturing firms ingest catalogue data to build internal procurement systems and identify the cheapest suppliers for bulk orders.

03
Catalogue Enrichment

B2B marketplaces scrape technical specifications and datasheet links to populate their own product detail pages with accurate attributes.

04
Supply Chain Forecasting

Analysts track stock availability and lead times across critical PPE and tooling categories to predict supply chain bottlenecks.

05
Market Share Analysis

Tool manufacturers monitor category search rankings and shelf-space allocation to evaluate brand visibility against private labels.

06
Distributor Margin Tracking

Brands audit retail prices against their wholesale costs to estimate distributor margins and enforce pricing policies.

Why DataFlirt

"Cromwell houses one of the most comprehensive MRO catalogues in the UK, but extracting technical specifications requires structured pipeline engineering."

Industrial distributors often rely on outdated manual scraping methods. DataFlirt automates the extraction of complex specification tables, stock levels, and tiered pricing structures using residential proxies and full JavaScript rendering. Your engineers get clean data, not maintenance tickets.

Technical Spec

Cromwell scraper: technical capabilities

Everything supported by our cromwell.co.uk scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions for dynamic stock and pricing elements
Supported
CAPTCHA bypass
Automated 2Captcha + CapSolver integration
Supported
Residential proxy rotation
UK ISP-grade residential IPs rotated per request
Supported
Technical table parsing
Extraction of HTML attribute tables into JSON key-value pairs
Supported
Datasheet URL extraction
Capture of direct links to PDF MSDS and technical documents
Supported
Stock level tracking
Capture of precise inventory counts and lead time text
Supported
Change detection (diffs)
Emit records only when fields change between runs
Supported
Webhook delivery
HTTP POST per record for real-time processing
Supported
Negotiated B2B account pricing
Gated customer-specific discount rates requiring corporate login
Partial
Customer order history
Historical purchase data locked behind user authentication
Partial
Infrastructure

Infrastructure powering the Cromwell pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows. Combined via scrapy-playwright middleware.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across UK regions. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.

Cloud-Native Orchestration

Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested objects
CSV
Flat file with typed columns
XLS
Excel compatible format for procurement teams
Parquet
Columnar format for data warehouses
AWS S3
Direct bucket delivery
Webhook
HTTP POST per record
API
REST endpoint for on-demand querying
PostgreSQL
Upsert into your existing schema
BigQuery
Streamed directly into your dataset
Snowflake
Stage and COPY INTO workflow
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About cromwell.co.uk scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Cromwell legal?

Scraping publicly available information from Cromwell is generally permissible under UK law. DataFlirt targets only public, non-authenticated product, pricing, and specification data. We do not extract personal data or circumvent authentication walls. Clients should review Cromwell ToS and consult legal counsel for specific use cases.

How do you handle technical specification variations?

MRO products have highly variable attributes. We parse the specification tables dynamically into JSON objects. We can also apply custom mapping logic to normalise field names so your database receives consistent schema regardless of the manufacturer.

How fresh is the stock and pricing data?

We can configure pipelines to run at your required cadence. Full catalogue refreshes typically run weekly, while targeted SKU lists can be monitored daily or hourly for critical stock level changes.

Can you extract data for specific brands only?

Yes. We can scope the pipeline to target specific brand pages (e.g., Kennedy, 3M, DeWalt) or specific category URLs rather than crawling the entire site.

What is the minimum viable engagement?

Our smallest packages start at a defined SKU list (typically 5,000 to 20,000 SKUs) with weekly delivery. For full catalogue extraction, we price based on volume and delivery frequency.

Can I request a sample dataset before committing?

Absolutely. We provide a sample run of up to 500 SKUs or specific category pages as part of the pre-engagement scoping process. You can validate schema fit and data quality before signing any contract.

$ dataflirt scope --new-project --source=cromwell.co.uk ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or a continuous price-monitoring feed across 400K SKUs. We scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in industrial and mro

Services

Data Extraction for Every Industry

View All Services →