SYSTEM all green source weidmueller.com queue 12,492 pages p99 latency 215ms dataflirt.com · scraper/weidmueller-com
RUN · 14 active pipelines · weidmueller.com live

Weidmüller data,
at warehouse scale.

We extract terminal blocks, relays, automation components, technical specs, and CAD metadata from Weidmüller. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake.

Components extracted
84K /run
Spec attributes mapped
1.2M /run
CAD links indexed
62K /run
Active pipelines
14
Uptime
99.94%
Data Dictionary

Every field we extract from weidmueller.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Component Base Data objects from weidmueller.com. All fields typed and schema-versioned.

order_numbertype_designationproduct_groupshort_descriptionlong_descriptiongtin_eanqty_per_packeclass_codeunspsc_codeimage_urlproduct_family_url
component_base data
● 200 OK
"order_number": "1020000000",
"type_designation": "WDU 2.5",
"product_group": "Terminal Blocks",
"gtin_ean": "4008190115354",
"qty_per_pack": 100,
"eclass_code": "27-14-11-20",
"unspsc_code": "39-12-14-10"
# order_numbertype_designationproduct_groupshort_descriptionlong_descriptiongtin_ean
1
2
3

Complete list of extractable fields for Technical Specifications objects from weidmueller.com. All fields typed and schema-versioned.

order_numberrated_voltage_vrated_current_awire_cross_section_mm2awg_minawg_maxmounting_typeoperating_temperature_minoperating_temperature_maxip_ratingdimensions_width_mmdimensions_height_mmdimensions_depth_mmweight_g
technical_specifications
● 200 OK
"order_number": "1020000000",
"rated_voltage_v": 800.0,
"rated_current_a": 24.0,
"wire_cross_section_mm2": 2.5,
"mounting_type": "TS 35",
"operating_temperature_min": -60,
"operating_temperature_max": 130,
"weight_g": 7.15
# order_numberrated_voltage_vrated_current_awire_cross_section_mm2awg_minawg_max
1
2
3

Complete list of extractable fields for Compliance & Approvals objects from weidmueller.com. All fields typed and schema-versioned.

order_numberrohs_statusrohs_datereach_statussvhc_substancece_markul_approvalcsa_approvalatex_certifiediec_standard
compliance_& approvals
● 200 OK
"order_number": "1020000000",
"rohs_status": "Conform",
"rohs_date": "2026-01-01",
"reach_status": "Conform",
"ce_mark": true,
"ul_approval": true,
"atex_certified": true
# order_numberrohs_statusrohs_datereach_statussvhc_substancece_mark
1
2
3

Complete list of extractable fields for Downloads & Assets objects from weidmueller.com. All fields typed and schema-versioned.

order_numberdatasheet_pdf_urlcad_step_urlcad_iges_urlcad_dxf_urleplan_macro_urluser_manual_urlcertificate_urlsoftware_urlimage_highres_url
downloads_& assets
● 200 OK
"order_number": "1020000000",
"datasheet_pdf_url": "https://catalog.weidmueller.com/catalog/Start.do?ObjectID=1020000000&page=ProductPdf",
"cad_step_url": "https://catalog.weidmueller.com/catalog/STEP/1020000000.stp",
"eplan_macro_url": "https://catalog.weidmueller.com/catalog/EPLAN/1020000000.edz",
"certificate_url": "https://catalog.weidmueller.com/catalog/Cert/CE_1020000000.pdf",
"image_highres_url": "https://catalog.weidmueller.com/catalog/Images/1020000000_high.jpg"
# order_numberdatasheet_pdf_urlcad_step_urlcad_iges_urlcad_dxf_urleplan_macro_url
1
2
3

Complete list of extractable fields for Related Products objects from weidmueller.com. All fields typed and schema-versioned.

order_numberaccessory_order_numbersalternative_order_numbersreplacement_order_numbersmatching_toolsmatching_markersend_plate_order_numberscross_connection_order_numbers
related_products
● 200 OK
"order_number": "1020000000",
"accessory_order_numbers": "['1050000000', '1060000000']",
"matching_tools": "['9008330000']",
"matching_markers": "['1609801044']",
"end_plate_order_numbers": "['1050000000']",
"cross_connection_order_numbers": "['1052560000', '1052660000']"
# order_numberaccessory_order_numbersalternative_order_numbersreplacement_order_numbersmatching_toolsmatching_markers
1
2
3

Capabilities

Industrial component extraction at scale

Weidmüller's catalogue is deeply nested and highly technical. Our pipelines parse complex engineering specifications, normalise MRO taxonomies, and extract asset URLs for direct integration into your PIM or ERP.

Full Component Extraction

Order numbers, type designations, descriptions, and packaging quantities scraped across all product groups.

Technical Spec Normalisation

Extract and standardise voltage, current, dimensions, and mounting types from dynamic HTML tables into flat schemas.

ECLASS & UNSPSC Mapping

Capture industry-standard classification codes to ensure immediate compatibility with procurement systems.

CAD & Asset Indexing

Extract direct URLs for STEP, IGES, DXF files, EPLAN macros, and high-resolution product images.

Compliance Data Mining

Capture RoHS, REACH, CE, UL, and ATEX certification statuses and download links for compliance auditing.

Accessory & Cross-Reference Mapping

Map parent components to compatible end plates, cross-connections, tools, and markers.

Regional Catalogue Support

Scrape locale-specific catalogues to capture regional availability and compliance variations.

Delta Updates

Run scheduled diffs to identify new product introductions, obsolete parts, and updated datasheets.

Datasheet Generation Links

Capture the dynamic URLs used to generate PDF datasheets on the fly from the Weidmüller product catalogue.

// engagement pipeline

From product group to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target product categories, ECLASS codes, or specific order number lists. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, handle regional routing, and write parsers for Weidmüller's technical tables.

Validation & QA
d 4–6

Schema validation, null-rate checks on critical specs (voltage, current), and asset URL verification before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Handling MRO catalogue complexity

Extracting data from industrial manufacturers requires handling deep taxonomies and inconsistent technical tables. Here is how we ensure data quality.

pipeline-monitor · weidmueller.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Deep category trees
Recursive taxonomy traversal

Weidmüller organises products into deeply nested hierarchies (e.g., Products > Connectivity > Terminal Blocks > W-Series). Our crawlers recursively map this taxonomy, ensuring no sub-category or variant is missed during extraction.

Dynamic spec tables
Heuristic attribute mapping

Technical specifications vary wildly between a terminal block and an industrial Ethernet switch. We use heuristic parsers to map varying HTML table rows into a normalised, columnar schema with consistent units of measurement.

Asset gating
Session management for downloads

Some CAD files and EPLAN macros require session tokens or specific request headers to generate the download link. Our Playwright orchestrators maintain valid sessions to extract direct, functional asset URLs.

Regional variations
Locale-specific crawling

Product availability and compliance data often differ between the EU, US, and Asian markets. We route requests through region-specific residential proxies to capture the exact catalogue presented to your target market.

Schema normalisation
Consistent data types

We convert string-based specifications (e.g., '2.5 mm²') into typed numeric fields (2.5) while preserving the unit in the schema definition, ensuring the data is immediately ready for mathematical operations in your ERP.

Applications

Who uses Weidmüller data — and how

Teams across industries use weidmueller.com data to build competitive products and smarter operations.

01
PIM Enrichment

Distributors populate their Product Information Management (PIM) systems with accurate specs, images, and ECLASS codes directly from the manufacturer.

02
MRO Procurement

Procurement teams build internal catalogues with accurate order numbers and replacement part data to streamline purchasing workflows.

03
Engineering CAD Libraries

Engineering teams ingest STEP and IGES model links to bulk-update their internal component libraries for CAD software.

04
Compliance Auditing

Manufacturers audit their Bill of Materials (BOM) against Weidmüller's latest RoHS, REACH, and SVHC declarations.

05
Competitor Benchmarking

R&D teams analyse Weidmüller's product portfolio specifications to benchmark their own connectivity and automation products.

06
Cross-Reference Engines

Distributors build cross-reference tools allowing customers to find Weidmüller equivalents for competitor terminal blocks and relays.

Why DataFlirt

"Weidmüller's online catalogue contains critical engineering data, but extracting millions of technical attributes requires a pipeline built for complex MRO taxonomies."

Industrial component scraping is an exercise in schema normalisation. Weidmüller's product pages feature deeply nested specification tables, dynamic CAD generation endpoints, and region-specific compliance documents. DataFlirt handles the extraction and normalisation, delivering clean tabular data ready for your ERP or PIM system.

Technical Spec

Weidmüller scraper — technical capabilities

Everything supported by our weidmueller.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions — required for dynamic asset generation and spec tables
Supported
CAPTCHA bypass
Automated 2Captcha + CapSolver integration
Supported
Residential proxy rotation
ISP-grade residential IPs to prevent rate limiting during deep catalogue crawls
Supported
Spec table normalisation
Converts variable HTML tables into fixed columnar schemas
Supported
CAD URL extraction
Extracts direct links for STEP, IGES, and DXF files
Supported
ECLASS & UNSPSC mapping
Captures standard classification codes for procurement integration
Supported
Multi-region catalogues
Supports DE, US, UK, and other regional Weidmüller domains
Supported
Direct CAD file download
We extract the URLs, but downloading terabytes of STEP files requires a separate S3 transfer agreement
Partial
Partner/Distributor pricing
Requires authenticated distributor portal credentials
Partial
Infrastructure

Infrastructure powering the Weidmüller pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows. Combined via scrapy-playwright middleware.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across DE/US/UK regions. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.

Cloud-Native Orchestration

Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested — schema versioned per run
CSV
Flat file with typed columns — Excel/Sheets compatible
XLS
Formatted Excel exports for procurement teams
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery — compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoints to query extracted Weidmüller data
BigQuery
Streamed directly into your dataset with schema auto-detect
Snowflake
Stage + COPY INTO workflow — incremental or full-replace
Postgres
Upsert into your existing schema with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About weidmueller.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Weidmüller's catalogue legal?

Scraping publicly available technical specifications, part numbers, and classification data is generally permissible. DataFlirt extracts only public, non-authenticated catalogue data. We do not circumvent authentication walls for distributor pricing. Clients should review Weidmüller's ToS and consult legal counsel for specific commercial use cases.

How do you handle the complex technical specification tables?

Weidmüller products span diverse categories, meaning spec tables vary constantly. We build heuristic parsers that map specific row headers (e.g., 'Rated voltage') to standardised schema columns, ensuring data consistency across the entire catalogue.

Can you extract CAD models and datasheets?

We extract the direct URLs for STEP files, IGES files, EPLAN macros, and PDF datasheets. By default, we deliver the URLs rather than the files themselves to prevent massive storage bloat, allowing your systems to download the assets as needed.

How often is the data updated?

Pipelines can be configured for daily, weekly, or monthly runs. For MRO catalogues, a weekly or monthly delta run is typical to capture new product introductions and updated compliance documents.

Do you capture ECLASS and UNSPSC codes?

Yes. Where Weidmüller provides standard classification codes (ECLASS, UNSPSC, or ETIM), we extract them to ensure the data integrates smoothly into standard procurement and ERP systems.

Can I get pricing data?

We extract publicly visible list prices if available on the regional site. However, we do not support logging into distributor portals to extract negotiated, account-specific pricing.

$ dataflirt scope --new-project --source=weidmueller.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump for your PIM or a continuous sync of compliance documents and CAD links — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in industrial and mro

Services

Data Extraction for Every Industry

View All Services →