We extract terminal blocks, relays, automation components, technical specs, and CAD metadata from Weidmüller. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Component Base Data objects from weidmueller.com. All fields typed and schema-versioned.
"order_number": "1020000000", "type_designation": "WDU 2.5", "product_group": "Terminal Blocks", "gtin_ean": "4008190115354", "qty_per_pack": 100, "eclass_code": "27-14-11-20", "unspsc_code": "39-12-14-10"
| # | order_number | type_designation | product_group | short_description | long_description | gtin_ean |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Technical Specifications objects from weidmueller.com. All fields typed and schema-versioned.
"order_number": "1020000000", "rated_voltage_v": 800.0, "rated_current_a": 24.0, "wire_cross_section_mm2": 2.5, "mounting_type": "TS 35", "operating_temperature_min": -60, "operating_temperature_max": 130, "weight_g": 7.15
| # | order_number | rated_voltage_v | rated_current_a | wire_cross_section_mm2 | awg_min | awg_max |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Compliance & Approvals objects from weidmueller.com. All fields typed and schema-versioned.
"order_number": "1020000000", "rohs_status": "Conform", "rohs_date": "2026-01-01", "reach_status": "Conform", "ce_mark": true, "ul_approval": true, "atex_certified": true
| # | order_number | rohs_status | rohs_date | reach_status | svhc_substance | ce_mark |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Downloads & Assets objects from weidmueller.com. All fields typed and schema-versioned.
"order_number": "1020000000", "datasheet_pdf_url": "https://catalog.weidmueller.com/catalog/Start.do?ObjectID=1020000000&page=ProductPdf", "cad_step_url": "https://catalog.weidmueller.com/catalog/STEP/1020000000.stp", "eplan_macro_url": "https://catalog.weidmueller.com/catalog/EPLAN/1020000000.edz", "certificate_url": "https://catalog.weidmueller.com/catalog/Cert/CE_1020000000.pdf", "image_highres_url": "https://catalog.weidmueller.com/catalog/Images/1020000000_high.jpg"
| # | order_number | datasheet_pdf_url | cad_step_url | cad_iges_url | cad_dxf_url | eplan_macro_url |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Related Products objects from weidmueller.com. All fields typed and schema-versioned.
"order_number": "1020000000", "accessory_order_numbers": "['1050000000', '1060000000']", "matching_tools": "['9008330000']", "matching_markers": "['1609801044']", "end_plate_order_numbers": "['1050000000']", "cross_connection_order_numbers": "['1052560000', '1052660000']"
| # | order_number | accessory_order_numbers | alternative_order_numbers | replacement_order_numbers | matching_tools | matching_markers |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Weidmüller's catalogue is deeply nested and highly technical. Our pipelines parse complex engineering specifications, normalise MRO taxonomies, and extract asset URLs for direct integration into your PIM or ERP.
Order numbers, type designations, descriptions, and packaging quantities scraped across all product groups.
Extract and standardise voltage, current, dimensions, and mounting types from dynamic HTML tables into flat schemas.
Capture industry-standard classification codes to ensure immediate compatibility with procurement systems.
Extract direct URLs for STEP, IGES, DXF files, EPLAN macros, and high-resolution product images.
Capture RoHS, REACH, CE, UL, and ATEX certification statuses and download links for compliance auditing.
Map parent components to compatible end plates, cross-connections, tools, and markers.
Scrape locale-specific catalogues to capture regional availability and compliance variations.
Run scheduled diffs to identify new product introductions, obsolete parts, and updated datasheets.
Capture the dynamic URLs used to generate PDF datasheets on the fly from the Weidmüller product catalogue.
Brief in. Clean data out.
Provide target product categories, ECLASS codes, or specific order number lists. We design the extraction schema together.
We configure Scrapy / Playwright crawlers, handle regional routing, and write parsers for Weidmüller's technical tables.
Schema validation, null-rate checks on critical specs (voltage, current), and asset URL verification before full launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Extracting data from industrial manufacturers requires handling deep taxonomies and inconsistent technical tables. Here is how we ensure data quality.
Weidmüller organises products into deeply nested hierarchies (e.g., Products > Connectivity > Terminal Blocks > W-Series). Our crawlers recursively map this taxonomy, ensuring no sub-category or variant is missed during extraction.
Technical specifications vary wildly between a terminal block and an industrial Ethernet switch. We use heuristic parsers to map varying HTML table rows into a normalised, columnar schema with consistent units of measurement.
Some CAD files and EPLAN macros require session tokens or specific request headers to generate the download link. Our Playwright orchestrators maintain valid sessions to extract direct, functional asset URLs.
Product availability and compliance data often differ between the EU, US, and Asian markets. We route requests through region-specific residential proxies to capture the exact catalogue presented to your target market.
We convert string-based specifications (e.g., '2.5 mm²') into typed numeric fields (2.5) while preserving the unit in the schema definition, ensuring the data is immediately ready for mathematical operations in your ERP.
Distributors populate their Product Information Management (PIM) systems with accurate specs, images, and ECLASS codes directly from the manufacturer.
Procurement teams build internal catalogues with accurate order numbers and replacement part data to streamline purchasing workflows.
Engineering teams ingest STEP and IGES model links to bulk-update their internal component libraries for CAD software.
Manufacturers audit their Bill of Materials (BOM) against Weidmüller's latest RoHS, REACH, and SVHC declarations.
R&D teams analyse Weidmüller's product portfolio specifications to benchmark their own connectivity and automation products.
Distributors build cross-reference tools allowing customers to find Weidmüller equivalents for competitor terminal blocks and relays.
"Weidmüller's online catalogue contains critical engineering data, but extracting millions of technical attributes requires a pipeline built for complex MRO taxonomies."
Industrial component scraping is an exercise in schema normalisation. Weidmüller's product pages feature deeply nested specification tables, dynamic CAD generation endpoints, and region-specific compliance documents. DataFlirt handles the extraction and normalisation, delivering clean tabular data ready for your ERP or PIM system.
Everything supported by our weidmueller.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows. Combined via scrapy-playwright middleware.
We maintain pools of residential ISP proxies across DE/US/UK regions. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.
Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About weidmueller.com scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available technical specifications, part numbers, and classification data is generally permissible. DataFlirt extracts only public, non-authenticated catalogue data. We do not circumvent authentication walls for distributor pricing. Clients should review Weidmüller's ToS and consult legal counsel for specific commercial use cases.
Weidmüller products span diverse categories, meaning spec tables vary constantly. We build heuristic parsers that map specific row headers (e.g., 'Rated voltage') to standardised schema columns, ensuring data consistency across the entire catalogue.
We extract the direct URLs for STEP files, IGES files, EPLAN macros, and PDF datasheets. By default, we deliver the URLs rather than the files themselves to prevent massive storage bloat, allowing your systems to download the assets as needed.
Pipelines can be configured for daily, weekly, or monthly runs. For MRO catalogues, a weekly or monthly delta run is typical to capture new product introductions and updated compliance documents.
Yes. Where Weidmüller provides standard classification codes (ECLASS, UNSPSC, or ETIM), we extract them to ensure the data integrates smoothly into standard procurement and ERP systems.
We extract publicly visible list prices if available on the regional site. However, we do not support logging into distributor portals to extract negotiated, account-specific pricing.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump for your PIM or a continuous sync of compliance documents and CAD links — we scope, build, and operate the pipeline. Tell us what you need.