We extract industrial product specs, MRO part numbers, CAD metadata, and lifecycle statuses from ABB. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Product Specifications objects from abb.com. All fields typed and schema-versioned.
"part_number": "1SDA066776R1", "product_id": "XT2N 160", "title": "XT2N 160 TMD 16-300 3p F F", "category": "Low Voltage Products and Systems", "sub_category": "Circuit Breakers", "lifecycle_status": "Active", "ean": "8015644006869", "net_weight": "1.1 kg"
| # | part_number | product_id | title | category | sub_category | lifecycle_status |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Technical Documentation objects from abb.com. All fields typed and schema-versioned.
"part_number": "1SDA066776R1", "doc_title": "Tmax XT Technical Catalogue", "doc_type": "Catalogue", "language": "English", "revision": "C", "publication_date": "2025-11-14", "file_size": "14.2 MB", "download_url": "https://search.abb.com/library/Download..."
| # | part_number | doc_title | doc_type | language | revision | publication_date |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Ordering Information objects from abb.com. All fields typed and schema-versioned.
"part_number": "1SDA066776R1", "order_code": "1SDA066776R1", "minimum_order_qty": 1, "selling_unit": "piece", "customs_tariff_number": "85362090", "country_of_origin": "Italy (IT)", "ean": "8015644006869"
| # | part_number | order_code | minimum_order_qty | selling_unit | customs_tariff_number | country_of_origin |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Environmental Data objects from abb.com. All fields typed and schema-versioned.
"part_number": "1SDA066776R1", "rohs_status": "Following EU Directive 2011/65/EU", "rohs_date": "2025-01-01", "weee_category": "5. Small Equipment (No External Dimension More Than 50 cm)", "reach_declaration": "Available", "scip_id": "8a7b6c5d-4e3f-2g1h-0i9j", "carbon_footprint": "12.4 kg CO2e"
| # | part_number | rohs_status | rohs_date | weee_category | reach_declaration | scip_id |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Distributor Locations objects from abb.com. All fields typed and schema-versioned.
"partner_name": "Industrial Automation Supply Co.", "partner_type": "Authorized Distributor", "address": "142 Manufacturing Blvd", "city": "Chicago", "country": "United States", "phone": "+1-555-0198", "website": "https://iasupply.example.com", "latitude": 41.8781, "longitude": -87.6298
| # | partner_name | partner_type | address | city | country | phone |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our ABB scraper navigates complex product configurators, nested category trees, and regional catalogues to deliver clean MRO datasets.
Extract highly variable technical attributes across robotics, motors, drives, and low-voltage categories into normalised schemas.
Monitor transition phases from Active to Classic, Limited, or Obsolete to inform procurement and engineering decisions.
Capture direct URLs and metadata for CAD files, installation manuals, declarations of conformity, and environmental profiles.
Map long ABB product IDs to standard EANs, customs tariff numbers, and minimum order quantities for supply chain systems.
Extract product availability and compliance data specific to local markets via geographically targeted proxies.
Scrape global channel partner networks to map authorised distributors, system integrators, and service providers.
Automate interactions with ABB's JavaScript-heavy product selection tools to extract valid component combinations.
Extract compliance declarations, WEEE categories, and SCIP database identifiers for sustainability reporting.
Run continuous pipelines that detect changes in lifecycle statuses or document revisions, pushing only updated records.
Brief in. Clean data out.
Provide target categories, part number lists, or regional requirements. We design the extraction schema.
We configure Playwright crawlers to navigate ABB configurators, handle pagination, and normalise spec tables.
Schema validation, null-rate checks, and cross-reference verification before full production launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Extracting MRO data requires handling deep JS configurators and massive spec tables. Here is how we maintain pipeline stability.
Many ABB product variants are only accessible via dynamic configurators. We deploy Playwright to execute JavaScript, select dropdowns, and render the final specification tables accurately.
A motor has entirely different attributes than a circuit breaker. We map thousands of unique attribute keys into a clean, predictable JSON schema, handling unit conversions and formatting variations.
ABB's technical library relies on complex search APIs and temporary tokens. Our pipeline interacts directly with these endpoints to extract permanent document URLs and accurate metadata.
Product availability and compliance data vary by region. We route requests through specific proxy nodes to capture the exact catalogue visible to users in Germany, the US, or India.
For MRO monitoring, we maintain a state file of all part numbers. The pipeline only flags records when a product shifts from 'Active' to 'Classic' or 'Obsolete', reducing data noise.
Procurement teams ingest part numbers and lifecycle statuses into ERP systems to identify obsolete components before they fail.
Engineering firms extract dimensional data, net weights, and CAD links to populate digital twin environments and BIM software.
Industrial manufacturers map ABB specifications against their own catalogues to build automated cross-reference tools.
Distributors monitor ordering codes, EANs, and customs tariff numbers to automate inventory ingestion and customs documentation.
Compliance teams extract RoHS, REACH, and carbon footprint data to meet environmental reporting requirements.
Market analysts scrape the global partner portal to map ABB's distribution footprint and identify regional service coverage gaps.
"ABB's digital catalogue contains millions of engineering specifications, but accessing that data programmatically requires navigating deep configurators and fragmented document portals."
Industrial data extraction requires more than simple HTTP requests. We deploy full browser automation to interact with ABB's product configurators, normalise highly variable specification tables, and index technical documentation accurately. DataFlirt handles the extraction infrastructure so your engineers can focus on integrating the data.
Everything supported by our abb.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and normalisation. Playwright executes JavaScript to navigate product configurators and render hidden specification tables.
We maintain pools of datacenter and residential proxies. Rotation happens per-request to ensure stable access to regional catalogues without rate limits.
Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About abb.com scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available product specifications and documentation from ABB is generally permissible under applicable law. DataFlirt targets only public, non-authenticated catalogue data. We do not extract personal data, circumvent authentication walls like MyABB, or violate GDPR. Clients should consult legal counsel for specific use cases.
We use Playwright to simulate browser sessions, executing JavaScript to interact with dropdowns, sliders, and selection toggles. This allows us to extract the final rendered specification tables for specific part configurations.
Yes. ABB's catalogue varies by country. We route our crawlers through geotargeted proxy nodes to extract exact product availability, compliance data, and ordering codes for your target regions.
We extract the metadata (title, language, revision, publication date) and the direct download URLs for all technical documentation. We do not host the files, but we provide the structured links for your systems to download.
For full catalogue extractions, we typically run weekly or monthly pipelines. For targeted lists of critical MRO parts, we can run daily checks to detect lifecycle status changes.
Absolutely. We provide a sample run of up to 1,000 part numbers or a specific product category as part of the pre-engagement scoping process to validate schema fit and data quality.
Our smallest packages start at a defined list of 10,000 part numbers or specific category trees. For full global catalogue extraction, we price based on volume, configurator complexity, and delivery frequency.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off catalogue dump or continuous lifecycle monitoring across 500K part numbers — we scope, build, and operate the pipeline. Tell us what you need.