We extract manufacturer profiles, textile machinery specifications, and industry news from Texdata. Delivered as clean JSON, CSV, or Parquet to S3 or BigQuery on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Company Profiles objects from texdata.com. All fields typed and schema-versioned.
"company_id": "TX-8492", "company_name": "Groz-Beckert KG", "country": "Germany", "website": "groz-beckert.com", "categories": "['Knitting Machinery', 'Needles']", "established_year": 1852, "employee_count": 9000
| # | company_id | company_name | country | city | website | |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Machinery Specs objects from texdata.com. All fields typed and schema-versioned.
"machine_id": "M-4921", "manufacturer": "Karl Mayer", "model_name": "HKS 3-M ON", "category": "Warp Knitting", "application": "Sportswear", "production_speed": "2800 rpm", "release_year": 2021
| # | machine_id | manufacturer | model_name | category | application | production_speed |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Industry News objects from texdata.com. All fields typed and schema-versioned.
"article_id": "N-99120", "headline": "ITMA 2027 to be held in Hannover", "publish_date": "2026-03-14", "category": "Trade Fairs", "tags": "['ITMA', 'Exhibition', 'Hannover']", "source_company": "CEMATEX"
| # | article_id | headline | author | publish_date | category | full_text |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Trade Fair Events objects from texdata.com. All fields typed and schema-versioned.
"event_id": "E-104", "event_name": "Techtextil 2026", "start_date": "2026-04-21", "end_date": "2026-04-24", "location": "Frankfurt", "venue": "Messe Frankfurt", "focus_areas": "['Technical Textiles', 'Nonwovens']"
| # | event_id | event_name | start_date | end_date | location | venue |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Product Categories objects from texdata.com. All fields typed and schema-versioned.
"category_id": "C-042", "category_name": "Spinning Machinery", "parent_category": "Textile Machinery", "supplier_count": 342, "machinery_count": 1205, "top_countries": "['Germany', 'Italy', 'China']"
| # | category_id | category_name | parent_category | description | supplier_count | machinery_count |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our Texdata scraper handles the entire portal: supplier directories, machinery technical specifications, and the industry news corpus - with pagination handling and document parsing built in.
Extract company names, contact details, product categories, and certifications across all global regions.
Parse technical specifications, production speeds, and power consumption metrics for textile machinery.
Download full article text, publication dates, author metadata, and tags from the daily news section.
Scrape upcoming event dates, locations, venue details, and complete exhibitor lists.
Extract hierarchical product and machinery categories to map the complete textile supply chain.
Isolate and validate email addresses, phone numbers, and corporate websites from supplier profiles.
Handle mixed German and English content natively, standardising field outputs across languages.
Parse tabular data from PDF brochures and specification sheets linked on manufacturer profiles.
Configure continuous pipelines at daily or weekly cadences to capture new supplier registrations and news articles.
Brief in. Clean data out.
Provide target categories, machinery types, or news sections. We design the extraction schema together.
We configure Scrapy crawlers, proxy rotation, and session management for texdata.com.
Schema validation, null-rate checks, and data type standardisation before full launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Extracting B2B directory data requires careful state management and normalisation. Here is how we ensure reliable delivery.
Texdata implements basic rate limiting and IP blocking. We use proxy pools with rotating IPs to distribute request volume, ensuring continuous extraction without triggering security thresholds.
Textile specifications on Texdata often mix German and English terminology. Our pipeline maps language-specific labels to a unified schema, ensuring consistent field names and units across the dataset.
Crucial machinery specifications are frequently locked inside PDF brochures. We integrate document parsing libraries to extract tabular data and text from linked PDFs directly into structured JSON.
Deep supplier directories require precise state management. Our crawlers traverse nested pagination without missing records or entering infinite loops, ensuring 100% coverage of target categories.
For the supplier database, we maintain a hash index of last-seen values per company. Subsequent runs only push diffs, reducing downstream processing load and storage costs.
Identify alternative textile manufacturers and raw material suppliers across global markets.
Track machinery upgrades and production capabilities of competing textile mills.
Analyse the news corpus and event focus areas to predict demand for technical textiles or sustainable fabrics.
Extract verified contact details for textile machinery manufacturers and fabric suppliers.
Compare technical specifications and power consumption metrics across different spinning or weaving machines.
Monitor trade fair exhibitor lists to identify networking targets and industry concentration.
"Texdata holds the technical specifications and supplier networks that run the global textile industry - but extracting it requires a systematic crawling infrastructure."
Textile supply chains rely on highly specific machinery and vetted suppliers. Manually compiling this data from Texdata is impossible at scale. DataFlirt automates the extraction of company directories, technical brochures, and news archives, delivering clean, structured datasets so your procurement and research teams can operate efficiently.
Everything supported by our texdata.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and retry logic. Playwright handles dynamic directory loading and search interfaces.
Integrated PDF parsing pipelines extract tabular specifications from manufacturer brochures linked within Texdata.
Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting.
Data delivered to where your team already works — no new tooling required.
About texdata.com scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available directory and news information is generally permissible. DataFlirt targets only public, non-authenticated supplier profiles, machinery specs, and news data. Clients should review Texdata terms of service and consult legal counsel for specific use cases.
We use automated document parsing libraries to extract tabular data from linked machinery brochures, converting unstructured PDF tables into clean JSON fields.
Yes, we extract publicly listed contact information, including emails, phone numbers, and corporate websites directly from the supplier profile pages.
We extract the raw text as published. We can implement translation APIs in the post-processing pipeline upon request to normalise output into a single language.
We can configure pipelines to poll the news section daily or hourly, extracting only newly published articles using hash-based diffing.
No. We only extract publicly accessible data. Premium reports requiring a paid Texdata subscription are not supported.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off supplier directory dump or continuous monitoring of textile machinery specs - we scope, build, and operate the pipeline. Tell us what you need.