We extract yarn weights, colourways, gauge metrics, pricing, and pattern requirements from Yarnspirations. Delivered as clean JSON, CSV, or Parquet to S3 or BigQuery on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Yarn Products objects from yarnspirations.com. All fields typed and schema-versioned.
"sku": "294009", "name": "Red Heart Super Saver Yarn", "brand": "Red Heart", "fibre_content": "100% Acrylic", "weight_category": "4 - Medium", "gauge_knit": "17 sts and 23 rows with a 5 mm knitting needle", "gauge_crochet": "12 sc and 15 rows with a 5.5 mm crochet hook", "base_price": 4.99
| # | sku | name | brand | fibre_content | weight_category | gauge_knit |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Colourways & Variants objects from yarnspirations.com. All fields typed and schema-versioned.
"parent_sku": "294009", "variant_sku": "294009-0311", "colour_name": "White", "colour_code": "0311", "dye_lot_required": false, "price": 4.99, "inventory_level": 450, "upc": "073650777458"
| # | parent_sku | variant_sku | colour_name | colour_code | dye_lot_required | image_url |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Free Patterns objects from yarnspirations.com. All fields typed and schema-versioned.
"pattern_id": "BRK0126-030278M", "title": "Bernat Blanket Crochet Bear", "craft_type": "Crochet", "skill_level": "Easy", "yarn_required": "Bernat Blanket", "hook_size": "8 mm (U.S. L/11)", "download_url": "https://www.yarnspirations.com/on/demandware.static/-/Sites-master-catalog-spinrite/default/pdf/BRK0126-030278M.pdf", "designer": "Yarnspirations Design Studio"
| # | pattern_id | title | craft_type | skill_level | yarn_required | hook_size |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Pricing & Promotions objects from yarnspirations.com. All fields typed and schema-versioned.
"sku": "294009-0311", "base_price": 4.99, "sale_price": 3.99, "discount_pct": 20, "bulk_discount_eligible": true, "clearance_flag": false, "currency": "USD", "timestamp": "2026-05-12T09:14:00Z"
| # | sku | base_price | sale_price | discount_pct | bulk_discount_eligible | clearance_flag |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Reviews & Ratings objects from yarnspirations.com. All fields typed and schema-versioned.
"review_id": "REV-992834", "sku": "294009", "rating": 5, "title": "Perfect for afghans", "body": "I have used this yarn for years. It washes well and holds its shape.", "author": "CrochetQueen99", "date": "2026-02-14", "verified_buyer": true
| # | review_id | sku | rating | title | body | author |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Yarnspirations structures data heavily around variants and patterns. Our pipeline parses technical textile specifications, colourway matrices, and PDF metadata directly into tabular formats.
Capture weight categories, fibre blends, care instructions, and exact gauge metrics for knitting and crochet across all brands.
Extract hundreds of colour variants per parent SKU, mapping colour codes, names, and dye lot requirements to exact stock levels.
Scrape craft type, skill level, required hook/needle sizes, and direct PDF download links for the entire free pattern database.
Monitor base prices, sale discounts, clearance flags, and bulk purchase tiers across the entire inventory.
Extract customer reviews, star ratings, and verified buyer tags to evaluate product reception and quality issues.
Filter and extract data specific to Bernat, Caron, Red Heart, Lily Sugar'n Cream, or Patons with brand-level attribution.
Track in-stock, out-of-stock, and low-stock indicators per colourway to model inventory depth.
Extract 'Yarns for this pattern' and 'Patterns for this yarn' relational data to map the product ecosystem.
Receive only records that have changed since the last run, reducing processing overhead for daily price and stock updates.
Brief in. Clean data out.
Provide category URLs, brand filters, or search terms. We design the extraction schema together.
We configure Scrapy crawlers, proxy rotation, and session management tailored to Yarnspirations' architecture.
Schema validation, null-rate checks, and variant mapping verification before full launch.
JSON, CSV, or Parquet pushed to your S3 bucket or BigQuery dataset on agreed cadence.
Extracting variant-heavy textile data requires handling dynamic DOM updates and strict request limits. Here is how we maintain data integrity.
Yarnspirations loads colour variants, stock status, and dynamic pricing via JavaScript. We use Playwright to execute these scripts and capture the rendered state, ensuring no variant is missed.
To prevent IP bans from aggressive crawling, we distribute requests across US-based residential proxies, managing concurrency and request delays to mimic legitimate traffic patterns.
Retailers update DOM structures for seasonal promotions. We build resilient extraction logic using multiple fallback selectors (CSS, XPath, JSON-LD) to maintain pipeline stability.
A single yarn line can have over 100 colourways. We flatten this hierarchy into structured, queryable rows, linking every variant back to its parent SKU and technical specifications.
Every run undergoes automated checks for missing prices, malformed gauge strings, or dropped variants. If anomaly thresholds are breached, the pipeline pauses and alerts our engineering team.
Craft retailers monitor Yarnspirations' pricing, discount frequencies, and clearance events to adjust their own promotional strategies.
Supply chain analysts track stock status across major colourways to identify supply shortages and forecast seasonal demand.
Market researchers analyse new pattern releases and popular yarn weights to identify emerging trends in the knitting and crochet communities.
Machine learning teams ingest pattern metadata and yarn requirements to train recommendation engines for craft enthusiasts.
Textile manufacturers aggregate fibre content and gauge data to benchmark their own product lines against industry standards.
Independent designers and small businesses track bulk discount eligibility and dye lot availability for large-scale project planning.
"Yarnspirations holds the definitive digital catalogue for craft textiles and pattern requirements, but extracting unified gauge and fibre metrics requires purpose-built pipelines."
Most teams underestimate the complexity of scraping textile variants. Handling hundreds of colourways per yarn line, extracting PDF pattern metadata, and normalising gauge metrics demands precise selector maintenance and JavaScript execution. DataFlirt absorbs this infrastructure overhead so your team can focus on data modelling.
Everything supported by our yarnspirations.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy manages orchestration and deduplication, while Playwright handles JavaScript execution for dynamic product variants and stock indicators.
We route traffic through high-quality residential proxies, managing concurrency and request headers to avoid triggering security blocks.
Pipelines run on AWS infrastructure managed by Apache Airflow, ensuring reliable scheduling, retries, and data delivery.
Data delivered to where your team already works — no new tooling required.
About yarnspirations.com scraping, legality, and pipeline operations.
Ask us directly →Yes. We can configure the crawler to target specific brand collections such as Bernat, Red Heart, or Caron, ignoring irrelevant product lines to reduce your data processing overhead.
Our pipeline captures the exact stock status for every variant. If a specific colourway is out of stock, it is recorded with its status flag rather than being omitted from the dataset.
We extract the direct URL to the PDF file. If required, we can configure an additional pipeline stage to download and store the actual PDF files in your S3 bucket.
We support daily or weekly runs for full catalogue refreshes. For targeted tracking of specific SKUs, we can configure hourly polling for stock and price changes.
Yes. We parse the technical specifications block on product and pattern pages, extracting specific knitting gauge, crochet gauge, and recommended tool sizes into structured fields.
We use a parent-child schema. The parent SKU holds the base specifications (fibre, weight, gauge), while child variants hold colour-specific data (colour code, image, stock status). This can be delivered as nested JSON or flattened CSV rows.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a full catalogue extraction or daily price monitoring across specific yarn lines, we manage the infrastructure. Tell us your requirements.