SYSTEM all green source toofaced.com queue 3,104 pages p99 latency 214ms dataflirt.com · scraper/toofaced-com
RUN · 14 active pipelines · toofaced.com live

Too Faced data,
ready for analysis.

We extract product lines, shade variants, ingredient lists, pricing, and reviews from toofaced.com. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products extracted
1,492 /run
Shade variants
8,941 /run
Review records
142K /total
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from toofaced.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Products objects from toofaced.com. All fields typed and schema-versioned.

product_idnamecategorysub_categorybase_pricecurrencydescriptionhow_to_useingredientsis_veganis_cruelty_free
products
● 200 OK
"product_id": "TF-9482",
"name": "Better Than Sex Mascara",
"category": "Makeup",
"sub_category": "Eyes",
"base_price": 29.0,
"currency": "USD",
"is_vegan": true,
"is_cruelty_free": true
# product_idnamecategorysub_categorybase_pricecurrency
1
2
3

Complete list of extractable fields for Shades objects from toofaced.com. All fields typed and schema-versioned.

product_idshade_nameshade_descriptionhex_codeswatch_image_urlin_stockskuprice_override
shades
● 200 OK
"product_id": "TF-1029",
"shade_name": "Cloud",
"shade_description": "Fairest with rosy undertones",
"hex_code": "#FAD6C9",
"in_stock": true,
"sku": "TF-1029-CLD",
"price_override": "None"
# product_idshade_nameshade_descriptionhex_codeswatch_image_urlin_stock
1
2
3

Complete list of extractable fields for Pricing objects from toofaced.com. All fields typed and schema-versioned.

product_idskubase_pricesale_pricediscount_pctcurrencyis_limited_editionpromotion_textscraped_at
pricing
● 200 OK
"product_id": "TF-9482",
"sku": "TF-9482-STD",
"base_price": 29.0,
"sale_price": 23.2,
"discount_pct": 20,
"currency": "USD",
"is_limited_edition": false,
"promotion_text": "20% Off Sitewide"
# product_idskubase_pricesale_pricediscount_pctcurrency
1
2
3

Complete list of extractable fields for Reviews objects from toofaced.com. All fields typed and schema-versioned.

review_idproduct_idreviewer_namestar_ratingreview_titlereview_textskin_typeage_rangerecommendationreview_date
reviews
● 200 OK
"review_id": "REV-847291",
"product_id": "TF-9482",
"star_rating": 5,
"review_title": "Holy Grail Mascara",
"skin_type": "Combination",
"age_range": "25-34",
"recommendation": true,
"review_date": "2023-11-14"
# review_idproduct_idreviewer_namestar_ratingreview_titlereview_text
1
2
3

Complete list of extractable fields for Categories objects from toofaced.com. All fields typed and schema-versioned.

category_idcategory_nameparent_categoryurlproduct_countbanner_image_urlmeta_titlemeta_description
categories
● 200 OK
"category_id": "CAT-EYES",
"category_name": "Eye Makeup",
"parent_category": "Makeup",
"url": "/shop/makeup/eyes",
"product_count": 48,
"meta_title": "Eye Makeup & Cosmetics | Too Faced",
"meta_description": "Shop cruelty-free eye makeup including mascara, eyeshadow palettes, and eyeliner."
# category_idcategory_nameparent_categoryurlproduct_countbanner_image_url
1
2
3

Capabilities

Cosmetics data extraction — structured and normalised

Our Too Faced scraper handles product variants, shade grids, and ingredient lists — standardising unstructured beauty data into queryable schemas.

Product Metadata Extraction

Extract names, descriptions, how-to-use instructions, and marketing copy for every item in the catalogue.

Shade & Swatch Mapping

Capture shade names, hex codes, undertone descriptions, and swatch image URLs across complex variant selectors.

Ingredient Parsing

Extract and normalise comma-separated ingredient lists into structured arrays for formulation analysis.

Price & Promotion Tracking

Monitor base prices, sale prices, sitewide discounts, and limited-time promotional banners.

Inventory Monitoring

Track out-of-stock status at the SKU level to monitor supply chain gaps and product popularity.

Review & Rating Mining

Extract star ratings, review text, and reviewer attributes like skin type and age range.

Vegan & Cruelty-Free Flags

Detect and structure certification badges and claims for ethical compliance tracking.

Cross-Sell Recommendations

Capture 'Frequently Bought Together' and 'Complete The Look' product associations.

High-Res Asset URLs

Extract CDN links for primary product images, lifestyle shots, and shade swatches.

Category Hierarchies

Map the site navigation structure to understand product taxonomy and placement.

// engagement pipeline

From catalogue to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target categories, product lists, or full-site requirements. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, proxy rotation, and variant hydration logic for toofaced.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, and shade mapping verification before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Handling beauty and cosmetic site structures

Extracting data from highly visual, JavaScript-heavy eCommerce sites requires specific infrastructure. Here is how we handle toofaced.com.

pipeline-monitor · toofaced.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
JavaScript rendering
Hydrating shade selectors

Beauty sites rely on complex JavaScript to swap images and prices when a user selects a shade. We use Playwright to execute these interactions, ensuring we capture exact SKU data for every variant, not just the default load state.

Asset extraction
High-resolution CDN mapping

We parse the DOM and JSON-LD to extract the highest resolution image URLs from the underlying CDN, bypassing compressed thumbnails.

Anti-bot layer
Residential proxy rotation

eCommerce platforms utilise WAFs to block automated traffic. We route requests through US-based residential proxies with realistic TLS fingerprints to maintain uninterrupted access.

Schema normalisation
Structuring ingredient lists

Ingredient lists are often unstructured text blocks. Our pipeline parses these blocks, strips marketing fluff, and emits clean arrays of individual chemical compounds and natural extracts.

Change detection
Tracking inventory shifts

We hash product states per run. When a shade goes out of stock or a price changes, we emit a diff record, providing a precise timeline of inventory velocity and promotional cycles.

Applications

Who uses Too Faced data — and how

Teams across industries use toofaced.com data to build competitive products and smarter operations.

01
Market Research

Beauty analysts track shade ranges, ingredient trends, and new product launches to identify market gaps.

02
Competitor Pricing

Retailers monitor direct-to-consumer pricing, bundle discounts, and sitewide promotions to adjust their own promotional calendars.

03
Assortment Planning

Merchandisers analyse category depth and variant counts to optimise their own brand portfolios.

04
Ingredient Analysis

Formulators track the inclusion of trending active ingredients and the exclusion of banned substances across product lines.

05
Sentiment Analysis

Brand managers ingest review text to correlate product ratings with specific skin types and age demographics.

06
Grey Market Monitoring

Authorised distributors track official MSRPs to identify unauthorised sellers undercutting prices on third-party marketplaces.

Why DataFlirt

"Beauty eCommerce relies on visual variants and ingredient matrices — data that requires precise DOM parsing to normalise into queryable warehouse tables."

Extracting from toofaced.com requires handling dynamic shade selectors, lazy-loaded image grids, and nested ingredient lists. DataFlirt manages the JavaScript rendering, proxy rotation, and schema normalisation so your data engineering team receives clean, structured output without maintaining the scrapers.

Technical Spec

Too Faced scraper — technical capabilities

Everything supported by our toofaced.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for shade selection and dynamic pricing
Supported
Residential proxy rotation
ISP-grade residential IPs to bypass WAF protections
Supported
Shade variant mapping
Extract all child SKUs, hex codes, and stock status per parent product
Supported
Ingredient list parsing
Normalise text blocks into queryable string arrays
Supported
Review pagination
Extract all historical reviews across paginated endpoints
Supported
Change detection (diffs)
Emit records only when price, stock, or copy changes
Supported
Webhook delivery
HTTP POST per record for immediate downstream ingestion
Supported
User account order history
Requires authenticated sessions and customer credentials
Partial
Loyalty point balances
Gated behind individual user authentication walls
Partial
Infrastructure

Infrastructure powering the extraction

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering for dynamic shade selectors and lazy-loaded assets.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies. Rotation happens per-request to prevent IP bans from retail WAFs.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested — schema versioned per run
CSV
Flat file with typed columns — Excel/Sheets compatible
XLS
Excel format for direct business analyst consumption
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery — compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoints to query your extracted datasets
PostgreSQL
Upsert into your existing schema with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About toofaced.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping toofaced.com legal?

Scraping publicly available information is generally permissible under applicable law. DataFlirt targets only public, non-authenticated product, pricing, and review data. We do not extract personal data or circumvent authentication walls.

How do you handle dynamic shade selectors?

We utilise Playwright to execute the JavaScript interactions required to surface variant-specific data, ensuring we capture the correct price, SKU, and image URL for every individual shade.

Can you extract ingredient lists?

Yes. We extract the raw ingredient text blocks and process them into structured arrays, separating active ingredients from base compounds where formatting allows.

How fresh is the pricing data?

Pipelines can be configured to run daily or intra-day. Change detection ensures you receive immediate updates when a price drops or a promotion goes live.

Do you download the product images?

We extract and deliver the high-resolution CDN URLs for the images. If direct binary download is required, we can configure a pipeline to sync assets to your S3 bucket.

How do you handle out-of-stock items?

We capture the inventory status boolean for every SKU. Out-of-stock items are recorded with their last known price and a flag indicating current unavailability.

$ dataflirt scope --new-project --source=toofaced.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Need a one-off catalogue dump or continuous price monitoring? We scope, build, and operate the pipeline. Tell us your requirements.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in beauty and skincare

Services

Data Extraction for Every Industry

View All Services →