SYSTEM all green source ipsy.com queue 12,845 pages p99 latency 185ms dataflirt.com · scraper/ipsy-com
RUN · 31 active pipelines · ipsy.com live

Ipsy beauty data,
at warehouse scale.

We extract product listings, brand catalogues, pricing signals, ingredient lists, and demographic-tagged reviews from Ipsy. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Products extracted
45.2K /day
Brand profiles
3.1K /run
Review records
1.2M /month
Active pipelines
31
Uptime
99.98%
Data Dictionary

Every field we extract from ipsy.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Product Listings objects from ipsy.com. All fields typed and schema-versioned.

product_idtitlebrand_namecategorysub_categorymsrpipsy_pricedescriptionhow_to_useingredientssize_mlis_full_sizeimage_urlsproduct_urlis_clean_beauty
product_listings
● 200 OK
"product_id": "p-128491",
"title": "Watermelon Glow Niacinamide Dew Drops",
"brand_name": "Glow Recipe",
"msrp": 35.0,
"ipsy_price": 18.0,
"category": "Skincare",
"is_full_size": true,
"is_clean_beauty": true
# product_idtitlebrand_namecategorysub_categorymsrp
1
2
3

Complete list of extractable fields for Brand Profiles objects from ipsy.com. All fields typed and schema-versioned.

brand_idbrand_namedescriptionwebsite_urlcountry_of_origincruelty_freeveganfeatured_productstotal_products_on_ipsybrand_logo_url
brand_profiles
● 200 OK
"brand_id": "b-4921",
"brand_name": "Tarte Cosmetics",
"country_of_origin": "USA",
"cruelty_free": true,
"vegan": false,
"total_products_on_ipsy": 42,
"website_url": "https://tartecosmetics.com"
# brand_idbrand_namedescriptionwebsite_urlcountry_of_origincruelty_free
1
2
3

Complete list of extractable fields for Reviews & Ratings objects from ipsy.com. All fields typed and schema-versioned.

review_idproduct_iduser_nicknamestar_ratingreview_textskin_typeskin_toneeye_colorhair_colorage_rangereview_datehelpful_votes
reviews_& ratings
● 200 OK
"review_id": "r-992814",
"product_id": "p-128491",
"star_rating": 5,
"skin_type": "Combination",
"skin_tone": "Medium",
"eye_color": "Brown",
"review_date": "2023-11-14T08:22:00Z"
# review_idproduct_iduser_nicknamestar_ratingreview_textskin_type
1
2
3

Complete list of extractable fields for Subscription Archives objects from ipsy.com. All fields typed and schema-versioned.

month_yearbox_typeproduct_idproduct_titlebrandis_full_sizeretail_valuethemecuratorimage_url
subscription_archives
● 200 OK
"month_year": "2023-10",
"box_type": "BoxyCharm",
"product_id": "p-88312",
"product_title": "Liquid Lash Extensions Mascara",
"brand": "Thrive Causemetics",
"is_full_size": true,
"retail_value": 25.0
# month_yearbox_typeproduct_idproduct_titlebrandis_full_size
1
2
3

Complete list of extractable fields for Drop Shop Pricing objects from ipsy.com. All fields typed and schema-versioned.

product_idtitlebranddrop_shop_pricemsrpdiscount_pctstock_statussale_event_namesale_start_datesale_end_date
drop_shop pricing
● 200 OK
"product_id": "p-128491",
"drop_shop_price": 12.0,
"msrp": 35.0,
"discount_pct": 65,
"stock_status": "in_stock",
"sale_event_name": "Mega Drop Shop",
"sale_end_date": "2023-11-20T23:59:59Z"
# product_idtitlebranddrop_shop_pricemsrpdiscount_pct
1
2
3

Capabilities

Everything you need from Ipsy — nothing you don't

Our Ipsy scraper navigates dynamic React frontends and drop-shop inventory systems to extract structured beauty catalogues, ingredient lists, and demographic-tagged reviews.

Full Product Data Extraction

Title, brand, MSRP, Ipsy pricing, description, usage instructions, and size variants — scraped at the product level.

Brand Intelligence

Extract brand origins, cruelty-free status, vegan certifications, and complete brand product portfolios hosted on Ipsy.

Ingredient Parsing

Capture complete INCI ingredient lists for chemical analysis, formulation tracking, and clean-beauty compliance checks.

Demographic Review Mining

Extract reviews correlated with user skin type, skin tone, eye colour, and age range — vital for targeted formulation research.

Subscription Box Archives

Historical data on Glam Bag, BoxyCharm, and Icon Box configurations, including curator themes and full-size vs sample ratios.

Drop Shop Pricing

Monitor flash sale events, discount depths, and inventory stock-outs during Mega Drop Shop windows.

Shade & Variant Mapping

Extract all available foundation, concealer, and lip shades tied to a single product ID with corresponding hex codes.

Clean Beauty Tagging

Identify products flagged under Ipsy's clean beauty standards, extracting specific free-from claims.

Scheduled + Streaming Modes

Run one-off bulk exports or configure continuous pipelines at daily cadences to track inventory changes.

// engagement pipeline

From brand list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target categories, brand names, or historical box months. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, handle React SPA routing, and manage session cookies for ipsy.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, and ingredient list normalisation before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our Ipsy pipeline handles the hard parts

Ipsy relies on heavy client-side rendering and aggressive caching. Here is how we extract reliable data without triggering rate limits.

pipeline-monitor · ipsy.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
SPA Rendering
Full Playwright execution for Next.js content

Ipsy's product pages and drop shop interfaces are heavily JavaScript-rendered. We run full Playwright browser sessions with JS execution and API interception to capture JSON payloads directly from Next.js hydration states.

Anti-bot layer
Residential proxy rotation

Frequent requests to Ipsy's catalogue trigger WAF blocks. We route all traffic through US-based residential ISP proxies with realistic browser fingerprints to maintain high concurrency without IP bans.

Dynamic Inventory
Flash sale tracking during Drop Shop

During Mega Drop Shop events, inventory states and prices change rapidly. Our pipelines scale concurrency dynamically to capture discount depths and out-of-stock signals before the event window closes.

Schema stability
Resilient selectors with fallback chains

Ipsy updates its frontend components frequently. We use multiple fallback chains per field — CSS selectors, XPath, and direct API response parsing — ensuring pipeline stability across deployments.

Monitoring & alerting
24/7 pipeline health

Every run emits structured logs to our observability stack. We alert on null-rate spikes in critical fields like ingredient lists or pricing, responding before data reaches your warehouse.

Applications

Who uses Ipsy data — and how

Teams across industries use ipsy.com data to build competitive products and smarter operations.

01
Market Research & Category Analysis

Beauty conglomerates track rising indie brands, product categories, and clean beauty adoption rates within subscription boxes.

02
Competitor Pricing Intelligence

Retailers monitor Ipsy's Drop Shop discounts to understand deep-discount strategies and MAP adherence by partner brands.

03
Ingredient & Formulation Trends

R&D teams parse INCI lists across thousands of SKUs to identify trending active ingredients and formulation shifts.

04
Consumer Sentiment & Review Mining

Brands analyse reviews filtered by skin type and tone to identify demographic-specific formulation flaws or marketing opportunities.

05
Brand Equity Monitoring

Investors track brand presence across Glam Bag and BoxyCharm tiers to gauge brand positioning and consumer reception.

06
Subscription Box Benchmarking

Competing subscription services analyse Ipsy's monthly curation value, full-size ratios, and brand partnerships.

Why DataFlirt

"Ipsy holds a uniquely structured dataset mapping beauty product reviews to specific skin tones, types, and concerns — invaluable for formulation research."

Extracting Ipsy data requires navigating heavily cached SPA frontends, dynamic inventory drops, and strict rate limits. DataFlirt manages the residential proxies, headless browser execution, and schema maintenance so your data science teams can focus on beauty trend analysis rather than pipeline engineering.

Technical Spec

Ipsy scraper — technical capabilities

Everything supported by our ipsy.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for Next.js hydration and dynamic loading
Supported
CAPTCHA bypass
Automated solver integration for aggressive WAF challenges
Supported
Residential proxy rotation
US-based ISP residential IPs rotated per request
Supported
Ingredient parsing
Extraction of full INCI lists from product description accordions
Supported
Demographic review extraction
Capture of skin type, tone, and age range tied to specific user reviews
Supported
Drop Shop pricing
Tracking flash sale discounts and inventory states during event windows
Supported
Personal Beauty Quiz results
User-specific quiz answers and algorithmic matches
Partial
User billing/subscription history
Private account-level billing and individual bag tracking
Partial
Infrastructure

Infrastructure powering the Ipsy pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering, API interception, and React hydration state extraction.

Residential Proxy Infrastructure

We maintain pools of US-based residential ISP proxies to bypass strict rate limits and geo-fencing applied to beauty catalogues.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting. State is stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested — schema versioned per run
CSV
Flat file with typed columns — Excel/Sheets compatible
Parquet
Columnar format for BigQuery, Snowflake, Athena
S3
Direct bucket delivery — compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoints for querying specific product or brand records
BigQuery
Streamed directly into your dataset with schema auto-detect
Snowflake
Stage + COPY INTO workflow — incremental or full-replace
Postgres
Upsert into your existing schema with conflict resolution
// faq

Common questions.

About ipsy.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Ipsy legal?

Scraping publicly available product, brand, and review data is generally permissible. DataFlirt targets only public catalogues and non-authenticated endpoints. We do not extract personal user data or circumvent authentication walls to access private billing information.

Can you track Mega Drop Shop pricing?

Yes. We can configure burst pipelines to run during specific flash sale windows, capturing deep discounts, MSRP comparisons, and inventory depletion rates before the shop closes.

Do you extract demographic data from reviews?

Yes. Ipsy reviews often include the user's self-reported skin type, skin tone, eye colour, and age range. We extract these fields alongside the review text and star rating for demographic analysis.

How do you handle Ipsy's dynamic React frontend?

We use Playwright to execute JavaScript and intercept background API calls, allowing us to extract clean JSON payloads directly from the Next.js application state rather than parsing complex HTML DOM structures.

Can I get historical Glam Bag data?

We can extract historical box configurations (Glam Bag, BoxyCharm, Icon Box) that remain publicly accessible on the platform, including curator themes and included product IDs.

What is the delivery cadence?

We support one-off historical exports, weekly catalogue refreshes, or daily delta updates depending on your analytical requirements and storage constraints.

$ dataflirt scope --new-project --source=ipsy.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a full ingredient database or continuous tracking of subscription box trends — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in beauty and skincare

Services

Data Extraction for Every Industry

View All Services →