SYSTEM all green source clearbit.com queue 12,943 domains p99 latency 218ms dataflirt.com · scraper/clearbit-com
RUN · 84 active pipelines · clearbit.com live

Clearbit data,
at warehouse scale.

We extract company profiles, technographics, funding rounds, and social intelligence from Clearbit. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Companies extracted
842K /day
Tech stack updates
3.1M /24h
Domain resolutions
1.4M /run
Active pipelines
84
Uptime
99.98%
Data Dictionary

Every field we extract from clearbit.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Company Profile objects from clearbit.com. All fields typed and schema-versioned.

domaincompany_namelegal_namedescriptiontagsindustrysectorfounded_yearlocationtimezonelogo_url
company_profile
● 200 OK
"domain": "stripe.com",
"company_name": "Stripe",
"description": "Financial infrastructure platform for the internet.",
"industry": "Financial Services",
"founded_year": 2010,
"location": "San Francisco, CA"
# domaincompany_namelegal_namedescriptiontagsindustry
1
2
3

Complete list of extractable fields for Firmographics objects from clearbit.com. All fields typed and schema-versioned.

domainemployee_countemployee_rangeestimated_annual_revenuerevenue_rangefiscal_year_endpublic_tickermarket_capnaics_codesic_code
firmographics
● 200 OK
"domain": "stripe.com",
"employee_range": "5000-10000",
"revenue_range": "$1B-$10B",
"public_ticker": "None",
"naics_code": "522320",
"sic_code": "7389"
# domainemployee_countemployee_rangeestimated_annual_revenuerevenue_rangefiscal_year_end
1
2
3

Complete list of extractable fields for Technographics objects from clearbit.com. All fields typed and schema-versioned.

domaintech_categorytechnology_namefirst_detectedlast_detectedproviderimplementation_typeactive_status
technographics
● 200 OK
"domain": "stripe.com",
"tech_category": "Analytics",
"technology_name": "Google Analytics",
"first_detected": "2015-04-12",
"active_status": true,
"provider": "Google"
# domaintech_categorytechnology_namefirst_detectedlast_detectedprovider
1
2
3

Complete list of extractable fields for Social & Web objects from clearbit.com. All fields typed and schema-versioned.

domainlinkedin_urltwitter_handlefacebook_urlcrunchbase_urlalexa_ranktraffic_metricsprimary_languagephone_number
social_& web
● 200 OK
"domain": "stripe.com",
"linkedin_url": "linkedin.com/company/stripe",
"twitter_handle": "stripe",
"crunchbase_url": "crunchbase.com/organization/stripe",
"alexa_rank": 1423,
"phone_number": "+1-888-902-3340"
# domainlinkedin_urltwitter_handlefacebook_urlcrunchbase_urlalexa_rank
1
2
3

Complete list of extractable fields for Funding Data objects from clearbit.com. All fields typed and schema-versioned.

domaintotal_funding_usdlatest_funding_datelatest_funding_amountlatest_funding_stageinvestor_namesvaluationacquisition_status
funding_data
● 200 OK
"domain": "stripe.com",
"total_funding_usd": 8700000000,
"latest_funding_date": "2023-03-15",
"latest_funding_stage": "Series I",
"investor_names": "['Sequoia Capital', 'Andreessen Horowitz']",
"acquisition_status": "Private"
# domaintotal_funding_usdlatest_funding_datelatest_funding_amountlatest_funding_stageinvestor_names
1
2
3

Capabilities

Company intelligence extracted precisely

Our Clearbit scraper targets domain-level firmographics, technology stacks, and social identifiers — handling rate limits, bot detection, and pagination automatically.

Firmographic Data Extraction

Capture revenue ranges, employee counts, NAICS codes, and founding dates for B2B segmentation.

Technographic Profiling

Extract active technology stacks, SaaS integrations, and hosting providers associated with the domain.

Social Intelligence

Scrape linked Twitter, LinkedIn, Facebook, and Crunchbase profiles for cross-platform enrichment.

Location & Hierarchy

Map HQ addresses, timezones, and geographic coordinates for territory planning.

Financial & Funding Data

Track public tickers, estimated revenue brackets, and recent funding rounds.

Domain Resolution

Resolve company names to canonical domains and extract primary web presence metrics.

Rate Limit Management

Distributed crawling architecture handles Clearbit's strict request throttling without IP bans.

Change Detection

Only emit records when a company's tech stack, employee count, or funding status changes.

Multi-Format Delivery

Push structured JSON or Parquet directly into BigQuery, Snowflake, or S3.

// engagement pipeline

From domain list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target domains, industry segments, or company names. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, proxy rotation, session management, and CAPTCHA handling for clearbit.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, and technographic coverage analysis before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Bypassing Clearbit's extraction defenses

Clearbit protects its data graph aggressively. Here is how our infrastructure maintains constant throughput.

pipeline-monitor · clearbit.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Residential proxies
ISP-grade IP rotation

Clearbit blocks datacentre IPs instantly. We route requests through high-reputation residential ISP proxies.

Rate limit circumvention
Distributed request architecture

We distribute requests across thousands of sessions to stay below threshold triggers and avoid 429 responses.

Header spoofing
Legitimate TLS fingerprints

Our HTTP clients mimic legitimate browser fingerprints, matching expected JA3/JA4 hashes to bypass WAF rules.

Delta extraction
Hash-based change detection

We hash firmographic states and only push updates when a company's data changes, reducing warehouse compute.

Schema resilience
Fallback JSON parsing

Clearbit updates its response structures. We use fallback JSON path extraction to prevent pipeline breaks.

Applications

Who uses Clearbit data — and how

Teams across industries use clearbit.com data to build competitive products and smarter operations.

01
B2B Lead Scoring

RevOps teams enrich inbound leads with firmographics to route high-value prospects to enterprise sales.

02
TAM Analysis

Strategy teams map total addressable market by extracting companies within specific revenue and NAICS brackets.

03
Competitor Intelligence

Product marketers track which companies are dropping or adopting competitor technology stacks.

04
Investment Sourcing

VC and PE firms monitor employee growth rates and funding stages to identify breakout startups.

05
ABM Campaigns

Marketing teams build hyper-targeted account lists based on exact technographic profiles.

06
Data Enrichment Platforms

Second-party data providers use our pipelines to update their internal company graphs.

Why DataFlirt

"Clearbit holds the definitive B2B data graph, but mapping millions of domains requires infrastructure that can handle aggressive rate limiting and fingerprinting."

Extracting firmographics at scale means managing thousands of residential proxies, handling complex JSON schema variations, and maintaining state across millions of domains. DataFlirt absorbs this operational overhead. We deliver clean, normalised company records directly to your warehouse so your data engineers can focus on modelling, not scraping.

Technical Spec

Clearbit scraper — technical capabilities

Everything supported by our clearbit.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Firmographic extraction
Revenue ranges, employee counts, and NAICS codes
Supported
Technographic detection
Active tech stack categories and specific providers
Supported
Social link scraping
LinkedIn, Twitter, and Crunchbase URL mapping
Supported
Residential proxy rotation
ISP-grade residential IPs to prevent blocks
Supported
Change detection (diffs)
Hash-based diff: only emit records with changed fields since last run
Supported
Webhook delivery
HTTP POST per record for real-time lead enrichment
Supported
Bulk domain resolution
Resolve raw company names to canonical domains
Supported
Contact-level emails
Individual employee email addresses and direct dials (requires Clearbit authenticated API)
Partial
Intent data signals
Clearbit Reveal web traffic intent (requires proprietary tracking pixel)
Partial
Infrastructure

Infrastructure powering the Clearbit pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Distributed Crawling

Scrapy orchestrates high-concurrency requests across millions of target domains, managed by Airflow.

Residential Network

Requests route through ISP-grade residential proxies to bypass IP bans and rate limits.

Cloud-Native Storage

Pipelines run on AWS ECS. State and deduplication hashes reside in managed PostgreSQL and Redis.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested — schema versioned per run
CSV
Flat file with typed columns — Excel/Sheets compatible
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery — compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoint for on-demand domain enrichment
XLS
Excel format for non-technical business teams
BigQuery
Streamed directly into your dataset with schema auto-detect
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About clearbit.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Clearbit legal?

Scraping publicly accessible company data is generally permissible. We target public firmographics and technographics, not private contact information. Clients must review Terms of Service.

How do you handle rate limits?

We distribute requests across a massive pool of residential IPs, ensuring no single IP exceeds threshold limits.

Can I provide a list of domains to enrich?

Yes. You can supply an S3 bucket or CSV of domains, and we will configure the pipeline to target only those entities.

Do you extract individual employee emails?

No. We extract domain-level firmographics and technographics. Individual PII and contact data requires an authenticated Clearbit API subscription.

How fresh is the technographic data?

We scrape the latest state presented by the platform. You can configure pipelines to run weekly or monthly to track technology adoption curves.

What happens if the schema changes?

Our extraction logic uses fallback JSON paths. If a structural change causes null-rate spikes, our monitoring alerts us and we patch the selectors.

$ dataflirt scope --new-project --source=clearbit.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need to enrich 10,000 domains or map millions of company profiles — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in business directories

Services

Data Extraction for Every Industry

View All Services →