SYSTEM all green source upcity.com queue 12,841 pages p99 latency 184ms dataflirt.com · scraper/upcity-com
RUN · 41 active pipelines · upcity.com live

UpCity agency data,
at warehouse scale.

We extract B2B service provider profiles, verified reviews, pricing models, and service matrices from UpCity. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your schedule.

Providers extracted
114K /run
Reviews processed
312K /month
Category scans
4.2K /day
Active pipelines
41
Uptime
99.98%
Data Dictionary

Every field we extract from upcity.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Agency Profiles objects from upcity.com. All fields typed and schema-versioned.

agency_idnameprofile_urllocationratingreview_countmin_project_sizeavg_hourly_rateemployee_countfounded_yeardescriptionwebsite_urlverified_statusexcellence_award_winner
agency_profiles
● 200 OK
"agency_id": "UC-84921",
"name": "Acme Digital Marketing",
"rating": 4.9,
"review_count": 42,
"min_project_size": "$5,000+",
"avg_hourly_rate": "$100 - $149",
"employee_count": "10 - 49",
"verified_status": true
# agency_idnameprofile_urllocationratingreview_count
1
2
3

Complete list of extractable fields for Services Matrix objects from upcity.com. All fields typed and schema-versioned.

agency_idservice_categoryservice_namepercentage_focusindustry_focusclient_focusframeworks_supportedscraped_at
services_matrix
● 200 OK
"agency_id": "UC-84921",
"service_category": "Marketing",
"service_name": "SEO",
"percentage_focus": 40,
"industry_focus": "['Healthcare', 'Finance']",
"client_focus": "['Midmarket', 'Small Business']"
# agency_idservice_categoryservice_namepercentage_focusindustry_focusclient_focus
1
2
3

Complete list of extractable fields for Verified Reviews objects from upcity.com. All fields typed and schema-versioned.

review_idagency_idreviewer_namereviewer_titlereviewer_companystar_ratingreview_datereview_textverified_statusproject_type
verified_reviews
● 200 OK
"review_id": "REV-99214",
"agency_id": "UC-84921",
"reviewer_title": "CMO",
"star_rating": 5.0,
"review_date": "2026-03-14",
"verified_status": true,
"project_type": "Website Redesign"
# review_idagency_idreviewer_namereviewer_titlereviewer_companystar_rating
1
2
3

Complete list of extractable fields for Portfolios objects from upcity.com. All fields typed and schema-versioned.

portfolio_idagency_idproject_titleclient_nameindustryproject_summarybudgetlaunch_dateimage_urls
portfolios
● 200 OK
"portfolio_id": "PF-1102",
"agency_id": "UC-84921",
"project_title": "FinTech App Launch",
"industry": "Financial Services",
"budget": "$50,000+",
"launch_date": "2025-11-01",
"image_urls": "['https://example.com/img1.jpg']"
# portfolio_idagency_idproject_titleclient_nameindustryproject_summary
1
2
3

Complete list of extractable fields for Search Rankings objects from upcity.com. All fields typed and schema-versioned.

keywordlocationrank_positionagency_nameagency_idsponsored_statusupcity_scorescraped_at
search_rankings
● 200 OK
"keyword": "SEO Agencies",
"location": "Chicago, IL",
"rank_position": 3,
"agency_id": "UC-84921",
"sponsored_status": false,
"upcity_score": 88
# keywordlocationrank_positionagency_nameagency_idsponsored_status
1
2
3

Capabilities

Everything you need from UpCity — nothing you don't

Our UpCity scraper handles the directory hierarchy: category pagination, provider profiles, verified review expansion, and service matrix extraction — with bot circumvention built in.

Full Profile Extraction

Extract agency name, website, contact details, employee count, minimum project size, and average hourly rate from every provider profile.

Verified Review Mining

Capture star ratings, review text, reviewer job titles, and project types across all paginated review tabs.

Service Matrix Parsing

Extract the exact percentage breakdown of services offered, industry focus, and client size focus for precise vendor matching.

Award & Certification Tracking

Identify UpCity Excellence Award winners and track verified partner status across specific technology stacks.

Location & Branch Mapping

Extract headquarters and secondary branch locations, including full address strings and geographic coordinates if available.

Category Rank Tracking

Monitor organic and sponsored positions for agencies across specific service categories and city-level directories.

Portfolio & Case Study Data

Extract project summaries, client names, budgets, and industry tags from agency portfolio sections.

Change Detection

Run recurring pipelines that only emit records when an agency updates their pricing, adds a new review, or changes service focus.

Scheduled Exports

Configure pipelines to run weekly or monthly to keep your B2B lead database or vendor management system fresh.

// engagement pipeline

From category list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target categories, locations, or specific agency URLs. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy crawlers, proxy rotation, and session management to traverse the UpCity directory hierarchy.

Validation & QA
d 4–6

Schema validation, null-rate checks, and normalisation of pricing tiers before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our UpCity pipeline handles the hard parts

Directory sites employ rate limiting and structural variations. Here is how we build resilient pipelines for UpCity data.

pipeline-monitor · upcity.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Rate limit evasion
Distributed crawling with residential IPs

Directory sites block datacenter IPs scanning thousands of profiles. We distribute requests across a pool of US-based residential proxies, pacing requests to mimic human browsing behaviour and avoid IP bans.

Dynamic DOM structures
Resilient selectors for varied profiles

Agency profiles on UpCity vary based on their subscription tier. Premium profiles have different DOM structures than free listings. Our selector strategy accounts for these variations to ensure consistent schema output.

Review pagination
Deep extraction of historical reviews

Agencies with hundreds of reviews require deep pagination. Our pipeline maintains session state to traverse all review pages, capturing historical sentiment without dropping records.

Nested service matrices
Normalising complex percentage breakdowns

UpCity displays service focus as visual bar charts or nested lists. We extract the underlying percentage values and normalise them into structured arrays for easy querying in your database.

Data normalisation
Standardised pricing and employee tiers

We clean and standardise string-based ranges (e.g., '$100 - $149/hr', '10 - 49 employees') into consistent formats, making the data immediately usable for filtering and analysis.

Applications

Who uses UpCity data — and how

Teams across industries use upcity.com data to build competitive products and smarter operations.

01
B2B Lead Generation

Sales teams targeting marketing agencies, IT firms, and accountants extract provider lists to enrich outbound campaigns.

02
Competitor Benchmarking

Agencies track competitor pricing models, service matrices, and review velocity to position themselves effectively in the market.

03
Partner Discovery

Software vendors identify top-rated implementation partners and agencies based on specific framework expertise and location.

04
Market Research

Analysts aggregate hourly rates and minimum project sizes across cities to map regional pricing trends in B2B services.

05
Reputation Management

Agencies ingest their own reviews and competitor reviews into BI tools for sentiment analysis and service improvement.

06
M&A Sourcing

Private equity firms screen potential acquisition targets by filtering for highly-rated agencies in specific niches with defined employee counts.

Why DataFlirt

"UpCity maps the B2B service ecosystem, but extracting structured intelligence on thousands of agencies requires dedicated pipeline infrastructure."

Directory scraping seems simple until you hit structural inconsistencies, rate limits, and nested review pagination. DataFlirt handles the proxy rotation, schema normalisation, and scheduling so your engineering team can focus on data modelling instead of crawler maintenance.

Technical Spec

UpCity scraper — technical capabilities

Everything supported by our upcity.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Playwright integration for dynamic charts and lazy-loaded profile elements
Supported
Residential proxy rotation
ISP-grade IPs to bypass directory rate limits
Supported
Review pagination
Full extraction of all historical reviews per agency
Supported
Category traversal
Automated discovery of agencies across all industry and location nodes
Supported
Service matrix normalisation
Conversion of visual service focus charts into structured percentage arrays
Supported
Change detection (diffs)
Only emit records when an agency updates their profile or receives a new review
Supported
Webhook delivery
HTTP POST per record for real-time CRM ingestion
Supported
User account credentials
Extraction of private dashboard analytics or internal lead data
Partial
Direct messaging
Automated submission of contact forms to agencies
Partial
Infrastructure

Infrastructure powering the UpCity pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows. Combined via scrapy-playwright middleware.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across US regions. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.

Cloud-Native Orchestration

Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested arrays
CSV
Flat file with typed columns
XLS
Excel format for non-technical teams
Parquet
Columnar format for data warehouses
AWS S3
Direct delivery to your cloud storage
Webhook
HTTP POST for event-driven architectures
API
REST endpoints to query your extracted datasets
BigQuery
Streamed directly into your dataset
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About upcity.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping UpCity legal?

Scraping publicly available directory information is generally permissible under applicable law, reinforced by the hiQ v. LinkedIn ruling. DataFlirt extracts only public, non-authenticated agency profiles and reviews. We do not bypass authentication walls or extract private user data.

How do you handle UpCity's rate limits?

We route requests through a large pool of residential proxies, ensuring our request volume per IP remains well below typical bot-detection thresholds. We also introduce randomised delays between requests.

Can you extract data for specific cities only?

Yes. We can scope the pipeline to target specific geographic URLs (e.g., SEO agencies in Austin, TX) or specific service categories, reducing unnecessary data extraction.

How do you format the service matrix percentages?

We extract the service focus percentages and output them as structured JSON arrays or nested columns in CSV, ensuring the numbers sum correctly and map to the specific service categories.

Can I get historical reviews?

Yes. Our initial run can paginate through the entire review history of an agency. Subsequent runs can be configured to only extract new reviews added since the last execution.

What is the delivery frequency?

Pipelines can be scheduled for one-off delivery, weekly updates, or monthly refreshes depending on how frequently you need the agency data updated in your systems.

Do you provide a sample of the UpCity data?

Yes. We offer a sample extraction of up to 100 agency profiles during the scoping phase so you can validate the schema and data quality before committing to a production pipeline.

$ dataflirt scope --new-project --source=upcity.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off agency directory dump or continuous tracking of B2B service providers — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in business directories

Services

Data Extraction for Every Industry

View All Services →