SYSTEM all green source italki.com queue 14,892 profiles p99 latency 184ms dataflirt.com · scraper/italki-com
RUN · 41 active pipelines · italki.com live

Italki teacher data,
at warehouse scale.

We extract tutor profiles, pricing structures, availability matrices, and review corpora from Italki. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Tutors extracted
28.4K /run
Price updates
112K /day
Reviews scraped
1.2M /month
Active pipelines
41
Uptime
99.98%
Data Dictionary

Every field we extract from italki.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Teacher Profiles objects from italki.com. All fields typed and schema-versioned.

tutor_iddisplay_nametutor_typeteaches_languagesspeaks_languagesaverage_ratingstudents_countlessons_countvideo_urlabout_me_text
teacher_profiles
● 200 OK
"tutor_id": "T849201",
"display_name": "Maria G.",
"tutor_type": "Professional Teacher",
"average_rating": 4.9,
"students_count": 412,
"lessons_count": 1845,
"teaches_languages": "['Spanish (Native)']"
# tutor_iddisplay_nametutor_typeteaches_languagesspeaks_languagesaverage_rating
1
2
3

Complete list of extractable fields for Pricing & Lessons objects from italki.com. All fields typed and schema-versioned.

tutor_idlesson_titlecategoryduration_minsprice_singleprice_packagetrial_availabletrial_price
pricing_& lessons
● 200 OK
"tutor_id": "T849201",
"lesson_title": "DELE Exam Preparation",
"category": "Test Preparation",
"duration_mins": 60,
"price_single": 25.0,
"price_package": 115.0,
"trial_available": true
# tutor_idlesson_titlecategoryduration_minsprice_singleprice_package
1
2
3

Complete list of extractable fields for Availability Calendar objects from italki.com. All fields typed and schema-versioned.

tutor_iddate_utcstart_time_utcend_time_utcstatusslot_idscraped_atoriginal_timezone
availability_calendar
● 200 OK
"tutor_id": "T849201",
"date_utc": "2026-05-14",
"start_time_utc": "14:00:00",
"end_time_utc": "15:00:00",
"status": "available",
"original_timezone": "Europe/Madrid",
"scraped_at": "2026-05-12T08:11:00Z"
# tutor_iddate_utcstart_time_utcend_time_utcstatusslot_id
1
2
3

Complete list of extractable fields for Reviews objects from italki.com. All fields typed and schema-versioned.

review_idtutor_idstudent_nameratingreview_textreview_datelanguage_taughtlessons_taken
reviews
● 200 OK
"review_id": "REV99281",
"tutor_id": "T849201",
"student_name": "James L.",
"rating": 5,
"review_text": "Maria is incredibly patient and explains grammar perfectly.",
"review_date": "2026-04-22"
# review_idtutor_idstudent_nameratingreview_textreview_date
1
2
3

Complete list of extractable fields for Search Results objects from italki.com. All fields typed and schema-versioned.

keywordrank_positiontutor_iddisplay_namebase_priceratingis_onlinescraped_at
search_results
● 200 OK
"keyword": "spanish native speaker",
"rank_position": 4,
"tutor_id": "T849201",
"base_price": 18.0,
"rating": 4.9,
"is_online": false,
"scraped_at": "2026-05-12T08:15:22Z"
# keywordrank_positiontutor_iddisplay_namebase_pricerating
1
2
3

Capabilities

Everything you need from Italki, nothing you do not

Our Italki scraper handles every layer of the platform: tutor profiles, dynamic availability calendars, complex pricing structures, and the review corpus. JavaScript rendering and anti-bot circumvention are built in.

Tutor Profile Extraction

Capture display names, tutor types, native languages, introduction text, video URLs, and overall statistics scraped at the profile level.

Dynamic Calendar Scraping

Extract available booking slots across a 30-day window. We hydrate the JavaScript calendar widgets and normalise all times to UTC.

Pricing & Package Parsing

Capture individual lesson rates, bulk package discounts, and trial lesson pricing across different lesson categories.

Language Pairing Mapping

Map the exact languages a tutor teaches against the languages they speak, including proficiency levels.

Review & Rating Mining

Extract full review text, star ratings, student details, and lesson counts paginated across all tutor reviews.

Search Rank Tracking

Track tutor visibility for specific language queries and filters, capturing organic rank positions.

Online Status Monitoring

Log the real-time online indicator for tutors to correlate availability with search visibility.

Timezone Normalisation

Italki displays times based on local browser settings. We enforce UTC standardisation across all extracted calendar data.

Scheduled & Streaming Modes

Run one-off bulk exports or configure continuous pipelines at hourly or daily cadences with change-detection diffing.

// engagement pipeline

From language filter to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target languages, tutor types, or specific profile URLs. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Playwright crawlers, proxy rotation, and session management to handle Italki dynamic calendars.

Validation & QA
d 4–6

Schema validation, null-rate checks, and timezone conversion tests before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our Italki pipeline handles the hard parts

Extracting accurate availability and pricing requires deep JavaScript execution and timezone management. Here is how we build resilient pipelines.

pipeline-monitor · italki.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Calendar hydration
Full Playwright execution for dynamic slots

Italki availability calendars are heavily JavaScript-rendered and require complex interaction to paginate through weeks. We run full Playwright browser sessions to trigger lazy-loads and extract all open booking slots.

Timezone normalisation
Converting local slots to UTC

The platform renders calendar times based on the client browser timezone. We force strict UTC contexts in our headless browsers to ensure all extracted availability data is standardised and comparable.

Anti-bot layer
Residential proxy rotation

Frequent requests to tutor search endpoints trigger rate limits. Our crawlers use residential ISP proxies with realistic browser fingerprints and randomised request timing to maintain pipeline stability.

Schema stability
Handling profile layout variations

Tutor profiles vary wildly based on tutor type and completed sections. Our selector strategy uses fallback chains to ensure missing optional fields do not break the extraction of core pricing and rating data.

Change detection
Only re-scrape what has changed

For tracking tutor availability, we maintain a hash index of last-seen slots. Subsequent runs only push diffs, reducing compute cost and downstream processing load.

Applications

Who uses Italki data and how

Teams across industries use italki.com data to build competitive products and smarter operations.

01
EdTech Competitor Analysis

Language learning platforms monitor Italki pricing structures, tutor counts, and lesson types to benchmark their own offerings.

02
Pricing Intelligence

Marketplaces track hourly rates by language and tutor type to optimise their own dynamic pricing algorithms.

03
Tutor Supply Monitoring

Aggregators track the volume of active tutors per language pair to identify supply shortages and recruitment opportunities.

04
Language Demand Forecasting

Analysts correlate review velocity and booked slots with specific languages to measure shifts in global language learning demand.

05
Market Expansion Research

Companies evaluate the density of native speakers offering lessons in emerging markets before launching localised services.

06
AI Training Data

ML teams use structured tutor profiles and review text to train matching algorithms and sentiment analysis models.

Why DataFlirt

"Italki holds the most comprehensive dataset on global language tutoring rates and availability, but none of it is queryable unless you build the pipeline."

Most teams underestimate the investment required: reliable Italki scraping requires residential proxies, full JavaScript rendering for dynamic calendars, daily selector maintenance, and complex timezone normalisation. DataFlirt absorbs that complexity so your engineers can focus on the analysis, not the infrastructure.

Technical Spec

Italki scraper technical capabilities

Everything supported by our italki.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for availability calendars and dynamic pricing tabs
Supported
CAPTCHA bypass
Automated 2Captcha + CapSolver integration for rate-limit walls
Supported
Residential proxy rotation
ISP-grade residential IPs rotated per request to prevent blocking
Supported
Availability calendar parsing
Extraction of 30-day booking windows for any tutor
Supported
Timezone conversion to UTC
All scraped times normalised to UTC regardless of proxy location
Supported
Video URL extraction
Capture of raw introduction video source URLs
Supported
Review pagination
Full review corpus extraction across all pages
Supported
Change detection (diffs)
Hash-based diff to only emit records with changed fields since last run
Supported
Student private messages
Gated data requires user authentication and violates terms
Partial
Booking transaction data
Private financial records between student and platform
Partial
Infrastructure

Infrastructure powering the Italki pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering, calendar pagination, and interaction flows.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies. Rotation happens per-request with sticky sessions where required to load continuous calendar views.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested schema versioned per run
CSV
Flat file with typed columns for quick analysis
XLS
Excel format for non-technical operations teams
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoints to query your extracted datasets
BigQuery
Streamed directly into your dataset with schema auto-detect
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About italki.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Italki legal?

Scraping publicly available information from Italki is generally permissible. DataFlirt targets only public, non-authenticated tutor profiles, pricing, calendars, and review data. We do not extract private student data or circumvent authentication walls.

How do you handle the dynamic availability calendars?

We use full Playwright browser sessions to render the JavaScript calendar widgets. Our crawlers paginate through the UI to extract all available slots and strictly convert local browser times to UTC for standardisation.

Can you track tutor pricing changes over time?

Yes. Every pipeline run produces timestamped snapshots. We maintain a time-series table per tutor for base rates, package prices, and trial lesson costs.

How fresh is the data?

Full catalogue refreshes at daily cadence complete within a 4-8 hour window. For specific subsets of high-volume tutors, we can configure hourly runs to track real-time calendar availability.

What is the minimum viable engagement?

Our smallest packages start at a defined set of languages or a specific list of tutor URLs with weekly delivery. We price based on volume and delivery frequency.

Can I request a sample dataset before committing?

Absolutely. We provide a sample run of up to 500 tutor profiles as part of the pre-engagement scoping process so you can validate schema fit and data quality.

$ dataflirt scope --new-project --source=italki.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off tutor catalogue dump or a continuous availability monitoring feed, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in education and courses

Services

Data Extraction for Every Industry

View All Services →