SYSTEM all green source globusjourneys.com queue 3,194 tours p99 latency 218ms dataflirt.com · scraper/globusjourneys-com
RUN · 42 active pipelines · globusjourneys.com live

Globus tour data,
at warehouse scale.

We extract escorted tour itineraries, dynamic pricing schedules, departure availability, and accommodation details from Globus. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Itineraries extracted
1,492 /run
Departure dates
18,340 /day
Price updates
32,105 /24h
Active pipelines
42
Uptime
99.98%
Data Dictionary

Every field we extract from globusjourneys.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Tour Overviews objects from globusjourneys.com. All fields typed and schema-versioned.

tour_idtour_nametravel_styledestination_regioncountries_visitedduration_daysbase_pricecurrencyratingreview_count
tour_overviews
● 200 OK
"tour_id": "GE",
"tour_name": "European Tapestry",
"travel_style": "Classic Tour",
"duration_days": 14,
"base_price": 3499.0,
"currency": "USD",
"rating": 4.6,
"review_count": 128
# tour_idtour_nametravel_styledestination_regioncountries_visitedduration_days
1
2
3

Complete list of extractable fields for Itinerary Details objects from globusjourneys.com. All fields typed and schema-versioned.

tour_idday_numberday_titledestination_cityactivitiesmeals_includedaccommodation_nameaccommodation_ratingoptional_excursions
itinerary_details
● 200 OK
"tour_id": "GE",
"day_number": 3,
"day_title": "Amsterdam to Rhine Cruise",
"destination_city": "Amsterdam",
"meals_included": "['Breakfast', 'Dinner']",
"accommodation_name": "Hotel NH Amsterdam",
"optional_excursions": "['Canal Cruise']"
# tour_idday_numberday_titledestination_cityactivitiesmeals_included
1
2
3

Complete list of extractable fields for Pricing & Departures objects from globusjourneys.com. All fields typed and schema-versioned.

tour_iddeparture_datereturn_dateavailability_statusprice_per_personsingle_supplementpromotionsguaranteed_departurebooking_url
pricing_& departures
● 200 OK
"tour_id": "GE",
"departure_date": "2025-05-14",
"return_date": "2025-05-27",
"availability_status": "Available",
"price_per_person": 3499.0,
"guaranteed_departure": true,
"promotions": "['Save $200 per couple']"
# tour_iddeparture_datereturn_dateavailability_statusprice_per_personsingle_supplement
1
2
3

Complete list of extractable fields for Accommodations objects from globusjourneys.com. All fields typed and schema-versioned.

tour_idhotel_namecitycountrystar_ratingamenitiesdescriptionimage_urlsnights_stayed
accommodations
● 200 OK
"tour_id": "GE",
"hotel_name": "Hotel NH Amsterdam",
"city": "Amsterdam",
"country": "Netherlands",
"star_rating": 4,
"nights_stayed": 2,
"amenities": "['Free WiFi', 'Fitness Centre']"
# tour_idhotel_namecitycountrystar_ratingamenities
1
2
3

Complete list of extractable fields for Reviews objects from globusjourneys.com. All fields typed and schema-versioned.

review_idtour_idreviewer_nametravel_dateoverall_ratingitinerary_ratingguide_ratingreview_texthelpful_votes
reviews
● 200 OK
"review_id": "REV-84920",
"tour_id": "GE",
"reviewer_name": "Sarah J.",
"travel_date": "2024-09-12",
"overall_rating": 5,
"guide_rating": 5,
"review_text": "Excellent tour director and perfectly paced itinerary.",
"helpful_votes": 12
# review_idtour_idreviewer_nametravel_dateoverall_ratingitinerary_rating
1
2
3

Capabilities

Everything you need from Globus Journeys

Our Globus scraper handles dynamic booking engine state, availability calendars, and regional pricing grids, with full JavaScript rendering and session management built in.

Full Itinerary Extraction

Day-by-day breakdown, meals, activities, and hotels. Scraped at the individual tour level with all text and metadata.

Live Departure Tracking

Monitor exact dates, guaranteed departures, and waitlist status for thousands of scheduled trips.

Dynamic Pricing Schedules

Capture base rates, seasonal adjustments, and single supplements directly from the booking matrix.

Travel Style Categorisation

Filter tours by Choice Touring, Undiscovered, or Independence by Globus classifications.

Accommodation Intelligence

Extract hotel properties mapped to specific tour days, including star ratings and amenities.

Optional Excursion Data

Scrape add-on activities, pricing, and duration metrics offered during free time on the tour.

Promotional Offer Tracking

Log early booking discounts, airfare credits, and seasonal sales tied to specific departure dates.

Cross-Brand Mapping

Link Globus tours to sister brands like Cosmos and Avalon Waterways for comparative analysis.

Scheduled + Streaming Modes

Run daily availability checks or full catalogue syncs weekly with change-detection diffing.

// engagement pipeline

From tour list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target regions, tour styles, or specific package URLs. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers and handle Globus booking engine state and session cookies.

Validation & QA
d 4–6

Schema validation, null-rate checks, and price-outlier detection before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket or Snowflake stage on agreed cadence.

Under the hood

How our Globus pipeline handles the hard parts

Travel sites rely on complex booking engines and regional pricing. Here is how we maintain data accuracy.

pipeline-monitor · globusjourneys.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Booking engine state
Sticky session management

Globus relies on complex session cookies to render date-specific pricing. We maintain sticky sessions across requests to extract accurate quotes and availability matrices without triggering timeouts.

JavaScript rendering
Full Playwright execution for calendars

Availability calendars load dynamically via XHR. We execute full Playwright browser sessions to intercept JSON payloads directly, capturing data that standard HTTP clients miss.

Regional pricing
Localised IP targeting

Prices change based on user IP and locale. Our proxy pools target specific geographic markets (US, UK, AU) to capture localised pricing and promotions accurately.

Schema stability
Resilient selectors

Travel booking layouts update seasonally. We use fallback chains for selectors to ensure itinerary extraction does not break when Globus updates its front-end code.

Change detection
Only re-scrape what changes

For large catalogues, we maintain a hash index of last-seen values. Subsequent runs only push diffs for availability and price drops, reducing compute cost and downstream processing load.

Applications

Who uses Globus data and how

Teams across industries use globusjourneys.com data to build competitive products and smarter operations.

01
Competitor Price Benchmarking

OTA and tour operators track Globus pricing and promotions to adjust their own package rates.

02
Inventory Aggregation

Travel agencies ingest real-time departure availability into their internal booking portals.

03
Market Research

Analysts monitor new itinerary launches and destination trends to identify shifts in consumer travel demand.

04
AI Travel Planners

ML teams use structured daily itineraries to train generative travel recommendation engines.

05
Yield Management

Revenue teams correlate guaranteed departure status with pricing adjustments to model tour operator yield.

06
Accommodation Sourcing

B2B hotel platforms analyse which properties Globus contracts to target similar operators.

Why DataFlirt

"Globus holds decades of structured itinerary data and pricing models. Extracting it requires navigating dynamic booking engines and regional session states."

Most teams underestimate the investment required: reliable travel scraping requires residential proxies, full JavaScript rendering for availability calendars, and daily selector maintenance. DataFlirt absorbs that complexity so your engineers can focus on the analysis, not the infrastructure.

Technical Spec

Globus scraper technical capabilities

Everything supported by our globusjourneys.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions for dynamic availability calendars
Supported
Dynamic pricing extraction
Captures base rates, supplements, and localized pricing
Supported
Regional proxy rotation
ISP-grade residential IPs for US/UK/AU market pricing
Supported
Cross-brand linking
Identifies Cosmos and Avalon Waterways equivalents
Supported
Day-by-day itinerary parsing
Extracts activities, meals, and accommodations per day
Supported
Change detection (diffs)
Hash-based diff for price and availability updates
Supported
Webhook delivery
HTTP POST per record for real-time inventory updates
Supported
Travel Agent Portal data
Gated B2B commission rates and agent-only inventory
Partial
MyGlobus account itineraries
Post-booking personalised documents and flight details
Partial
Infrastructure

Infrastructure powering the Globus pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration. Playwright handles JavaScript rendering for booking engines. Combined via scrapy-playwright middleware.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across US/UK/AU regions to capture localised travel pricing and avoid rate limits.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling for daily availability checks. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested arrays
CSV
Flat file with typed columns
XLS
Excel compatible format for business teams
Parquet
Columnar format for data warehouses
AWS S3
Direct bucket delivery
Webhook
HTTP POST per record for real-time updates
API
REST endpoints for on-demand querying
BigQuery
Streamed directly into your dataset
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About globusjourneys.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Globus legal?

Scraping publicly available tour information is generally permissible under applicable law. DataFlirt targets only public, non-authenticated itinerary, pricing, and availability data. We do not extract personal data or circumvent agent authentication walls.

How do you handle the dynamic booking engine?

We use full Playwright browser sessions with sticky cookie management to maintain state across the booking flow, ensuring accurate extraction of date-specific pricing.

Can you extract pricing for different countries?

Yes, our residential proxy pools allow us to target specific locales, capturing regional pricing variations and local market promotions.

How fresh is the availability data?

Daily pipeline runs capture exact departure availability and guaranteed status within a 4-6 hour window across the entire catalogue.

Do you extract sister brands like Cosmos?

Yes, we can target Cosmos and Avalon Waterways using similar pipeline architectures, linking equivalent tours where applicable.

What is the minimum viable engagement?

Our smallest packages start at a defined region or tour style with weekly delivery. Contact us with your specific data requirements for a custom quote.

$ dataflirt scope --new-project --source=globusjourneys.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off itinerary dump or a continuous availability feed across 3,000 tours, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in tourism and travel guides

Services

Data Extraction for Every Industry

View All Services →