Python Web Scraper & Data Collection System

Employer
Web
Project parameters
Type of cooperationOne-time project
SectionSoftware development
Prepaymentwithout prepayment
Payment methodsCash, Bank transfer
Acceptance of requestsfrom yesterday, 22:30 until Aug 31, 2026
Project description
We need a production-grade web scraping system to collect structured data from 5-8 target websites on a daily schedule.
1. Scraper Architecture
— Scrapy framework with rotating proxies and user-agent pool
— Playwright/Selenium integration for JS-heavy pages
— CAPTCHA-bypass layer (2captcha or anti-captcha API)
— Rate limiting and polite crawling (robots.txt compliance)
2. Data Pipeline
— Raw HTML → structured JSON extraction
— Data validation and deduplication logic
— Normalization of prices, dates, phone numbers
— PostgreSQL storage with schema design
— Incremental updates (only new/changed records)
3. Monitoring & Ops
— Health dashboard: success rate, items per run, errors
— Slack/email alerts on failure or data anomalies
— Docker Compose deployment
— Cron-based scheduling with retry logic
Target Sites
— Confidential (NDA required) — details shared after agreement
— Mix of static HTML and dynamic SPA sites
Deliverables
— Full source code with tests
— Docker Compose production config
— Schema migrations
— Runbook and monitoring setup
1. Scraper Architecture
— Scrapy framework with rotating proxies and user-agent pool
— Playwright/Selenium integration for JS-heavy pages
— CAPTCHA-bypass layer (2captcha or anti-captcha API)
— Rate limiting and polite crawling (robots.txt compliance)
2. Data Pipeline
— Raw HTML → structured JSON extraction
— Data validation and deduplication logic
— Normalization of prices, dates, phone numbers
— PostgreSQL storage with schema design
— Incremental updates (only new/changed records)
3. Monitoring & Ops
— Health dashboard: success rate, items per run, errors
— Slack/email alerts on failure or data anomalies
— Docker Compose deployment
— Cron-based scheduling with retry logic
Target Sites
— Confidential (NDA required) — details shared after agreement
— Mix of static HTML and dynamic SPA sites
Deliverables
— Full source code with tests
— Docker Compose production config
— Schema migrations
— Runbook and monitoring setup