Python Web Scraper & Data Collection System

Employer
Digital
Project parameters
Type of cooperationOne-time project
SectionSoftware development
Prepaymentwithout prepayment
Payment methodsCash, Bank transfer
Acceptance of requestsfrom yesterday, 21:46 until Aug 30, 2026
Project description
What We Need
We need a production-grade web scraping system to collect structured data from 5-8 target websites on a daily schedule.
Scope of Work
1. Scraper Architecture
— Scrapy framework with rotating proxies and user-agent pool
— Playwright/Selenium integration for JS-heavy pages
— CAPTCHA-bypass layer (2captcha or anti-captcha API)
— Rate limiting and polite crawling (robots.txt compliance)
2. Data Pipeline
— Raw HTML → structured JSON extraction
— Data validation and deduplication logic
— Normalization of prices, dates, phone numbers
— PostgreSQL storage with schema design
— Incremental updates (only new/changed records)
3. Monitoring & Ops
— Health dashboard: success rate, items per run, errors
— Slack/email alerts on failure or data anomalies
— Docker Compose deployment
— Cron-based scheduling with retry logic
Target Sites
— Confidential (NDA required) — details shared after agreement
— Mix of static HTML and dynamic SPA sites
Deliverables
— Full source code with tests
— Docker Compose production config
— Schema migrations
— Runbook and monitoring setup
Budget: $550 fixed. Timeline: 21 days.
We need a production-grade web scraping system to collect structured data from 5-8 target websites on a daily schedule.
Scope of Work
1. Scraper Architecture
— Scrapy framework with rotating proxies and user-agent pool
— Playwright/Selenium integration for JS-heavy pages
— CAPTCHA-bypass layer (2captcha or anti-captcha API)
— Rate limiting and polite crawling (robots.txt compliance)
2. Data Pipeline
— Raw HTML → structured JSON extraction
— Data validation and deduplication logic
— Normalization of prices, dates, phone numbers
— PostgreSQL storage with schema design
— Incremental updates (only new/changed records)
3. Monitoring & Ops
— Health dashboard: success rate, items per run, errors
— Slack/email alerts on failure or data anomalies
— Docker Compose deployment
— Cron-based scheduling with retry logic
Target Sites
— Confidential (NDA required) — details shared after agreement
— Mix of static HTML and dynamic SPA sites
Deliverables
— Full source code with tests
— Docker Compose production config
— Schema migrations
— Runbook and monitoring setup
Budget: $550 fixed. Timeline: 21 days.