Skip to content

Python Web Scraper & Data Collection System

Content
Employer

Content

> 10 projects
Project parameters
Type of cooperationOne-time project
Prepaymentwithout prepayment
Payment methodsCash, Bank transfer
Acceptance of requestsfrom until Aug 31, 2026
Project description
What We Need

We need a production-grade web scraping system to collect structured data from 5-8 target websites on a daily schedule.

Scope of Work

1. Scraper Architecture
— Scrapy framework with rotating proxies and user-agent pool
— Playwright/Selenium integration for JS-heavy pages
— CAPTCHA-bypass layer (2captcha or anti-captcha API)
— Rate limiting and polite crawling (robots.txt compliance)

2. Data Pipeline
— Raw HTML → structured JSON extraction
— Data validation and deduplication logic
— Normalization of prices, dates, phone numbers
— PostgreSQL storage with schema design
— Incremental updates (only new/changed records)

3. Monitoring & Ops
— Health dashboard: success rate, items per run, errors
— Slack/email alerts on failure or data anomalies
— Docker Compose deployment
— Cron-based scheduling with retry logic

Target Sites
— Confidential (NDA required) — details shared after agreement
— Mix of static HTML and dynamic SPA sites

Deliverables
— Full source code with tests
— Docker Compose production config
— Schema migrations
— Runbook and monitoring setup

Budget: $550 fixed. Timeline: 21 days.
Project author: Content