# Distributed Web Scraper with Anti-Detection

A distributed web scraping system that mimics human behavior to avoid detection. The system consists of a central Flask server that coordinates multiple remote worker nodes, each running their own scraping service.

## Architecture

- **Central Server**: Routes scraping requests to available workers
- **Worker Nodes**: Independent services that perform the actual scraping
- **Communication**: HTTP-based API with authentication
- **Data Flow**: Direct response from worker to client via server proxy

## Features

- Distributed scraping across multiple machines
- Anti-detection measures:
  - Random user agents
  - Variable viewport sizes
  - Random scrolling behavior
  - Delays between requests
  - Browser fingerprint masking
- Content summarization using Sumy
- RESTful API for job submission
- Health checking of worker nodes
- Secure communication with API keys

## Setup

### Central Server

1. Install server dependencies:
```bash
pip install -r requirements-server.txt
```

2. Configure workers in config.json:
```json
{
    "server": {
        "host": "0.0.0.0",
        "port": 5000
    },
    "workers": [
        {
            "id": "worker1",
            "endpoint": "http://worker1.example.com:5001",
            "api_key": "worker1_secret_key"
        }
    ]
}
```

3. Start the server:
```bash
python server.py
```

### Worker Nodes

1. Install worker dependencies on each worker machine:
```bash
pip install -r requirements-worker.txt
```

2. Create worker-specific configuration (worker_config.json):
```json
{
    "id": "worker1",
    "host": "0.0.0.0",
    "port": 5001,
    "api_key": "worker1_secret_key",
    "user_agent": "Mozilla/5.0...",
    "viewport": {"width": 1920, "height": 1080},
    "scroll_behavior": {
        "min_scrolls": 2,
        "max_scrolls": 5,
        "scroll_delay_min": 1000,
        "scroll_delay_max": 3000
    }
}
```

3. Start the worker:
```bash
python worker.py
```

4. Download NLTK data (required for summarization):
```python
import nltk
nltk.download('punkt')
```

## Usage

1. Submit URL for scraping and receive result:
```bash
curl -X POST http://localhost:5000/scrape \
  -H "Content-Type: application/json" \
  -d '{"url": "https://example.com"}'
```

2. Check system status:
```bash
curl http://localhost:5000/status
```

## API Endpoints

### Worker Endpoints

- GET `/health`: Check worker health status
  - Requires: X-API-Key header
  - Returns: Worker health information

- POST `/scrape`: Submit URL for scraping
  - Requires: X-API-Key header
  - Request body: `{"url": "https://example.com"}`
  - Returns: Scraped data or error information

### Central Server Endpoints

- POST `/scrape`: Submit URL for scraping and receive result
  - Request body: `{"url": "https://example.com"}`
  - Returns: Scraped data or error information
  ```json
  {
    "status": "success",
    "data": {
      "url": "https://example.com",
      "content": "Markdown formatted content...",
      "summary": "Summarized content...",
      "timestamp": "2023-..."
    }
  }
  ```

- GET `/status`: Get system status
  - Returns: Worker status and health information

## Result Format

Each scraping response contains:
- URL: The scraped page URL
- Content: Full page content in markdown format
- Summary: Text summary generated by Sumy
- Timestamp: When the page was scraped

## Anti-Detection Features

The scraper implements several measures to avoid detection:
1. Distributed workload across multiple physical machines
2. Random scrolling behavior with variable speeds and distances
3. Different user agents and viewport sizes per worker
4. Delays between requests with random intervals
5. Browser fingerprint masking

## Security

- Worker authentication via API keys
- HTTPS recommended for production deployment
- Worker health monitoring
- Automatic worker failover
