197 lines
5.9 KiB
Markdown
197 lines
5.9 KiB
Markdown
# Playwright API Server
|
|
|
|
A FastAPI-based server that uses Playwright to scrape websites and extract various types of data.
|
|
|
|
## Recent Migration (v3.0)
|
|
|
|
### Migration from Puppeteer to Playwright
|
|
|
|
This project has been successfully migrated from Puppeteer (pyppeteer) to Playwright for better performance, reliability, and maintainability.
|
|
|
|
#### Key Benefits of Playwright
|
|
|
|
- **Better performance**: More efficient browser management and faster page loads
|
|
- **Improved stability**: Better handling of modern web applications
|
|
- **Enhanced features**: Better support for modern web standards
|
|
- **Active development**: More frequent updates and better community support
|
|
- **Multi-browser support**: Can easily switch between Chromium, Firefox, and WebKit
|
|
|
|
#### Migration Changes
|
|
|
|
- **Dependencies**: Updated from `pyppeteer` to `playwright`
|
|
- **Browser management**: Improved browser pool with better resource management
|
|
- **API compatibility**: All existing endpoints remain the same
|
|
- **Docker image**: Updated to use Playwright browsers instead of Chrome
|
|
|
|
## Recent Fixes (v2.0)
|
|
|
|
### Issues Fixed
|
|
|
|
- **Resource temporarily unavailable (Errno 11)** - Fixed by implementing proper browser cleanup and resource limits
|
|
- **100% CPU usage** - Fixed by reducing browser pool size, adding operation limits, and improving cleanup
|
|
- **Server hanging** - Fixed by adding timeouts, emergency cleanup, and better error handling
|
|
|
|
### Key Improvements
|
|
|
|
#### 1. Resource Management
|
|
|
|
- **Reduced browser pool size**: From 5 to 3 browsers maximum
|
|
- **Shorter browser TTL**: From 1 hour to 30 minutes
|
|
- **Concurrent operation limits**: Maximum 5 concurrent operations
|
|
- **Queue size limits**: Prevent unbounded growth
|
|
|
|
#### 2. Better Cleanup
|
|
|
|
- **Automatic browser recycling**: Browsers are recycled every 30 minutes
|
|
- **Emergency cleanup**: Force cleanup all browsers when needed
|
|
- **Proper page cleanup**: Ensure pages are closed after each operation
|
|
- **Signal handlers**: Graceful shutdown on SIGTERM/SIGINT
|
|
|
|
#### 3. Monitoring & Debugging
|
|
|
|
- **System status endpoint**: `/status` - Monitor CPU, memory, file descriptors
|
|
- **Browser pool monitoring**: Track active browsers and their ages
|
|
- **Emergency cleanup endpoint**: `/emergency-cleanup` - Force cleanup when needed
|
|
- **Monitoring script**: `monitor.py` - Real-time system monitoring
|
|
|
|
#### 4. Error Handling
|
|
|
|
- **Timeout handling**: 30s timeout for getting browsers from pool
|
|
- **Emergency recovery**: Auto-cleanup when timeouts occur
|
|
- **Better exception handling**: More detailed error logging
|
|
|
|
## Environment Variables
|
|
|
|
```bash
|
|
# Required
|
|
API_KEY=your-api-key-here
|
|
|
|
# Optional (with defaults)
|
|
MAX_BROWSERS=3 # Maximum browser instances (default: 3)
|
|
BROWSER_TTL=1800 # Browser time-to-live in seconds (default: 1800 = 30min)
|
|
MAX_CONCURRENT_OPERATIONS=5 # Max concurrent operations (default: 5)
|
|
CACHE_EXPIRY_HOURS=36 # Cache expiry in hours (default: 36)
|
|
CLEANUP_CRON=0 3 * * * # Cache cleanup schedule (default: daily at 3 AM)
|
|
RATE_LIMIT_MINUTE=60 # Requests per minute (default: 60)
|
|
```
|
|
|
|
## API Endpoints
|
|
|
|
### Core Endpoints
|
|
|
|
- `GET /` - Visit URL and get HTML content
|
|
- `GET /seo` - Extract SEO information
|
|
- `GET /meta` - Extract meta tags and Open Graph data
|
|
|
|
### Management Endpoints
|
|
|
|
- `HEAD /` - Health check
|
|
- `GET /status` - System status and browser pool information
|
|
- `POST /emergency-cleanup` - Force cleanup all browsers
|
|
- `GET /cache/stats` - Cache statistics
|
|
- `GET /cache/clear` - Clear all cache
|
|
|
|
## Usage Examples
|
|
|
|
### Basic Usage
|
|
|
|
```bash
|
|
# Visit a URL
|
|
curl -H "X-API-Key: your-api-key" "http://localhost:8000/?url=https://example.com"
|
|
|
|
# Extract SEO data
|
|
curl -H "X-API-Key: your-api-key" "http://localhost:8000/seo?url=https://example.com"
|
|
|
|
# Get system status
|
|
curl -H "X-API-Key: your-api-key" "http://localhost:8000/status"
|
|
```
|
|
|
|
### Monitoring
|
|
|
|
```bash
|
|
# Monitor system every 30 seconds
|
|
python monitor.py your-api-key 30
|
|
|
|
# Monitor with custom URL
|
|
python monitor.py your-api-key 30 http://your-server:8000
|
|
|
|
# Trigger emergency cleanup
|
|
python monitor.py your-api-key 30 http://localhost:8000 cleanup
|
|
```
|
|
|
|
## Docker Usage
|
|
|
|
```bash
|
|
# Build the image
|
|
docker build -t playwright-api .
|
|
|
|
# Run with environment variables
|
|
docker run -d \
|
|
--name playwright-api \
|
|
-p 8000:8000 \
|
|
-e API_KEY=your-api-key \
|
|
-e MAX_BROWSERS=3 \
|
|
-e BROWSER_TTL=1800 \
|
|
-v /path/to/cache:/db \
|
|
playwright-api
|
|
```
|
|
|
|
## Testing the Migration
|
|
|
|
To verify that the Playwright migration works correctly:
|
|
|
|
```bash
|
|
# Run the migration test script
|
|
python test_playwright_migration.py
|
|
```
|
|
|
|
This will test:
|
|
|
|
- Playwright installation and basic functionality
|
|
- Browser utilities module
|
|
- Browser service functionality
|
|
|
|
## Troubleshooting
|
|
|
|
### High CPU Usage
|
|
|
|
1. Check system status: `GET /status`
|
|
2. If browser pool is at capacity, trigger cleanup: `POST /emergency-cleanup`
|
|
3. Consider reducing `MAX_BROWSERS` or `MAX_CONCURRENT_OPERATIONS`
|
|
|
|
### Resource Errors (Errno 11)
|
|
|
|
1. The system now automatically handles this with better cleanup
|
|
2. Monitor with `python monitor.py` to track resource usage
|
|
3. If persistent, trigger emergency cleanup
|
|
|
|
### Server Hanging
|
|
|
|
1. Check for stuck operations with monitoring script
|
|
2. Trigger emergency cleanup
|
|
3. Restart the container if needed
|
|
|
|
### Playwright Installation Issues
|
|
|
|
1. Ensure Playwright browsers are installed: `playwright install chromium`
|
|
2. Check Docker build logs for browser installation
|
|
3. Verify system dependencies are installed
|
|
|
|
## Performance Tips
|
|
|
|
1. **Use caching**: The API caches results for 36 hours by default
|
|
2. **Monitor resources**: Use the monitoring script to track usage
|
|
3. **Adjust limits**: Tune `MAX_BROWSERS` and `MAX_CONCURRENT_OPERATIONS` based on your server capacity
|
|
4. **Regular cleanup**: The system automatically recycles browsers every 30 minutes
|
|
|
|
## Logs
|
|
|
|
The server provides detailed logging for:
|
|
|
|
- Browser creation and cleanup
|
|
- Resource usage warnings
|
|
- Error conditions
|
|
- Cache operations
|
|
|
|
Monitor logs to identify patterns and adjust configuration accordingly.
|