Playwright API Server
A FastAPI-based server that uses Playwright to scrape websites and extract various types of data.
Recent Migration (v3.0)
Migration from Puppeteer to Playwright
This project has been successfully migrated from Puppeteer (pyppeteer) to Playwright for better performance, reliability, and maintainability.
Key Benefits of Playwright
- Better performance: More efficient browser management and faster page loads
- Improved stability: Better handling of modern web applications
- Enhanced features: Better support for modern web standards
- Active development: More frequent updates and better community support
- Multi-browser support: Can easily switch between Chromium, Firefox, and WebKit
Migration Changes
- Dependencies: Updated from
pyppeteertoplaywright - Browser management: Improved browser pool with better resource management
- API compatibility: All existing endpoints remain the same
- Docker image: Updated to use Playwright browsers instead of Chrome
Recent Fixes (v2.0)
Issues Fixed
- Resource temporarily unavailable (Errno 11) - Fixed by implementing proper browser cleanup and resource limits
- 100% CPU usage - Fixed by reducing browser pool size, adding operation limits, and improving cleanup
- Server hanging - Fixed by adding timeouts, emergency cleanup, and better error handling
Key Improvements
1. Resource Management
- Reduced browser pool size: From 5 to 3 browsers maximum
- Shorter browser TTL: From 1 hour to 30 minutes
- Concurrent operation limits: Maximum 5 concurrent operations
- Queue size limits: Prevent unbounded growth
2. Better Cleanup
- Automatic browser recycling: Browsers are recycled every 30 minutes
- Emergency cleanup: Force cleanup all browsers when needed
- Proper page cleanup: Ensure pages are closed after each operation
- Signal handlers: Graceful shutdown on SIGTERM/SIGINT
3. Monitoring & Debugging
- System status endpoint:
/status- Monitor CPU, memory, file descriptors - Browser pool monitoring: Track active browsers and their ages
- Emergency cleanup endpoint:
/emergency-cleanup- Force cleanup when needed - Timeout-based cleanup:
/force-cleanup-old- Force close old browser/page instances - Monitoring script:
monitor.py- Real-time system monitoring
4. Timeout Management
- Configurable timeout:
BROWSER_INSTANCE_TIMEOUT_MINUTES(default: 10 minutes) - Automatic cleanup: Old browser and page instances are automatically force closed
- Page tracking: All page instances are tracked with creation timestamps
- Health check integration: Timeout cleanup runs every 2 minutes as part of health checks
5. Error Handling
- Timeout handling: 30s timeout for getting browsers from pool
- Emergency recovery: Auto-cleanup when timeouts occur
- Better exception handling: More detailed error logging
Environment Variables
# Required
API_KEY=your-api-key-here
# Optional (with defaults)
MAX_BROWSERS=3 # Maximum browser instances (default: 3)
BROWSER_TTL=1800 # Browser time-to-live in seconds (default: 1800 = 30min)
MAX_CONCURRENT_OPERATIONS=5 # Max concurrent operations (default: 5)
BROWSER_INSTANCE_TIMEOUT_MINUTES=10 # Force close old browser/page instances after N minutes (default: 10)
CACHE_EXPIRY_HOURS=36 # Cache expiry in hours (default: 36)
CLEANUP_CRON=0 3 * * * # Cache cleanup schedule (default: daily at 3 AM)
RATE_LIMIT_MINUTE=60 # Requests per minute (default: 60)
API Endpoints
Core Endpoints
GET /- Visit URL and get HTML contentGET /seo- Extract SEO informationGET /meta- Extract meta tags and Open Graph dataGET /resulting-url- Get the final URL after navigation (handles redirects)
Management Endpoints
HEAD /- Health checkGET /status- System status and browser pool informationPOST /emergency-cleanup- Force cleanup all browsersPOST /force-cleanup-old- Force cleanup old browser/page instances (based on timeout)GET /cache/stats- Cache statisticsGET /cache/clear- Clear all cache
Usage Examples
Basic Usage
# Visit a URL
curl -H "X-API-Key: your-api-key" "http://localhost:8000/?url=https://example.com"
# Extract SEO data
curl -H "X-API-Key: your-api-key" "http://localhost:8000/seo?url=https://example.com"
# Get resulting URL (handles redirects)
curl -H "X-API-Key: your-api-key" "http://localhost:8000/resulting-url?url=https://example.com"
# Get system status
curl -H "X-API-Key: your-api-key" "http://localhost:8000/status"
# Force cleanup old browser/page instances
curl -X POST -H "X-API-Key: your-api-key" "http://localhost:8000/force-cleanup-old"
Monitoring
# Monitor system every 30 seconds
python monitor.py your-api-key 30
# Monitor with custom URL
python monitor.py your-api-key 30 http://your-server:8000
# Trigger emergency cleanup
python monitor.py your-api-key 30 http://localhost:8000 cleanup
Docker Usage
# Build the image
docker build -t playwright-api .
# Run with environment variables
docker run -d \
--name playwright-api \
-p 8000:8000 \
-e API_KEY=your-api-key \
-e MAX_BROWSERS=3 \
-e BROWSER_TTL=1800 \
-v /path/to/cache:/db \
playwright-api
Testing the Migration
To verify that the Playwright migration works correctly:
# Run the migration test script
python test_playwright_migration.py
This will test:
- Playwright installation and basic functionality
- Browser utilities module
- Browser service functionality
Troubleshooting
High CPU Usage
- Check system status:
GET /status - If browser pool is at capacity, trigger cleanup:
POST /emergency-cleanup - Force cleanup old instances:
POST /force-cleanup-old - Consider reducing
MAX_BROWSERSorMAX_CONCURRENT_OPERATIONS - Adjust
BROWSER_INSTANCE_TIMEOUT_MINUTESto a lower value (e.g., 5 minutes)
Resource Errors (Errno 11)
- The system now automatically handles this with better cleanup
- Monitor with
python monitor.pyto track resource usage - If persistent, trigger emergency cleanup
Server Hanging
- Check for stuck operations with monitoring script
- Trigger emergency cleanup
- Restart the container if needed
Playwright Installation Issues
- Ensure Playwright browsers are installed:
playwright install chromium - Check Docker build logs for browser installation
- Verify system dependencies are installed
Performance Tips
- Use caching: The API caches results for 36 hours by default
- Monitor resources: Use the monitoring script to track usage
- Adjust limits: Tune
MAX_BROWSERSandMAX_CONCURRENT_OPERATIONSbased on your server capacity - Regular cleanup: The system automatically recycles browsers every 30 minutes
Logs
The server provides detailed logging for:
- Browser creation and cleanup
- Resource usage warnings
- Error conditions
- Cache operations
Monitor logs to identify patterns and adjust configuration accordingly.