ocr api
Build and Push Docker Images / build-and-push (push) Failing after 2m5s

This commit is contained in:
2025-07-17 16:00:42 +02:00
parent 7860aa3579
commit fed33b7623
9 changed files with 893 additions and 2 deletions
+207
View File
@@ -0,0 +1,207 @@
# OCR API
A FastAPI-based OCR (Optical Character Recognition) service that can extract text from PDF files and images using Tesseract.
## Features
- **PDF Processing**: Convert PDF files to images and extract text using OCR
- **Image Processing**: Direct OCR processing of image files (PNG, JPG, JPEG, BMP, TIFF)
- **Multi-language Support**: Supports English, French, German, Spanish, and Italian
- **Confidence Scoring**: Returns confidence scores for OCR results
- **RESTful API**: Clean, documented API endpoints
- **Docker Support**: Easy deployment with Docker and Docker Compose
- **File Validation**: Built-in file type and size validation
- **Background Processing**: Efficient file cleanup and processing
## Quick Start
### Using Docker (Recommended)
1. **Clone and build the container:**
```bash
docker-compose up --build
```
2. **Access the API:**
- API Documentation: http://localhost:8000/docs
- Health Check: http://localhost:8000/health
### Local Development
1. **Install system dependencies:**
```bash
# Ubuntu/Debian
sudo apt-get update
sudo apt-get install tesseract-ocr tesseract-ocr-eng poppler-utils libpoppler-cpp-dev pkg-config
# macOS
brew install tesseract poppler
# Arch Linux
sudo pacman -S tesseract tesseract-data-eng poppler
```
2. **Install Python dependencies:**
```bash
pip install -r requirements.txt
```
3. **Run the application:**
```bash
python main.py
```
## API Endpoints
### Health Check
- **GET** `/health` - Check if the service is running
### PDF OCR
- **POST** `/ocr/pdf` - Extract text from PDF files
- **File**: PDF file (max 50MB)
- **Response**: Extracted text with metadata
### Image OCR
- **POST** `/ocr/image` - Extract text from image files
- **File**: Image file (PNG, JPG, JPEG, BMP, TIFF)
- **Response**: Extracted text with metadata
## Usage Examples
### Using curl
**Process a PDF file:**
```bash
curl -X POST "http://localhost:8000/ocr/pdf" \
-H "accept: application/json" \
-H "Content-Type: multipart/form-data" \
-F "file=@document.pdf"
```
**Process an image file:**
```bash
curl -X POST "http://localhost:8000/ocr/image" \
-H "accept: application/json" \
-H "Content-Type: multipart/form-data" \
-F "file=@image.png"
```
### Using Python
```python
import requests
# Process PDF
with open('document.pdf', 'rb') as f:
files = {'file': f}
response = requests.post('http://localhost:8000/ocr/pdf', files=files)
result = response.json()
print(f"Extracted text: {result['text']}")
print(f"Processing time: {result['processing_time']}s")
print(f"Confidence: {result['confidence']}%")
```
### Using JavaScript/Fetch
```javascript
// Process PDF
const formData = new FormData();
formData.append("file", fileInput.files[0]);
fetch("http://localhost:8000/ocr/pdf", {
method: "POST",
body: formData,
})
.then((response) => response.json())
.then((data) => {
console.log("Extracted text:", data.text);
console.log("Processing time:", data.processing_time);
console.log("Confidence:", data.confidence);
});
```
## Response Format
```json
{
"text": "Extracted text content...",
"pages": 3,
"processing_time": 2.45,
"confidence": 85.2
}
```
## Configuration
### Environment Variables
- `PYTHONUNBUFFERED=1` - Ensures Python output is not buffered
- `PYTHONDONTWRITEBYTECODE=1` - Prevents Python from writing .pyc files
### Tesseract Configuration
The Docker image includes multiple language packs:
- English (eng)
- French (fra)
- German (deu)
- Spanish (spa)
- Italian (ita)
To add more languages, modify the Dockerfile and add additional `tesseract-ocr-*` packages.
## Performance Tips
1. **Image Quality**: Higher DPI (300) provides better OCR accuracy but slower processing
2. **File Size**: Larger files take longer to process
3. **Text Quality**: Clear, high-contrast text yields better results
4. **Language**: Specify the correct language for better accuracy
## Troubleshooting
### Common Issues
1. **Tesseract not found**: Ensure Tesseract is installed on your system
2. **PDF processing errors**: Check if Poppler utilities are installed
3. **Memory issues**: Reduce DPI or process smaller files
4. **Poor OCR quality**: Ensure images are clear and have good contrast
### Logs
The application logs processing information. Check Docker logs:
```bash
docker-compose logs ocr-api
```
## Development
### Project Structure
```
ocr-api/
├── main.py # FastAPI application
├── requirements.txt # Python dependencies
├── Dockerfile # Docker configuration
├── docker-compose.yml # Docker Compose configuration
└── README.md # This file
```
### Adding Features
1. **New OCR engines**: Modify the `process_pdf_with_ocr` function
2. **Additional file formats**: Add new endpoints and processing logic
3. **Custom preprocessing**: Add image preprocessing steps before OCR
4. **Batch processing**: Implement queue-based processing for multiple files
## License
This project is open source and available under the MIT License.