

| Model | OCR-Bench Accuracy (%) | Multilingual Accuracy (%) | Layout / Table Understanding (%) |
|---|---|---|---|
| Next OCR | 99.0 | 96.8 | 95.3 |
| PaddleOCR | 95.2 | 93.9 | 95.3 |
| Deepseek OCR | 90.6 | 87.4 | 86.1 |
| Tesseract | 92.0 | 88.4 | 72.0 |
| EasyOCR | 90.4 | 84.7 | 78.9 |
| Google Cloud Vision / DocAI | 98.7 | 95.5 | 93.6 |
| Amazon Textract | 94.7 | 86.2 | 86.1 |
| Azure Document Intelligence | 95.1 | 93.6 | 91.4 |
| Model | Handwriting (%) | Scene Text (%) | Complex Tables (%) |
|---|---|---|---|
| Next OCR | 92 | 96 | 91 |
| PaddleOCR | 88 | 92 | 90 |
| Deepseek OCR | 80 | 85 | 83 |
| Tesseract | 75 | 88 | 70 |
| EasyOCR | 78 | 86 | 75 |
| Google Cloud Vision / DocAI | 90 | 95 | 92 |
| Amazon Textract | 85 | 90 | 88 |
| Azure Document Intelligence | 87 | 91 | 89 |
1from transformers import AutoTokenizer, AutoModelForVision2Seq
2import torch
3
4model_id = "Lamapi/next-ocr"
5
6tokenizer = AutoTokenizer.from_pretrained(model_id)
7model = AutoModelForVision2Seq.from_pretrained(model_id, torch_dtype=torch.float16)
8
9img = Image.open("image.jpg")
10
11# ATTENTION: The content list must include both an image and text.
12messages = [
13 {"role": "system", "content": "You are Next-OCR, an helpful AI assistant trained by Lamapi."},
14 {
15 "role": "user",
16 "content": [
17 {"type": "image", "image": img},
18 {"type": "text", "text": "Read the text in this image and summarize it."}
19 ]
20 }
21]
22
23# Apply the chat template correctly
24prompt = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
25inputs = processor(text=prompt, images=[img], return_tensors="pt").to(model.device)
26
27with torch.no_grad():
28 generated = model.generate(**inputs, max_new_tokens=256)
29
30print(processor.decode(generated[0], skip_special_tokens=True))| Feature | Description |
|---|---|
| 🖼️ High-Accuracy OCR | Extracts text from images, documents, and screenshots reliably. |
| 🇹🇷 Multilingual Support | Works with 30+ languages including Turkish. |
| ⚡ Lightweight & Efficient | Optimized for resource-constrained environments. |
| 📄 Layout & Math Awareness | Handles tables, forms, and mathematical formulas. |
| 🏢 Reliable Outputs | Suitable for enterprise document workflows. |
| Specification | Details |
|---|---|
| Base Model | Qwen 3 |
| Parameters | 8 Billion |
| Architecture | Vision + Transformer (OCR LLM) |
| Modalities | Image-to-text |
| Fine-Tuning | OCR datasets with multilingual and math/tabular content |
| Optimizations | Quantization-ready, FP16 support |
| Primary Focus | Text extraction, document understanding, mathematical OCR |
Next OCR — Compact OCR + math-capable AI, blending accuracy, speed, and multilingual document intelligence.