Skip to content
Mehedi Hasan Sarkar
All case studies

PDF textbooks into exam content with OCR and AI

Learning · AI pipelineNeovotech Ltd, 2024 to 2026

Problem

An AI-assisted learning platform needed its library of printed textbooks as structured content. That means headings, equations, tables, multiple-choice questions and creative questions. All other features use this content, from AI-generated exams to the AI tutor.

Doing this by hand is slow. Doing it with OCR (reading text from page images) and AI is also slow. It also fails in the middle. A page can time out. The AI quota can reach its limit. A worker can stop halfway through a long book. If one failure means starting the whole book again, the pipeline never finishes.

Solution

Book PDFs are large, so they never go through the API server. The browser uploads each file directly to Cloudflare R2 storage with a signed URL.

After the upload, the work continues as background jobs on Redis queues. Every job has a visible status: queued, processing, retrying, succeeded or failed.

  1. Direct upload to storage
  2. Queued job
  3. OCR, page by page
  4. Typed content blocks
  5. Structured with Gemini
  6. Chapters readyOnly the failed chapter regenerated
A failure in one stage does not stop the whole book.

A Python worker downloads the PDF, creates thumbnails and runs OCR page by page. It turns the raw output into typed blocks. These are headings, paragraphs, lists, equations, tables, images, multiple-choice questions and creative questions.

Structured extraction and chapter summaries use Google Gemini with JSON-constrained output. This means Gemini must return data in a fixed JSON shape. So later steps always receive predictable data.

I designed the pipeline to expect failure. Each page has a hard timeout, which is a strict time limit. Jobs retry after calculated delays. When the AI quota is reached, jobs wait longer before they try again. If a worker stops in the middle of a job, the job recovers automatically. A failed chapter can be regenerated by itself, without the rest of the book. Problems are reported through notifications and alerts.

The trade-off. A separate worker behind a queue has a cost. It means two languages, a queue, and more parts to run and deploy than doing OCR inside the API. I accepted this because OCR fails in ways an API should not. A page hangs, the AI quota runs out, or a worker crashes in the middle of a book. All of these problems stay inside the worker. So the API still answers uploads and status checks, whatever the OCR is doing.

Impact

  • One bad page no longer fails a whole book.
  • A crashed worker resumes instead of starting over.
  • Later steps always receive predictable, typed content.
  • Editors can regenerate one chapter instead of the whole book.

Related case studies

Have a similar problem?

Tell me what is breaking or what you need built. I will reply with how I would approach it.

Let's talk