ArchiveLMGet beta access
Back to ArchiveLMBeta

OCR API for historical documents

Send one page image, get back its text exactly as written. The same models that read the documents in ArchiveLM, without the app around them. Built for the pages general-purpose OCR gets wrong: old type, handwriting, and languages other than English.

The API is in beta. It's open to approved ArchiveLM beta accounts: request access, then create a key in your dashboard. Endpoints and prices may change before launch. Beta customers keep their rate for 12 months after launch.

Pricing

Free

$0
100 pages / month

Self-serve key, no card.

Pay-as-you-go

$0.025
per page

$25 per 1,000 pages after your free 100. Add a card in your dashboard; billed monthly.

Volume

from $0.02
per page

20,000+ pages a month. Quoted.

You're billed for successful pages only. A page we fail to read returns 502 ocr_failed and isn't billed, though it does count toward your monthly page limit.

What it's good at, and what it isn't

  • Letters, manuscripts and handwritten records
  • Books, pamphlets, legal and parliamentary records
  • Any language or script. Nothing is translated or modernised
  • Pages where unreadable text must be marked [ILLEGIBLE], not guessed
  • Dense multi-column newspapers — use the ArchiveLM app, which reads them in segments
  • Article structuring, layout, bounding boxes or confidence scores — you get plain text
  • Multi-page PDFs — send one page per call
  • Cheap bulk OCR of modern print — Google Document AI or Mistral OCR are more than 10× cheaper for that, and we'd use them too

Quickstart

1. Sign in and open Dashboard → OCR API. Create a key. It's shown once, so store it. To go past 100 pages a month, add a card on the same page.

2. Send a page.

curl -X POST https://www.archivelm.com/api/v1/ocr \
  -H "X-API-Key: alm_..." \
  -H "Content-Type: application/json" \
  -d '{"image_url": "https://example.org/scans/page-001.jpg", "language_hint": "German"}'

Response:

{
  "request_id": "…",
  "text": "Berliner Abendpost …",
  "chars": 8412,
  "pages": 1,
  "model_tier": "standard",   // "enhanced" when a second, stronger pass was needed
  "retried": false
}

Endpoints

EndpointDoes
POST /ocrRead one page: image_url or image_base64, optional language_hint (a language name, e.g. German)
GET /pingCheck a key; returns your plan, pages used this month and your limit. Free.

Base URL: https://www.archivelm.com/api/v1. Authenticate with the X-API-Key header.

Limits

  • JPEG, PNG, WebP or a single-page PDF. Type is checked from the file, not the name.
  • image_url: HTTPS, up to 10 MB, no redirects. Use it for anything over ~3 MB.
  • image_base64: up to ~3 MB of image (the encoded request must stay under 4.5 MB).
  • Two requests at a time per account. A call can take a few minutes when the second pass runs, so set your client timeout to 300 s.
  • Errors: 400 bad input, 401 bad key, 402 free pages used up (add billing), 413 file too large, 415 not JSON, 429 too many concurrent requests or monthly cap reached, 502 page not read (not billed), 503 busy or daily capacity reached.

Your documents

We keep nothing. The image is read, the text goes back to you, and both are discarded. We log only metadata: when you called, file size, how long it took and whether it worked. Pages are read by Google's Gemini models (see our subprocessors).

Send us your hardest pages

Tell us what you're reading and roughly how many pages. We'll show you where it fails before you build on it.

Get beta access

Questions? hello@archivelm.com

Terms · Privacy · hello@archivelm.com