সরকারি নথি

Government document OCR for Bangladesh

Citizen applications, land records (পর্চা/খতিয়ান), registrations and legacy files — much of it handwritten Bangla that no foreign OCR can touch. Digitized on your own infrastructure.

Why public-sector digitization stalls

📜

Handwritten legacy records

Decades of registries, mutations and applications written by hand in Bangla — invisible to generic OCR. Our handwriting recognition AI reads them.
🔐

Security requirements

Citizen data cannot go to foreign clouds. Fully offline, air-gapped deployment fits public-sector security policy.
📈

Archive-scale volumes

Lakhs of pages per district. Batch processing on a single GPU server handles thousands of pages daily — scale by adding machines.

What gets digitized

🌾

Land records

পর্চা, খতিয়ান, mutation records — dag and khatian numbers preserved in Bangla numerals, searchable at last.
📝

Citizen applications

Filled application forms extracted field-by-field into structured data for e-governance systems.
📚

Registries & gazettes

Printed registers, gazettes and notices — mixed Bangla-English pages read in one pass.
🗂️

Legacy archives

Bound volumes scanned to PDF, then converted to searchable text page by page.

Start with a pilot batch

Send us a representative sample — we'll digitize it and hand you the accuracy report, so the decision is based on your documents, not our claims.

Frequently Asked Questions

Can it read old land records like পর্চা and খতিয়ান?

Yes — legacy land records are largely handwritten Bangla, which is our core specialty. Document condition matters: clear pages score higher, faded or damaged pages are flagged with lower confidence rather than guessed silently.

Our documents cannot leave the office. What are our options?

Air-gapped on-premise deployment: the system runs on your own hardware with no internet connection at all. Documents, results and even model improvements never leave your infrastructure.

We have a backlog of lakhs of pages. Is that feasible?

Yes — batch processing on a single GPU server handles thousands of pages daily, so archive-scale digitization projects are realistic on modest hardware.

Does it preserve Bangla numerals in dag/khatian numbers?

Yes — ০১২৩৪৫৬৭৮৯ stays Bangla in the output, exactly as written. Nothing is silently converted, which matters for legal record fidelity.

How accurate is it on old documents?

~99% on clearly printed pages and ~88% character accuracy on handwriting, with confidence flags marking anything uncertain so staff verify only the flagged minority.