Bangla birth-certificate OCR at bKash
A first-of-its-kind multilingual OCR for Bangladeshi birth certificates, served on NVIDIA Triton for customer onboarding at bKash.
Summary
- Role
- ML Engineer, then Senior ML Engineer (Data Science)
- Timeline
- Apr 2023 – Oct 2025
- Stack
- Python · ONNX Runtime · NVIDIA Triton
- Links
- Private (work project)
3×
weekly successful onboarding, at most
vs. before the new document types were supported
+8%
OCR performance
vs. the previous production model
100k+
requests per week
served on Triton with ONNX Runtime
Work project. Details are generalised, with no proprietary code, data or screenshots.
Problem
Customer onboarding at bKash depends on reading identity documents accurately. Bangladeshi birth certificates mix Bangla and English text and come in more than one layout, and no off-the-shelf OCR read them reliably. A certificate the system can't read is an onboarding that doesn't complete.
Constraints
- Two scripts on one page. Bangla and Latin text sit side by side, so detection and recognition had to handle both.
- Layout variety. Certificates exist in several formats. Supporting new ones later turned out to be the biggest single gain.
- Production scale. The service sits inside a live onboarding flow and handles 100k+ requests a week.
- Sensitive data. Birth certificates are personal identity documents. This page uses no real samples or screenshots, and the diagram is redrawn and generalised.
Approach and architecture
Certificate image
input
Captured during customer onboarding
- NVIDIA Triton · ONNX Runtime
Pre-processing
stage 1
Orientation, crop and contrast
Text detection
stage 2
Locate text lines and fields
Recognition
stage 3
Bangla and English text, side by side
Field extraction & validation
stage 4
Map text to fields and check formats
Onboarding service
output
Structured fields for customer verification
Diagram description
A certificate image from the onboarding flow goes through pre-processing, text detection and Bangla/English recognition. These stages run as ONNX models on NVIDIA Triton. The recognised text is mapped to structured fields and validated before it reaches the onboarding service.
Decision
Build a dedicated multilingual pipeline
Detection, recognition and field extraction were designed around the certificate itself: two scripts, fixed fields, and a known set of layouts.
Considered instead
- A general-purpose OCR engineNothing available read this document's mix of Bangla and English reliably, which is why the result was a first of its kind.
Decision
Export the models to ONNX and serve them on NVIDIA Triton
Triton runs ONNX Runtime models behind a standard inference API, with dynamic batching and model versioning built in. That is what let the OCR scale to 100k+ requests a week.
Considered instead
- Running the models inside the application processTies model execution to the web layer and leaves batching, versioning and scaling as custom code.
Decision
Grow coverage by adding document types
As Senior ML Engineer I extended the system to new document types. That raised OCR performance by 8% and is what moved weekly successful onboarding up to 3×.
Considered instead
- Tuning the existing model on the original layouts onlyCertificates in unsupported layouts failed however well the known layouts were read.
Results
3×
weekly successful customer onboarding, at most
Baseline: weekly successful onboardings before the new document types
+8%
OCR performance after adding new document types
Baseline: the previous production model
100k+
requests per week in production
Served on NVIDIA Triton with ONNX Runtime
The first version made birth-certificate onboarding work end to end. The 2025 extension to new document types raised OCR performance by 8% over the previous production model and lifted weekly successful customer onboarding up to 3×.
What I'd do differently
- Build an evaluation set per document type from day one, so every new layout gets a before-and-after number instead of being discovered in production.
- Report field-level accuracy next to onboarding success, so a regression in one field shows up before it shows up in business numbers.