Skip to content
Abdullah Al Rafi

Bangla birth-certificate OCR at bKash

A first-of-its-kind multilingual OCR for Bangladeshi birth certificates, served on NVIDIA Triton for customer onboarding at bKash.

Summary

Role
ML Engineer, then Senior ML Engineer (Data Science)
Timeline
Apr 2023 – Oct 2025
Stack
Python · ONNX Runtime · NVIDIA Triton
Links
Private (work project)

3×

weekly successful onboarding, at most

vs. before the new document types were supported

+8%

OCR performance

vs. the previous production model

100k+

requests per week

served on Triton with ONNX Runtime

Work project. Details are generalised, with no proprietary code, data or screenshots.

On this page
  1. Summary
  2. Problem
  3. Constraints
  4. Approach and architecture
  5. Results
  6. What I'd do differently

Problem

Customer onboarding at bKash depends on reading identity documents accurately. Bangladeshi birth certificates mix Bangla and English text and come in more than one layout, and no off-the-shelf OCR read them reliably. A certificate the system can't read is an onboarding that doesn't complete.

Constraints

  • Two scripts on one page. Bangla and Latin text sit side by side, so detection and recognition had to handle both.
  • Layout variety. Certificates exist in several formats. Supporting new ones later turned out to be the biggest single gain.
  • Production scale. The service sits inside a live onboarding flow and handles 100k+ requests a week.
  • Sensitive data. Birth certificates are personal identity documents. This page uses no real samples or screenshots, and the diagram is redrawn and generalised.

Approach and architecture

  1. Certificate image

    input

    Captured during customer onboarding

  2. NVIDIA Triton · ONNX Runtime
    1. Pre-processing

      stage 1

      Orientation, crop and contrast

    2. Text detection

      stage 2

      Locate text lines and fields

    3. Recognition

      stage 3

      Bangla and English text, side by side

  3. Field extraction & validation

    stage 4

    Map text to fields and check formats

  4. Onboarding service

    output

    Structured fields for customer verification

Fig. 1 · Generalised OCR pipeline, redrawn. Stage boundaries are illustrative; production details are confidential.
Diagram description

A certificate image from the onboarding flow goes through pre-processing, text detection and Bangla/English recognition. These stages run as ONNX models on NVIDIA Triton. The recognised text is mapped to structured fields and validated before it reaches the onboarding service.

Decision

Build a dedicated multilingual pipeline

Detection, recognition and field extraction were designed around the certificate itself: two scripts, fixed fields, and a known set of layouts.

Considered instead

  • A general-purpose OCR engineNothing available read this document's mix of Bangla and English reliably, which is why the result was a first of its kind.

Decision

Export the models to ONNX and serve them on NVIDIA Triton

Triton runs ONNX Runtime models behind a standard inference API, with dynamic batching and model versioning built in. That is what let the OCR scale to 100k+ requests a week.

Considered instead

  • Running the models inside the application processTies model execution to the web layer and leaves batching, versioning and scaling as custom code.

Decision

Grow coverage by adding document types

As Senior ML Engineer I extended the system to new document types. That raised OCR performance by 8% and is what moved weekly successful onboarding up to 3×.

Considered instead

  • Tuning the existing model on the original layouts onlyCertificates in unsupported layouts failed however well the known layouts were read.

Results

3×

weekly successful customer onboarding, at most

Baseline: weekly successful onboardings before the new document types

+8%

OCR performance after adding new document types

Baseline: the previous production model

100k+

requests per week in production

Served on NVIDIA Triton with ONNX Runtime

The first version made birth-certificate onboarding work end to end. The 2025 extension to new document types raised OCR performance by 8% over the previous production model and lifted weekly successful customer onboarding up to 3×.

What I'd do differently

  • Build an evaluation set per document type from day one, so every new layout gets a before-and-after number instead of being discovered in production.
  • Report field-level accuracy next to onboarding success, so a regression in one field shows up before it shows up in business numbers.