DocLens reads bank statements, KYC documents, GST invoices, income tax returns and payslips, turns them into structured and auditable data, derives the credit numbers an underwriter uses, and flags document tampering. It runs on your own infrastructure, with no per page fee and no document leaving your network.
Capabilities
Every value carries its page, the text it came from, the method that produced it and a confidence score. Nothing returns a bare number, so an underwriter or an auditor can always ask where a figure came from.
Table and text layouts parsed, every row reconciled against the running balance. Salary detected by regularity, obligations and EMIs identified, bounces and charges separated, average balances and FOIR inputs derived.
PAN, Aadhaar, driving licence, passport, voter ID and Udyam. Structure and checksums validated on capture, including the Aadhaar Verhoeff digit and the GSTIN base thirty six check.
Tax invoices reconciled arithmetically, ITR, Form 26AS and GST returns mapped to the fields a credit team actually uses. Statutory schemas no open tool covers for India.
Assessed income, monthly obligations, FOIR, bounce conduct, balance behaviour, category mix and circular transaction flags, all computed from reconciled transactions rather than estimated.
Editor and metadata forensics, incremental save detection, font substitution in numerals, running balance continuity, duplicate rows, round amount clustering and leading digit analysis. Each finding carries a plain reason.
Confidence scoring routes each file to straight through, review or refer. Low confidence fields are listed with the reason, ready for a human to confirm before the file proceeds.
How it works
The same path runs for every document, and every stage is inspectable.
PDF and image to pages. OCR runs only where a page needs it.
Weighted evidence, in one readable file, not an opaque model.
One module per document family, each field with its provenance.
Aadhaar, PAN, GSTIN, IFSC, CIN and Udyam checked by arithmetic.
Income, obligations, FOIR and conduct from reconciled rows.
Ten authenticity checks produce a risk score with reasons.
Straight through, review or refer, with the reason attached.
Why DocLens
The hard part of this problem is not the optical character recognition. It is the layer above it: Indian bank layouts, credit derivations an underwriter can defend, statutory schemas, and forensics tuned to the documents that arrive in a real loan file. That layer is what DocLens is.
Self hosted by design. Documents are processed inside your network. This posture fits the DPDP Act and RBI data localisation expectations.
The core has no metered cost. Process one file or a hundred thousand for the same running cost.
Every figure traces to a page, a snippet and a method. A model risk committee can read the classification rules.
PAN, Aadhaar, GSTIN, IFSC, ITR, Form 26AS and GST returns are first class, not an afterthought bolted onto a global template.
The full Aadhaar number is discarded at capture. Only the last four digits are retained, because a lending file does not need more.
One endpoint returns the full structured result. It sits behind your own gateway and feeds your LOS or LMS.
Security and data
This preview is access controlled. In a production deployment the service runs entirely within your environment.
Entry needs a code. The page and the API both refuse requests without it.
This preview processes uploads in memory and does not retain them beyond the working result.
Run the same engine inside your own network so no document leaves your control.
This is a private preview of DocLens. Enter the access code you were given to open the workspace and process a document.