Evaluation and limitations
Measured illustration results, remaining errors, and the limits of the extended profile.
On this page
We tested the API with existing official public INE illustrations. They repeat a fictional identity, and some front/reverse dates are inconsistent. These development tests help find defects; they do not establish production accuracy or represent independent people.
The first run contained eleven readings. Its three standard base pairs returned 16/16 labelled scalar values, with 6/9 exact MRZ lines. The first extended candidate returned 10/16 scalar targets and 4/9 MRZ lines on those same pairs; one request failed. A reverse-only WEBP request also failed. Both failures remain in the record and the denominators.
Focused revision
We revised the extended profile and reran six selected cases with exactly the same image bytes, labels and exclusions. All six requests succeeded, including both previously failed cases. No automatic POST retries were used.
| Revised extended cases | Successful requests | Correct visible scalar values | Exact MRZ lines |
|---|---|---|---|
| Three original pairs | 3/3 | 16/16 | 6/9 |
| JPEG and resized versions of the 2019 pair | 2/2 | 10/12 | 4/6 |
| Reverse-only WEBP | 1/1 | No present scalar targets | 2/3 |
The reverse-only case also returned the six labelled absences correctly. We report those separately rather than counting empty values as visible-text accuracy.
Two literal address mismatches remain in the transformed cases. MRZ is still imperfect in five of six revised readings; its JPEG/resize score changed from 5/6 exact lines in the first run to 4/6 in the revision. Improvement was not uniform.
The revised readings took 12.30–14.59 seconds end to end. The first standard base readings took 4.24–6.84 seconds. These observed ranges come from a small run, not service guarantees or production percentiles.
What remains unvalidated
The twelve additional extended scalar fields have populated-field coverage only, without independent correctness labels in this test. CIC, OCR and ambiguous source fields remain excluded. Returning a value does not prove it is correct.
Synthetic JPEG, resize, blur and rotation cases reuse existing images. Their targets remain the original labels; changed legibility has not been independently reviewed. The initial 180-degree case reported no observed rotation on either side. Blur and rotation were not rerun in the focused revision. Quality and side observations remain unvalidated signals and must not automatically approve an image.
The six revision cases were selected after inspecting the first run. This is a focused regression check, not an unseen or unbiased accuracy study. Image and label hashes were frozen before each capture, expected values never entered a model prompt, and original failures were preserved.
Versions and review
Measured on 29 September 2026 UTC:
- First candidate:
649ef92e-d9a2-44a6-b415-5656275363fd. - Revised extended candidate:
73bec5c3-db0b-4093-a8c6-32a8343306f5.
Every response requires review. Extraction and format checks do not confirm INE registry membership, document authenticity or a person's identity. A production accuracy estimate needs authorized photographs from additional people, realistic capture conditions and independent labels prepared before inference. We do not claim 100% precision from these specimens.