Measured engineering

Built to be inspected. Measured on named controls.

RSEEN treats source linkage, extraction, formulas, conflict handling and evidence decisions as measurable controls. Each measurement record names the corpus, exact numerator and denominator, build identity, exclusions and failures. Test volume shows regression breadth; real-file performance is measured separately on an agreed corpus.

How the institution's method is encoded
What the intelligence does

Credit intelligence that shows its work.

RSEEN Core's credit intelligence reads Arabic and English credit documents, keeps every material fact bound to the page it came from, and computes the analysis with defined, inspectable formulas; the credit decision is the institution's.

RSEEN's document intelligence was designed for Arabic from the start, not added as a translation layer over an English product: Arabic and English documents are read in the original, through native text where it exists and OCR where the evidence is an image. Extraction accuracy is measured on the institution's own corpus.

Engineering baseline · synthetic / development corpus

11,026 backend tests, counted rather than rounded, at Core revision 5c2ef63cc (27 July 2026): collected regression breadth, not a one-run pass result, and not re-verified at the current Core revision. A recount is scheduled. In the 19 July 2026 development baseline, 30/30 extraction-fixture controls, 101/101 formula-DSL controls and 232/232 conflict-component controls passed; the sets overlap. Pilot qualification is measured separately on the institution's agreed corpus.

30/30 extraction-fixture controls passed
101/101 formula DSL controls passed
232/232 conflict component controls passed
95/95 scale-authority controls passed · units resolved before promotion
70/70 same-process deterministic renderer outputs byte-identical

Shared-responsibility boundary

A security boundary your team can inspect.

Data never leaves Saudi Arabia. Keys stay with the institution. The map shows where each layer runs and who owns the decision.

Platform controls Deployment choices

Institution boundary

Decision and operating authority Identity, authority, keys, data policy and the credit decision

RSEEN platform boundary

RSEEN Core Complete standalone interface for documents, evidence, computation, conflict, review and outputs
Approved Playbook Source hierarchy, calculations, evidence requirements, exceptions and authority model
Workflow modules · planned Planned workflow modules, scoped with each institution

Optional providers and connectors

Documented routes Inference, OCR, messaging, signing and client systems

Built into RSEEN

  • Request bound to case, provenance and execution route
  • Traceable case result or clear stop reason
  • Credit tables computed by formula, with any withheld grade visible
  • RSEEN Core works as a standalone product; further modules are planned and scoped with each institution

Configured for the institution

  • Identity, roles, key custody and storage
  • Provider eligibility, residency, egress and logs
  • Backup, restore and recovery objectives
  • Support, break-glass, retention, exit and connectors

RSEEN keeps an append-only, attributable audit trail of the supported credit workflow—who did what, when, and against which evidence. It is written for its readers: internal audit, the credit risk committee, and the SAMA supervisory examiner following one figure from source to decision.

The map keeps responsibility clear from evidence and computation through workflow and credit decision.

How RSEEN is judged

Seven measures. No blended verdict.

The institution sets thresholds before a locked run. Every measure reports its named corpus, exact numerator and denominator, unevaluable population, required slices, run identity and known limitations. Synthetic and pilot results remain separate; an unsupported or failing slice is never averaged away.

M1

Citation precision and recall

Measures whether each citation resolves to the right source and supports the material fact, and whether every material fact that needs support has it.

M2

Extraction error by document class

Reports correct, wrong, missing, spurious and unevaluable fields separately by language, acquisition mode, document class and field type.

M3

Formula reconciliation

Compares engine and end-to-end outputs with an independently specified oracle, while retaining missing or unavailable inputs in the result.

M4

Conflict coverage

Measures known conflicts caught with both values attributable through reviewer resolution, plus false flags on clean control pairs.

M5

Withhold correctness

Checks both paths: insufficient-evidence cases pause for resolution; sufficient-evidence cases proceed to the independently expected grade.

M6

Deterministic repeatability

Compares rendered bytes across repeated same-process calls within named deterministic renderer controls.

M7

Throughput and latency

Reports queue, extraction, analysis and total time by file class on named hardware and concurrency, including retries and failures.

Qualification uses a locked synthetic or anonymised pilot corpus, institution-approved answer keys and thresholds, and a recorded run identity.

Arabic product surfaces are not a claim about Arabic extraction accuracy; accuracy is measured on the agreed institutional corpus.

Measurement system

From run identity to accountable result.

The harness fixes what is measured, separates evidence lanes, exercises controls and clean twins, repeats named deterministic checks, and keeps every failure in the denominator.

01 Instrument

Name the object before measuring it.

A run pins source hashes, answer-key identity, build, configuration, hardware and corpus membership before outputs are opened. The record also identifies the exact control lane, from extraction and formula fixtures to deterministic renderer and interface checks.

A result without a pinned identity is not scored.

02 Separate

Engineering evidence is not qualification.

Development fixtures test implementation behaviour. Locked synthetic cases measure controlled holdouts. Anonymised pilot files measure the institution's declared document mix. Their results and denominators remain separate.

A development fixture never enters a qualification denominator.

03 Exercise

Test the hard path and its clean twin.

Controlled pairs test a clean source against one changed factor: a conflict, missing evidence, OCR corruption, formula edge case or unsupported citation. Positive and negative paths are counted together.

Coverage and false-positive behaviour are measured together.

04 Repeat

Fixed renderer inputs must remain byte-identical.

The deterministic renderer baseline produced 70/70 byte-identical outputs across 7 named controls and 10 same-process calls per control.

Exact evidence for the named deterministic renderer boundary.

05 Account

Every result keeps its denominator.

Wrong, missing, abstained, low-confidence, failed, retried and unevaluable cases retain distinct labels inside the frozen population. A zero denominator is reported as not evaluable, never as a pass.

Failures remain visible; they are not removed after the run.

The run boundary is measured too
  • Every measured run records the software build, configuration, feature flags, hardware and concurrency.
  • Source, answer-key and output artifacts are bound by hashes; identity proves custody, while correctness is measured separately.
  • Queue time, active processing, retries, recovery and terminal failures remain part of the operational result.
  • Production credit data remains in Saudi Arabia and the institution holds its keys; the run manifest records the deployed boundary.
Evidence carried with every measured result
A rate without its measurement record is not accepted as evidence.
Corpus named lane, classes and case counts
Ratio exact numerator, denominator and NE count
Identity build, configuration and run ID
Slices document, language and acquisition mode
Exclusions pre-registered reasons remain visible
Failures case-level examples and disposition
Limits known and unmeasured populations
Owner threshold and decision ownership

Engineering tests establish bounded implementation behaviour on their named fixtures. Measured performance for an institution is established on its locked synthetic or anonymised pilot corpus with a recorded run identity.

Inspect the measurement boundary.

Review the corpus design, metric contracts, run identity and failure ledger, then define the institution's thresholds for a measured pilot.

Request a fit assessment