Measured engineering

Every capability is tested, measured and replayed.

RSEEN treats source linkage, extraction, formulas, conflict handling, evidence decisions and replay as measurable contracts. Each run records the corpus, exact numerator and denominator, build identity, exclusions and failures. Test volume shows regression breadth; it does not stand in for real-file accuracy.

How the institution's method is encoded
Engineering baseline · synthetic / development corpus

Bounded controls run on 2026-07-19. Not a qualification, pilot, production or audit result.

8,334 backend tests · counted historical regression baseline
30/30 extraction fixture controls passed
101/101 formula DSL controls passed
232/232 conflict component controls passed
4/4 · 0/2 insufficient cases withheld · sufficient cases wrongly withheld
70/70 deterministic renders byte-identical

Shared-responsibility boundary

A security boundary your team can inspect.

Data never leaves Saudi Arabia, and the institution holds the keys. The map shows where each layer runs and who owns the decision.

Platform controls Deployment choices

Institution boundary

Decision and operating authority Identity, authority, keys, data policy and the credit decision

RSEEN platform boundary

RSEEN Core Complete standalone interface for documents, evidence, computation, conflict, review and outputs
Approved Playbook Source hierarchy, calculations, evidence requirements, exceptions and authority model
RSEEN OCT · expansion scope Queues, routing, committee workflow, conditions, handoff and origination control

Optional providers and connectors

Documented routes Inference, OCR, messaging, signing and client systems

Built into RSEEN

  • Request bound to case, provenance and execution route
  • Traceable case result or clear stop reason
  • Credit tables computed by formula, with any withheld grade visible
  • RSEEN Core works as a standalone product; RSEEN OCT expands the same platform when required

Configured for the institution

  • Identity, roles, key custody and storage
  • Provider eligibility, residency, egress and logs
  • Backup, restore and recovery objectives
  • Support, break-glass, retention, exit and connectors

The map keeps responsibility clear from evidence and computation through workflow and credit decision.

How RSEEN is judged

Seven measures. No blended verdict.

The institution sets thresholds before a locked run. Every measure reports its named corpus, exact numerator and denominator, unevaluable population, required slices, run identity and known limitations. Synthetic and pilot results remain separate; an unsupported or failing slice is never averaged away.

M1

Citation precision and recall

Measures whether each citation resolves to the right source and supports the material fact, and whether every material fact that needs support has it.

M2

Extraction error by document class

Reports correct, wrong, missing, spurious and unevaluable fields separately by language, acquisition mode, document class and field type.

M3

Formula reconciliation

Compares engine and end-to-end outputs with an independently specified oracle, while retaining missing or unavailable inputs in the result.

M4

Conflict coverage

Measures known conflicts caught with both values attributable through reviewer resolution, plus false flags on clean control pairs.

M5

Withhold correctness

Checks both paths: insufficient-evidence cases pause for resolution; sufficient-evidence cases proceed to the independently expected grade.

M6

Replay stability

Compares scores, grades, evidence state, reasons and rendered bytes across repeated runs with frozen inputs and configuration.

M7

Throughput and latency

Reports queue, extraction, analysis and total time by file class on named hardware and concurrency, including retries and failures.

Qualification requires a locked synthetic or anonymised pilot corpus, institution-approved answer keys and thresholds, and a reproducible run. The baseline above does not meet that boundary.

Measurement system

From run identity to accountable result.

The harness fixes what is measured, separates evidence lanes, exercises controls and clean twins, replays frozen inputs, and keeps every failure in the denominator.

01 Instrument

Name the object before measuring it.

A run pins source hashes, answer-key identity, build, configuration, hardware and corpus membership before outputs are opened. The same record distinguishes live execution, persisted replay, fixture replay and interface smoke.

A result without a pinned identity is not scored.

02 Separate

Engineering evidence is not qualification.

Development fixtures test implementation behaviour. Locked synthetic cases measure controlled holdouts. Anonymised pilot files measure the institution's declared document mix. Their results and denominators remain separate.

A development fixture never enters a qualification denominator.

03 Exercise

Test the hard path and its clean twin.

Controlled pairs test a clean source against one changed factor: a conflict, missing evidence, OCR corruption, formula edge case or unsupported citation. Positive and negative paths are counted together.

Coverage and false-positive behaviour are measured together.

04 Replay

Frozen inputs must replay cleanly.

Replay compares structured values, score, grade, evidence state, reasons and rendered bytes. The first bounded deterministic baseline produced 70/70 byte-identical outputs across 7 controls and 10 calls each.

This is deterministic-renderer evidence, not end-to-end qualification.

05 Account

Every result keeps its denominator.

Wrong, missing, abstained, low-confidence, failed, retried and unevaluable cases retain distinct labels inside the frozen population. A zero denominator is reported as not evaluable, never as a pass.

Failures remain visible; they are not removed after the run.

The run boundary is measured too
  • Every measured run records the software build, configuration, feature flags, hardware and concurrency.
  • Source, answer-key and output artifacts are bound by hashes; identity proves custody, while correctness is measured separately.
  • Queue time, active processing, retries, recovery and terminal failures remain part of the operational result.
  • Production credit data remains in Saudi Arabia and the institution holds its keys; the run manifest records the deployed boundary.
Evidence carried with every measured result
A rate without its measurement record is not accepted as evidence.
Corpus named lane, classes and case counts
Ratio exact numerator, denominator and NE count
Identity build, configuration and run ID
Slices document, language and acquisition mode
Exclusions pre-registered reasons remain visible
Failures case-level examples and disposition
Limits known and unmeasured populations
Owner threshold and decision ownership

Engineering tests establish bounded implementation behaviour on their named fixtures. Only a sealed, reproducible run on a locked synthetic or anonymised pilot corpus can support a measured performance claim for that corpus.

Inspect the measurement boundary.

Review the corpus design, metric contracts, run identity and failure ledger, then define the institution's thresholds for a measured pilot.

Request a fit assessment