Institution boundary
Decision and operating authority Identity, authority, keys, data policy and the credit decisionEvery capability is tested, measured and replayed.
RSEEN treats source linkage, extraction, formulas, conflict handling, evidence decisions and replay as measurable contracts. Each run records the corpus, exact numerator and denominator, build identity, exclusions and failures. Test volume shows regression breadth; it does not stand in for real-file accuracy.
How the institution's method is encodedBounded controls run on 2026-07-19. Not a qualification, pilot, production or audit result.
Shared-responsibility boundary
A security boundary your team can inspect.
Data never leaves Saudi Arabia, and the institution holds the keys. The map shows where each layer runs and who owns the decision.
RSEEN platform boundary
Optional providers and connectors
Documented routes Inference, OCR, messaging, signing and client systemsBuilt into RSEEN
- Request bound to case, provenance and execution route
- Traceable case result or clear stop reason
- Credit tables computed by formula, with any withheld grade visible
- RSEEN Core works as a standalone product; RSEEN OCT expands the same platform when required
Configured for the institution
- Identity, roles, key custody and storage
- Provider eligibility, residency, egress and logs
- Backup, restore and recovery objectives
- Support, break-glass, retention, exit and connectors
The map keeps responsibility clear from evidence and computation through workflow and credit decision.
Seven measures. No blended verdict.
The institution sets thresholds before a locked run. Every measure reports its named corpus, exact numerator and denominator, unevaluable population, required slices, run identity and known limitations. Synthetic and pilot results remain separate; an unsupported or failing slice is never averaged away.
Citation precision and recall
Measures whether each citation resolves to the right source and supports the material fact, and whether every material fact that needs support has it.
Extraction error by document class
Reports correct, wrong, missing, spurious and unevaluable fields separately by language, acquisition mode, document class and field type.
Formula reconciliation
Compares engine and end-to-end outputs with an independently specified oracle, while retaining missing or unavailable inputs in the result.
Conflict coverage
Measures known conflicts caught with both values attributable through reviewer resolution, plus false flags on clean control pairs.
Withhold correctness
Checks both paths: insufficient-evidence cases pause for resolution; sufficient-evidence cases proceed to the independently expected grade.
Replay stability
Compares scores, grades, evidence state, reasons and rendered bytes across repeated runs with frozen inputs and configuration.
Throughput and latency
Reports queue, extraction, analysis and total time by file class on named hardware and concurrency, including retries and failures.
Qualification requires a locked synthetic or anonymised pilot corpus, institution-approved answer keys and thresholds, and a reproducible run. The baseline above does not meet that boundary.
From run identity to accountable result.
The harness fixes what is measured, separates evidence lanes, exercises controls and clean twins, replays frozen inputs, and keeps every failure in the denominator.
Name the object before measuring it.
A run pins source hashes, answer-key identity, build, configuration, hardware and corpus membership before outputs are opened. The same record distinguishes live execution, persisted replay, fixture replay and interface smoke.
A result without a pinned identity is not scored.
Engineering evidence is not qualification.
Development fixtures test implementation behaviour. Locked synthetic cases measure controlled holdouts. Anonymised pilot files measure the institution's declared document mix. Their results and denominators remain separate.
A development fixture never enters a qualification denominator.
Test the hard path and its clean twin.
Controlled pairs test a clean source against one changed factor: a conflict, missing evidence, OCR corruption, formula edge case or unsupported citation. Positive and negative paths are counted together.
Coverage and false-positive behaviour are measured together.
Frozen inputs must replay cleanly.
Replay compares structured values, score, grade, evidence state, reasons and rendered bytes. The first bounded deterministic baseline produced 70/70 byte-identical outputs across 7 controls and 10 calls each.
This is deterministic-renderer evidence, not end-to-end qualification.
Every result keeps its denominator.
Wrong, missing, abstained, low-confidence, failed, retried and unevaluable cases retain distinct labels inside the frozen population. A zero denominator is reported as not evaluable, never as a pass.
Failures remain visible; they are not removed after the run.
- Every measured run records the software build, configuration, feature flags, hardware and concurrency.
- Source, answer-key and output artifacts are bound by hashes; identity proves custody, while correctness is measured separately.
- Queue time, active processing, retries, recovery and terminal failures remain part of the operational result.
- Production credit data remains in Saudi Arabia and the institution holds its keys; the run manifest records the deployed boundary.
Engineering tests establish bounded implementation behaviour on their named fixtures. Only a sealed, reproducible run on a locked synthetic or anonymised pilot corpus can support a measured performance claim for that corpus.
Inspect the measurement boundary.
Review the corpus design, metric contracts, run identity and failure ledger, then define the institution's thresholds for a measured pilot.
Request a fit assessment