← Part 26

AQ-TRUST-001

Defines the six compounding components of external trust with the mechanism the programme runs and the check an outsider performs without our cooperation, then the thirteen-metric framework with baseline and target-setting methods rather than invented targets.

Charter artifact · gated by Part 26 · revision 0, draft for internal review

AQ-TRUST-001 — External trust architecture and demonstration metrics

DEFINITION IN USE

External trust in this programme means one thing only: whether a competent outsider can check a stated claim without the programme's cooperation. It is not perception, reputation or communication. Every mechanism below is written so that it produces an artefact an outsider can act on; a mechanism that only produces an impression is not recorded here.

The six components compound rather than add. A programme strong in five and absent in one is checked at the level of the missing one, because the missing component is where an outsider stops being able to proceed independently. Nothing described here has been built or exercised; this is a defined process at Concept level, intended to be running before AQ-GATE-05.

THE SIX COMPONENTS

ComponentOperational mechanism the programme runsCheck an outsider performs unaidedFailure signature
C-1 Reproducible workEvery published quantity ships an artefact bundle: written procedure, raw capture, uncertainty budget dated before the run, DC-1 tag set, code or configuration hash, and environment record.Take the bundle, re-execute the procedure, compare the conclusion — without asking us anything.A figure with no bundle, or a bundle whose budget is dated after the run.
C-2 Transparent limitationsEach document carries a "what this does not show" section naming the assumptions that were not tested, plus the live state of CF-1..CF-7 and F-1..F-14 with owners.Diff the limitation section against the claim section; anything claimed and not limited is either proven or overstated.Limitations written only as future work, or a failure mode with no current state.
C-3 Independent scrutinyNamed external reviewers receive read access to the Evidence Fabric slice behind a claim, and unscripted question time at every demonstration from DM-2 upward. Reviewer time is paid from programme funds, so scrutiny does not wait on external capital or partnership.Ask a question we did not select, on a slice we did not curate, and see whether it resolves to an evidence id at the time.Scrutiny available only through curated question sets or summarised extracts.
C-4 Consistent claim disciplineEvery externally released sentence is bound before release to an evidence level, a gate, and a Fabric id, in a claim register. Release is blocked on an unbound sentence.Sample any public sentence, follow its id into the Fabric, and confirm the evidence supports exactly that sentence and no wider one.A public sentence with no id; an id that supports a narrower statement than the sentence makes.
C-5 Documented correctionsAppend-only correction log. Each entry records what was claimed, what is now known, who identified it, its severity class, and what changed in the process so the class of error cannot recur silently. Public text is never edited without a log entry and an archived prior version.Compare an archived earlier release against the current one; every difference of substance must have a log entry.A changed claim with no entry; entries that record the change but not who found it.
C-6 Repeatable capabilityAny capability statement requires the same procedure executed by a second operator on a separate day, both runs recorded in full including the run that went badly.Request the second run's record and compare dispersion and configuration against the first.One golden run; a second run whose configuration differs and is not logged as differing.

Figure 1 — The outsider's check path; every hop must exist before the next is usable

  PUBLIC SENTENCE
        |  (C-4 claim register)
        v
  EVIDENCE ID ---------------------> CORRECTION LOG (C-5)
        |                                    ^
        |  (C-3 read access)                 | entry required for
        v                                    | any change of substance
  EVIDENCE FABRIC SLICE (L6) ---------------+
        |
        |  (C-1 bundle: procedure + raw + budget + tags + hash + env)
        v
  RE-EXECUTION BY OUTSIDER
        |
        +--> same conclusion  -> claim stands at its stated level
        +--> different result -> C-5 entry, externally identified
        +--> cannot proceed   -> record WHICH hop was missing;
                                 under DC-2 this is 'undetermined',
                                 never 'unconfirmed but probably fine'

  C-2 runs beside every hop: the limitation section states which
  hops were never attempted.  C-6 supplies the second run without
  which no capability sentence is released at all.

THIRTEEN-METRIC FRAMEWORK

No numeric target is stated in this document, and none may be added before the corresponding baseline exists. A threshold asserted before first measurement is a MODELLED value presented as a measured one, which DC-1 forbids and which would make the first campaign an exercise in meeting a number the programme invented. Each metric therefore declares how its baseline is obtained and how its target will later be established from that baseline. Until then the target field reads undetermined, which is a value, not a gap.

IDCategoryWhat is countedBaseline sourceHow the target will be set
TQ-1TechnicalShare of published quantities carrying an uncertainty budget dated before the run.First bench campaign after AQ-GATE-06.Definitional: every quantity. The measured figure is the escape rate and where escapes occur.
TQ-2TechnicalDispersion of the same quantity across repeats by different operators, in the quantity's own units.First three repeats under C-6.Set only after the instrument's own noise floor is characterised separately; a target set earlier would encode instrument noise as system behaviour.
TQ-3TechnicalFraction of decision cycles in which one or more inputs resolve to "undetermined" (DC-2).AQ-GATE-05 simulation campaign, re-baselined on hardware at AQ-GATE-06.Stated as a maximum for a named input set, derived from measured sensor availability. Never "as low as possible" — suppressing this figure would mean defaulting missing inputs to benign.
SC-1ScientificAssumptions named in a published result, against assumptions identified by review afterwards.Review of the first three published results.Set as an acceptable count of review-discovered assumptions per result once review data exists.
SC-2ScientificPublished results whose outcome did not support the design intent, over all published results.Observed, never designed.No quota is ever set. A floor is introduced only if the observed value is zero across a stated window, since a sustained zero is evidence of filtering rather than of correctness.
RP-1ReproducibilityArtefact bundles re-executed to the same conclusion by someone outside the build team.First external re-execution attempt.Set from attempt data together with the attempt count needed for the figure to mean anything; reported with its denominator always.
RP-2ReproducibilityBundles containing all six required elements listed under C-1.First public release.Definitional: all six present. This is a completeness check, not a performance figure, so it needs no measured target.
OB-1Observer confidenceUnscripted demonstration questions resolved to an evidence id at the time, against deferred, against unanswerable.First DM-2 walkthrough.Set after three sessions from the observed deferral mix, and reported per category rather than as a single ratio.
OB-2Observer confidenceDeviations between the witnessed run and the internal run of record (procedure, configuration, trust state entered).First externally witnessed run.Definitional: zero undeclared deviations. Declared deviations are counted and characterised, not targeted.
TR-1TraceabilityConsequential transitions reconstructable from the Evidence Fabric alone by a reviewer without operator assistance (DC-5).AQ-GATE-07 reconstruction exercise.Pass criterion is binary and definitional: all of them. The measured figures are reviewer time and query count per transition, whose targets are set from the first exercise.
TR-2TraceabilityValues carrying a MEASURED / DERIVED / MODELLED tag (DC-1).First release containing quantities.Definitional: every value. The tracked figure is the leak rate and the surface on which leaks appear.
CI-1Claim integrityExternal sentences with a resolvable evidence id, and sentences whose id supports a narrower statement than the sentence makes.First release passing the claim register.Definitional for coverage; the over-reach count is reported per release with the offending sentences quoted.
CI-2Claim integrityCorrections identified by the programme itself, against corrections identified externally, per severity class.First ten correction-log entries.See below. No target is set for the ratio itself; the target is set on the process that produces it.

CI-2 AS THE PRIMARY INDICATOR

MOST INFORMATIVE SINGLE INDICATOR

The ratio of self-identified to externally-identified corrections says more about a programme's credibility than any technical metric it reports. A programme finding its own errors is performing the checking it claims to perform. A falling ratio means outsiders are finding what internal review did not, and that is a statement about the review process, not about the individual error. The ratio is therefore reviewed at every gate alongside the gate's own evidence.

HOW CI-2 IS READ HONESTLY

The ratio is trivially gamed by logging inconsequential corrections, so every entry carries a severity class and the ratio is reported per class, never as a single number. It is always published with both denominators: a high ratio over four total corrections carries no information. A programme at Concept level with no corrections logged has an undetermined ratio — under DC-2 that is recorded as undetermined and not read as a clean record. L8 analysis may cluster log entries and propose that a pattern exists; a named human classifies severity and writes the entry, because L8 is advisory only and an advisory system grading its own programme's honesty is not a check.

The framework as a whole is reviewed at each gate from AQ-GATE-05 onward. A metric that cannot be baselined at its stated gate is not quietly dropped: it is carried forward with its missing baseline named, which is the same treatment an unmet gate receives in AQ-CRED-001.