Index Methodology
Experimental · proof-of-architecture testnet · last reviewed 2026-08-02
Scale
The index is expressed on a 0–100 scale. The bound is applied in the scoring engine as min(max(ets, 0.0), 100.0) and has been unchanged since the first commit in the repository, 1c6b586 (2026-03-19).
Values published for May 2026 (0.284–0.375) are therefore fractions of 100 — approximately 0.3 out of 100 — not fractions of 1. Any description of this index as being on a 0–1 scale is incorrect, and this page is the authoritative statement of the scale.
Definition history
The aggregate has been redefined four times. Every change is listed with the commit that introduced it, so any published figure can be matched to the definition in force when it was computed.
v0 · 2026-03-19 · 1c6b586
Mean ETS across active subjects, excluding those whose confidence band was 'insufficient' and those without a snapshot.
v1 · 2026-05-02 … 2026-05-04 · bb6ea47 · 6879b50 · 4d6350b
Moved to SQL aggregation; the confidence-band exclusion was removed, so all scored subjects count.
v2 · 2026-06-29 · 2026-07-01 · e4e88da · 034ebae
Subjects with a zero score excluded from the mean; a one-sided damper was added (floor = 0.95 × previous), bounding decreases only.
v3 · 2026-08-02 · beca2a1
Damper made symmetric (±5% per cycle). The one-sided form allowed the index to drift upward without a matching change in evidence.
v4 · 2026-08-02 · 6502437
Stale evidence is counted at its floor weight instead of being dropped. Dropping it removed the lowest-weighted items from a mean, so a decaying evidence base raised the score — measured at 3.274× inflation on 2026-08-02 versus 1.000× in May.
Re-computation of 2026-08-02
Following v3 and v4, the existing population was re-scored in a single batch on 2026-08-02: 8 cycles, moving the aggregate from 3.0526 to 2.4252, at which point successive cycles changed it by less than 0.5%.
These 8 cycles were executed in one batch and do not represent 8 days of operation. The batch is recorded in the audit log as a migration re-computation. Scheduled operation is one cycle per day.
The May 2026 cycle re-derives from the database with the current engine at 0.000% deviation for all twelve principals, confirming that these changes apply forward only and do not restate historical figures.
Current trajectory
No new evidence is currently being ingested, so existing evidence continues to age and the index continues to decline toward the floor weights that bound it. Projected analytically from the decay floors:
| Date | Index | Basis |
|---|---|---|
| 2026-08-02 | 2.4252 | current, after the v4 re-computation |
| 2026-08-24 | 2.0111 | projected |
| 2026-09-04 | 1.8354 | projected |
| t → ∞ | 0.0549 | asymptote, set by the SF-07 decay floor |
The index does not decay to zero: the asymptote is set by the decay floor of the dominant signal family.
Why every subject is in the lowest band
All scored subjects sit in the suppressed band. That is the correct reading of the present evidence base, not a fault in the scoring: coverage is 2 of 7 signal families, from a single source, at the two lowest evidence tiers. A band above suppressed requires an index of 10.0 or higher.
This is not a change of behaviour. Across the entire recorded history — all 21,981 snapshots, including the May 2026 window whose figures are published — no snapshot has ever reached 10.0. Band-level differentiation has never been demonstrated by this testnet; the published May result is a numeric range, and it remains accurate.
The mechanism does discriminate where the inputs differ: of twelve principals, the one whose staleness ratio differs (54% against 77%) is separated from the other eleven. The population contains exactly two distinct evidence profiles and the index produces exactly two distinct values. Widening that requires more sources, more signal families and higher evidence tiers — a supply problem, not a scoring problem.
Reproducing these figures
The scoring engine takes an explicit reference time, so a historical cycle can be re-derived deterministically rather than against the current clock. One known limitation: the anti-gaming multiplier is written onto an evidence row while a burst is in progress and persists there, so it cannot be re-derived from a later instant — a faithful replay restores the stored multipliers. Making the engine a pure function of evidence and time is core-hardening work.
A scoring run whose output is published cannot fall back to substituted coefficients: the engine refuses to start unless the deployed coefficient set is in force.
Observational analytics only. Not a rating, certification, or advisory service. All outputs carry EXPERIMENTAL status.
