FORGE is a research and evidence framework. It is not a personality, an agent, or an assistant. It does not speak, decide, or hold an identity. It is the procedure that reasoning parties submit to.


Changelog from v0.1

v0.2 exists because a real mission produced two failures that v0.1 did not prevent, and because an outside instrument turned out to contain three rules worth adopting. Under the freeze discipline, a version change is a commit with a stated reason. This is the reason.

ChangeOrigin
§4.1 Reconciliation before contradictionWe logged a 3.3x gap between two institutions as a contradiction. It was not one. We were comparing two different concepts.
§4.2 Definitional blindness as a distinct failure modeDiscovered when an indicator turned out to be structurally incapable of registering a policy change, not merely late.
§3.1 Instrument-nature confidence capAdopted from an external course instrument that capped survey-based indicators at MEDIUM confidence regardless of data quality.
§3.2 Lag-insensitive instrumentsOur own correction, caught mid-mission, after we flagged a WTO bound rate as stale while simultaneously writing that bound rates cannot go stale.
§7.1 Observation-level provenanceAdopted from the same external instrument, which carried provenance one level deeper than we did.
§8.1 No black-box scoreAdopted; independently converged with our own auditability requirement.

Prime Directive

Never protect the hypothesis. Protect the search for truth.

Every rule below is a mechanical consequence of that sentence. Where a rule and the Directive appear to conflict, the rule is wrong and gets rewritten.


1. Roles and the shape of authority

NodeWhat it isCategory
DirectionThe human principal. Sets objectives, supplies reality.Entity
Model AA reasoning party.Entity
Model BA second reasoning party, different provider, adversarial by design.Entity
FORGEThe procedure.Method
DIRECTION — objectives, judgment, reality
    ├── Model A ── works under: FORGE
    └── Model B ── works under: FORGE

FORGE does not hang off either model. Both submit to the same method. If it disciplined only one, the tribunal would be asymmetric and the adversarial pass would be theatre.


2. The two rules that define the method

Rule 1. Epistemic authority does not belong to the model, and does not belong to the principal either

Mission Control resolves actions. Evidence resolves beliefs.

Direction decides what to do. It does not decide what is true. Where Model A asserts A and Model B produces strong evidence for not-A, the ledger records CONTESTED. Direction may still act on A for commercial reasons. What is not legitimate is a record that reads TRUE BECAUSE THE PRINCIPAL PICKED IT. If the principal can break epistemic ties, commercial pressure launders itself into evidence status.

Rule 2. The software denies what the model grants itself

A constitution without enforcement is literature.

Uncertainty propagation. A node may not hold a status stronger than the weakest of its material dependencies, over a single namespace covering claims, evidence and calculations.

allowed_status = min(requested_status, min(status of each material dependency))

Reserved fields. The author writes requested_status. The validator writes effective_status and clearance.final_status. A ledger that pre-fills a validator-owned field is rejected in its entirety before a single threshold is evaluated.


3. Evidence model: two axes, not one ladder

ProvenanceValidation state
P0model-generatedV0unverified
P1secondary sourceV1internally consistent
P2primary sourceV2deterministically verified
P3independent corroborationV3validated by observation
V4experimentally tested
V5independently reproduced

Institutional authority does not raise provenance. A press release from a prestigious institution is P1, not P2. Prestige is not provenance.

3.1 Confidence caps on the nature of the instrument (new in v0.2)

Data quality and instrument nature are different constraints, and the second one binds independently of the first.

An instrument can be perfectly accurate and still not support a confident read of the present. Cap confidence on what the instrument is, before assessing how well it was measured.

A survey-based indicator published every few years cannot support HIGH confidence in a trend, even when its latest observation is recent, its coverage is complete and peers are available. A tariff schedule compiled with a multi-year lag cannot support HIGH confidence about the current legal regime, even when the number is exactly right for its own year.

Declare the cap alongside the indicator, before running the analysis, so it cannot be adjusted after seeing the result.

3.2 Lag-insensitive instruments (new in v0.2)

The mirror of 3.1, and it must be declared with equal care, because an age penalty applied indiscriminately destroys good evidence.

Some quantities do not decay with observation age, because they can only change through a formal, observable act.

A WTO bound tariff rate observed three years ago is not stale. It is unchanged, unless a renegotiation occurred, and whether one occurred is itself checkable. Applying a generic age penalty to such an instrument is not conservatism; it is an error that discards the most reliable evidence in the set.

The test is not "how old is the observation" but "what would have to have happened for this to have changed, and did it happen".


4. Disagreement between sources

The naive handling is to resolve by precedence: prefer source A over source B by institutional rank. That produces one number and destroys the information that two competent institutions disagreed, which is frequently the most informative object in the evidence base.

FORGE records and classifies instead. But before classifying, reconcile.

4.1 Reconciliation before contradiction (new in v0.2)

A disagreement is not a contradiction until the definitions have been reconciled. Test within-publisher before asserting between-publisher.

This rule exists because we broke it. Two institutions reported the same-sounding quantity for the same country with a 3.3x gap, and we logged it as an institutional contradiction with a HIGH severity. It was not a contradiction. One reported the rate that ignores preferential agreements; the other reported the rate that uses them. Both were correct about different things, and the comparison was our error.

The procedure that catches it, in order:

  1. Read the metadata definition of each series, not its label. Labels collide; definitions do not. "Applied tariff" names at least two distinct quantities.
  2. Eliminate the publisher as a variable. Ask whether one publisher publishes both concepts. If it does, compare within that publisher first. A gap that survives is institutional. A gap that collapses was definitional.
  3. Only then classify.

When reconciliation succeeds, record it as resolved and state explicitly that neither source was negated. This matters: the discipline that stops an analyst from discarding an inconvenient official source is the same discipline that must stop them from claiming a contradiction that does not exist.

Reconciliation is usually worth more than the disagreement was. In the case above, the reconciled gap turned out to be the measured value of a country's trade-agreement network, which was a better finding than the contradiction would have been.

4.2 Classification taxonomy

Five classes. They are not the same thing and must not be pooled, because the remedy differs for each.

ClassDefinitionRemedy
ContradictionSame concept, same period, different publishers, different values, surviving reconciliationReport both; adopt neither as "the" value; state which was used and why
Definitional divergenceSame label, different definitionCarry both as separate quantities; never quote one as "the" measure
Vintage mismatchValues of different ages presented as equally currentStamp every cell with its observation year; flag on age
Measure versus realityThe indicator moved opposite to the phenomenon it claims to measureReport as a finding, not as an anomaly to be smoothed
Definitional blindnessThe indicator is structurally incapable of registering the phenomenon, by construction rather than by lagSee 4.3

4.3 Definitional blindness (new in v0.2)

Staleness is curable by patience. Blindness is not. Do not treat them as the same defect.

An indicator is stale when it will eventually describe the present, once the publisher catches up. An indicator is blind when it will never describe the phenomenon, because of how it is constructed.

The case that produced this rule: a policy change was aimed precisely at the segment of trade that a weighted indicator systematically under-weights. Waiting for the next data release would not have surfaced it, because the weights are dominated by the segment the policy deliberately left alone.

Test for blindness: if the publisher updated to today's data tonight, would this indicator move? If the honest answer is no, and the phenomenon is real, the indicator is blind and must be replaced rather than refreshed.


5. Verification is a tuple, not a property

Verification is a property of (source, verifier, method, date), never of the source alone.

A field of the form verified_url: true is malformed and must not exist in a schema.

  • A 403 is logged as a 403. It does not mean the source does not exist. It means that

verifier could not reach it by that method on that date.

  • A failed retrieval by one party is not evidence against the source.
  • Every unreached source carries an upgrade route: the specific, cheaper action that

would move it to a higher tier. An unreached source without an upgrade route is an excuse, not a finding.

  • This yields an argument for multi-party architecture independent of adversarial

diversity: access diversity. A second party does not only think differently, it reaches different things.


6. Independence, and what does not count as it

  • Repetition is not independence. One hundred runs of the same model under different

seeds are one observation.

  • Complementarity gain is the metric for a second verifier: what did it reach that the

first could not?

  • Self-reported and externally audited violation rates are different quantities. A rate

that falls because the detector degraded looks identical to one that falls because behaviour improved. Use a fixed-size adversarial sample, set before the run.


7. No phantom capability

specified ≠ implemented ≠ validated

Every capability carries one of three states and the state is part of the record: implemented and running; specified only; conceptual. A system that speaks about specified capabilities in the present tense is describing a system that does not exist.

7.1 Observation-level provenance (new in v0.2)

FORGE previously carried provenance at the level of the claim. That is one level too shallow for quantitative work.

Every observation retains: entity, entity code, period, indicator, indicator code, unit, expected frequency, source, verifier, method, retrieval date.

Claim-level provenance answers "where did this argument come from". Observation-level provenance answers "where did this cell come from", which is the question an auditor actually asks, and the one that makes a workbook reproducible rather than merely sourced.

Declare expected frequency per indicator and evaluate coverage against that expectation, not against the calendar. An irregular series is not deficient for having fewer observations than there are years.


8. Mission structure and clearance

A mission ledger contains claims with dependencies, evidence with verification tuples, deterministic calculations, a contradiction register, a friction record, and a mandatory artifact_path. A mission without a produced artifact cannot reach PASS; its ceiling is NO_CLEARANCE, which is distinct from failure.

Missions are chosen so that they double as real work wherever possible. A method that only runs on synthetic exercises never meets the conditions it was built for.

8.1 No black-box score (adopted in v0.2)

Every flag records the data used, the rule applied, and the reason it fired. Where more than one rule could fire, record which one did.

And the corollary, which is a refusal rather than a requirement:

Do not generate an arbitrary composite score. A number that aggregates incommensurable dimensions through undeclared weights is not a measurement. Where a composite is genuinely required, its weights are declared before the data is seen and its sensitivity to alternative weightings is reported alongside it.

8.2 Clearance outcomes

CodeStateMeaning
0PASSEpistemic and friction thresholds both met
1FAIL_EPISTEMICEvidence standards not met
2FAIL_FRICTIONCost of operation exceeded its ceiling
3FAIL_BOTHBoth
4NO_CLEARANCENothing was submitted for judgement
5SCHEMA_REFUSEDReserved field authored, or structural violation

FAIL_FRICTION is a first-class failure. A check that cries wolf gets ignored, and a method nobody runs has an effective accuracy of zero.


9. The freeze

Thresholds live in a version-controlled file, never in prose.

FORGE cannot pass by redefining PASS.

Any threshold change is a commit with a stated reason, and that commit invalidates the pilot in progress. Criteria are not changed after a run begins, not to rescue a result and not to reflect something learned mid-run. When a mission fails, the thresholds do not move; the failure is recorded, the root cause investigated, and a new version born carrying the lesson.


10. Reusable method lessons

  1. A materiality rule anchored to a moving average never fires under slow drift. If the average tracks the trend, deviation from the average stays at zero while the underlying quantity walks away. Pair every moving-average test with a level test against a fixed base, and record which path fired.
  2. A constitution without enforcement is literature.
  3. Repetition is not independence.
  4. A violation rate that falls because the detector got worse is indistinguishable from one that falls because behaviour improved. Fix the sample size in advance.
  5. A check that cries wolf gets ignored. That is a friction failure, not a cosmetic one.
  6. Verify the frame before verifying the data. Eligibility rules, scope constraints and assignment boundaries are cheap to check and expensive to get wrong. Correct figures inside the wrong frame are still a wrong answer.
  7. A conclusion that survives a complete replacement of its comparison set is stronger than one validated only against the sample originally chosen.
  8. Every unreached source carries its upgrade route.
  9. A disagreement is not a contradiction until the definitions have been reconciled. Read definitions, not labels. Eliminate the publisher as a variable before blaming the publisher.
  10. Staleness and blindness are different defects with different remedies. Ask whether the indicator would move if the publisher updated tonight.
  11. An age penalty applied indiscriminately destroys the best evidence in the set. Ask what would have to have happened for the value to change, not how old it is.

Operating summary

Six commitments bind anyone working under FORGE:

  1. State the evidence tier of every load-bearing figure, and never quote a Tier 3 assumption as though it were verified.
  2. Where you cannot verify, say so and leave the item open, with the upgrade route named. An open item is a finding. A filled gap is a fabrication.
  3. Never hold a conclusion stronger than its weakest dependency.
  4. When two of your own readings disagree, record the disagreement rather than selecting the more convenient one.
  5. Before asserting that two sources contradict each other, prove you are comparing the same quantity.
  6. Cap confidence on what the instrument is, not only on how well it was measured, and exempt the instruments that genuinely cannot go stale.