Ir al contenido

The published archive is an output, and the build refuses to overwrite content it cannot reproduce

Esta página aún no está disponible en tu idioma.

The published archive is an output, and the build refuses to overwrite content it cannot reproduce

Section titled “The published archive is an output, and the build refuses to overwrite content it cannot reproduce”

seed/scripts/populate_archive.py ends by copying the repository profile over the published archive’s MANIFEST.json (:736, a literal shutil.copyfile). The per-site metadata.csv is regenerated from the profile’s path map, so it is an output of the same process and cannot preserve anything the copy removes. apply_upgrade.py’s header already recorded the consequence: “a single re-run would have restored 916 tiny images and dropped all 72 videos — regardless of the state of any metadata.csv.”

So the direction of truth was already decided by the code, and never enforced. Measured 2026-08-26, before this change:

repo profilepublished archive
assets20042005
with field_values19042005

One ordinary build would have deleted one published asset and stripped 12,097 valuesfield_values from 1,947 assets (100 losing all of them, 1,847 losing 3–11 each) plus mature on the same set. Zero assets were richer in the repo, so the drift was strictly one-directional: the archive had been edited, or enriched by tooling that wrote only to it, and the profile never caught up.

⚠️ Nothing compared the two before writing, and nothing ever had. The drift survived because the build is not run often, and the one person who might have noticed was the one who would have run it.

The profile is the source of truth. The published archive is an output. The build enforces that by refusing to produce a run it cannot justify.

  1. Before writing, compare. If the destination holds content the source does not, the build refuses and exits non-zero, naming what would be lost.
  2. The remedy is to fix the SOURCE, never the destination. Editing the archive is undone by the next build by construction, so the error points at apply_upgrade.py — reconcile the profile, then publish.
  3. Duplicate ids are refused NON-overridably. A manifest holding two records under one id has no correct interpretation — aa seed keys on the stable id and silently takes whichever it reads last — so unlike a loss, there is no version of it that is somebody’s intended change. No flag forces it through.
  4. A deliberate removal remains possible, via an explicit --allow-regression, which states in its output that the loss is real and unrecoverable. The gate is against silent destruction, not against intent.
  5. The guard runs in --dry-run too. A dry run that passes while a real run would destroy data is worse than no dry run at all.
  • The 12,097 drifted values were carried back into the profile before the guard shipped. A guard in front of an unreconciled source would simply have blocked the tool forever, so repair is part of the decision, not a follow-up.
  • ⚠️ The guard proves nothing until it has been seen to REFUSE. It was verified in both directions — against the pre-repair profile it reports 12,097 losses and exits non-zero; against the repaired one, zero. A guard only ever observed permitting a run is untested, which is the failure ADR 0095’s 2026-08-26 amendment records at length.
  • Content-level drift is a different question from presence and is only partly addressed here: a missing key, an empty value and a differing value are three cases. Presence is enforced; divergence in value is tracked separately (#1294, #1295).
  • This does not make the archive backed up. It makes one specific destruction impossible. The published dataset still has no backup, and every other path that writes to that share is still unguarded.

Amendment, 2026-08-26 (#1294, #1295): a MEASUREMENT is not content, and this ADR does not govern it

Section titled “Amendment, 2026-08-26 (#1294, #1295): a MEASUREMENT is not content, and this ADR does not govern it”

The CHANGED_VALUE case above was left “tracked separately”, and the separate tracking found that the two issues were not the same kind of question at all.

#1294 — 160 site_a file_size_bytes where the profile and the share disagreed. The instinct this ADR creates is “the profile is the source of truth, so the share is wrong.” That instinct is not applicable, and following it was how sprint 14 nearly shipped a profile that would have made the next build refuse its own input.

A byte count is a MEASUREMENT of a file the pipeline produces, not a value the profile is free to assert. There is exactly one right answer — what kenney_hq.py build makes from the committed manifest and the pack — and the profile’s job is to describe it. This ADR governs which records exist and what values they carry. It has nothing to say about arithmetic.

Measured against a rebuilt pool: 150 of site_a’s 260 replacement rows and 472 of site_b’s 656 named a size the file does not have, and site_a’s published share agreed with the rebuilt pool on 776 of 777 records. The repository was the stale side. newSize is the size of a render, #630 and #685 both changed what frame a vector is drawn into, and nothing ever re-derived the numbers — they were measured once, by hand. Re-measurement is now a command (kenney_hq.py sizes), report-only and non-zero on drift by default so it can stand as a gate.

And it was visible without the share or the pool. balance-assets.site_a.json was emitted after those fixes and had been contradicting the replacements docs about 115 pool files inside the repository the whole time. Two committed documents naming one pool file must agree about its size; that is a test now.

#1295 — the gate could not see the pass. apply_upgrade.py --check had a term for every pass except apply_replacements, because that pass returned records processed, not records modified260/260 on every run, upgraded or not. A number that is never zero cannot be a drift signal, so the pass was left out rather than fixed, and a profile with drifted replacements passed the pre-publish gate for as long as it took someone to notice by hand.

⚠️ This is the second consequence above, arriving from the other direction: a gate only ever observed permitting a run is untested. The refusal now has a constructed-drift test that drives the real script, watches it fail, repairs the profile and watches it pass.

What a future reader should take from this. When the profile and the archive disagree, ask first what kind of value it is:

the value is…who is authoritativeexample
content — a record, a field value, a flagthe profile (this ADR)field_values, mature, which assets exist
a measurement of a file the pipeline producesthe artifact the pipeline makesfile_size_bytes on a pool render
a claim about bytes staged from elsewhereneither, until the bytes are checkedthe 11 video records — see below

The third row is unresolved and is not a byte count. Eleven site_a video records claim a size their staged file does not have, and probing the origins each record names returns exactly the profile’s number — while four of them carry a metadata.sha256 that matches the smaller staged file. Those records describe two different artifacts at once, and no rule in this ADR picks between them: it is a decision about what the published dataset ships. populate_archive.py’s pre-staged branch only checks size > 0, so nothing will surface it on its own.


Amendment, 2026-08-27 (#1311, #1312): the guard cannot see a corrupted measurement

Section titled “Amendment, 2026-08-27 (#1311, #1312): the guard cannot see a corrupted measurement”

The amendment above split content from measurement and said the artifact is authoritative for the second. manifest_guard does not implement that split, and sprint 14d found the gap.

manifest_guard.py:34-43 refuses MISSING_RECORD, MISSING_KEY and EMPTIED_VALUE, and reports CHANGED_VALUE as “NOT a loss … Reported, never refused” on the reasoning that an edit is what a change looks like and the profile is the source of truth for edits.

That reasoning is correct for content and wrong for measurements. Measured on dev before PR #1311: twelve records where the profile and the published manifest disagreed on file_size_bytes, and in all twelve the manifest matched the bytes on disk while the profile claimed larger, totalling 2,690,105,638 bytes. populate_archive.py:841 copies the profile over MANIFEST.json, so the next publish would have replaced correct measurements with wrong ones, and the guard would have reported it and proceeded.

The same property that makes a correction safe makes a corruption invisible. The sprint-14d brief cited CHANGED_VALUE approvingly as proof that re-measuring four hashes was safe. That was true, and its inverse was equally true and unstated.

⚠️ And the split is per record class, not global. metadata.sha256 is a measurement for hq records and identity for internet-root ones: sanitize_and_assemble.py:1517 mints the asset id from it and :1558-1560 derive three timestamps from it. A brief that ruled “re-measure the four hashes” and closed the question would have moved ids on the next assembly, and only the implementing agent’s refusal to follow a closed instruction stopped it.

So this ADR’s category is a property of the FIELD IN A RECORD CLASS, not of the field. Deciding “is this content, a measurement, or identity” has to be asked per class, and a rule that answers it once for a field name is wrong.

Tracked as #1312. Not fixed here: naively promoting CHANGED_VALUE to a loss would refuse every legitimate edit and make the guard unusable, which is the failure mode of over-correcting a gate.


Amendment, 2026-08-27 (#1318, #1312, #1313): the split is implemented, and its axis is source_root

Section titled “Amendment, 2026-08-27 (#1318, #1312, #1313): the split is implemented, and its axis is source_root”

The amendment above named the gap and left it open, because promoting CHANGED_VALUE to a loss would refuse every legitimate edit. Sprint 14e closed it without that cost.

CHANGED_VALUE is unchanged. A new CORRUPTED_MEASUREMENT verdict sits beside it, and the axis that selects between them is the record’s source_root, never the field name:

source_rootwhyverdict on a disagreement
site, torrent_import, internetthe bytes are staged at the destination with no reproducible source, so the destination is the artifactrefused
local, hq, packcopied from a source the profile is built against, so the share can lag itreported, permitted

This is the previous amendment’s “per record class, not per field” rule given a mechanism. It is measured rather than asserted: against a kenney-hq pool built fresh on 2026-08-27 (945 vectors rendered, 86 bitmaps copied), all 656 of studio-b’s hq records match the profile and only 264 match the published manifest.

⛔ The direction error this ADR’s own author then made

Section titled “⛔ The direction error this ADR’s own author then made”

That 656-versus-264 measurement exists because the sprint brief asked for the opposite of the correct thing. Site_b’s published manifest disagreed with the profile on 392 file_size_bytes; the manifest matched the bytes on disk on all 392 and the profile on none; the brief concluded the profile was stale and asked for it to be reconciled to its files. Doing so would have overwritten 392 correct values with stale ones, and the next build would then have refused its own input.

Matching its own bytes proves only that a copy is SELF-CONSISTENT. A stale copy agrees with itself perfectly. All 392 records are hq, copied from the pool, so the authority is the pool.

So this ADR’s rule needs its sharper form: the artifact is what the pipeline PRODUCES, not where the pipeline WRITES. A destination is downstream of the artifact and inherits its staleness silently. The record’s source_root is what points at the artifact, which is why the verdict above keys on it.

⚠️ Worth recording that this ADR’s measurement-versus-content rule was written one day earlier by the same author who then applied it to the wrong noun. A rule stated at the level of “which side is authoritative” is not usable until it also says how to find the side.

The larger site_b defect, which no issue had named

Section titled “The larger site_b defect, which no issue had named”

While the 392 were being disputed, 6,806 field_values across all 1,306 site_b records existed at the share and not in the profile, in the same eleven keys as #1275’s. Since populate_archive.py:736 copies the profile over MANIFEST.json, an ordinary run would have stripped every one. The new guard reported 6,806 losses; manifest-reconcile.site_b.json carries them back (6,806 filled, 0 overwritten, 1,306 ids unchanged in both directions) and the guard now reports 0.

⚠️ The generalisable miss: a sibling artifact’s known defect was not tested for. Site_a’s defect was missing field_values; site_b was measured for file_size_bytes instead, one field was checked, and the finding was generalised from it. When two artifacts come off one pipeline, test the second for the first one’s defect before reporting whatever the first probe happened to find.


Amendment, 2026-09-21 (#1319, ADR 0098): a committed identity migration is not a deletion

Section titled “Amendment, 2026-09-21 (#1319, ADR 0098): a committed identity migration is not a deletion”

manifest_guard.compare keyed both sides by id and filed every destination id absent from the source as MISSING_RECORD. Two migrations (#1293, #1310; ADR 0098) had moved 511 post ids onto values derived from each post’s own content, the published wall still carried the old ids, and so the guard read the pipeline’s own work as a deletion. Measured read-only against the live share: posts.json reported 175 MISSING_RECORD on site_a and 336 on site_b, every one an old_id in seed/upgrades/post-id-migration.studio-a.json or .studio-b.json whose new_id the current profile holds, and 0 uncovered; MANIFEST.json reported 0 losses on both sites. Decision 1 refused a publish that would have lost nothing, and --allow-regression (Decision 4) was the only way through, which is the wrong tool: it waves through every loss, not the one thing that is not a loss.

A recorded identity migration is not a deletion, and the evidence is the committed document.

  1. A destination record whose id is a recorded old_id, and whose new_id is present in the source, is MIGRATED_RECORD: not a loss, not an addition, reported on its own line of the report.
  2. The evidence is the pipeline’s reconciliation document, seed/upgrades/post-id-migration.<stem>.json, which migrate_post_ids.py writes for exactly this reader. It is located from the posts profile alone (<stem>.posts.json), and its profile field must name the file being guarded. Nothing is inferred: not from a title, not from a member set, not from a resemblance. An id the document does not record stays MISSING_RECORD.
  3. The consumer validates the document one-to-one before it compares anything, and refuses non-overridably when it cannot: unparseable; a missing or malformed move; a profile naming another file; one old_id recorded twice; two old_id values landing on one new_id; an id on both sides (an uncomposed chain); a mapped new_id the source does not hold; an old_id the source still holds. An absent or invalid mapping is a loss, never a migration.
  4. Migration excuses nothing carried across the move. The moved record is compared against its new self with only the identity key excluded: a MISSING_KEY or EMPTIED_VALUE across the move refuses exactly as on an unmoved record, CHANGED_VALUE stays report-only, and the id transition itself never surfaces as a change.
  5. Duplicate ambiguity still refuses, non-overridably, on either side: two source records under one id (Decision 3), or two old_id values claiming one new_id in the document.

Decisions 1 to 5 and the two measurement amendments above are untouched, and MANIFEST.json comparisons produce the verdicts they produced before. Witness, with the document consumed: site_a reports 175 migrations and 0 losses, site_b 336 and 0, with 0 identity-key changes on either, and the same MANIFEST.json verdicts as before (0 losses; 42 and 440 reported changes).