Add tools/build_site.py: runs the full pipeline in dependency order
(build_content -> build_listings -> seo_inject --all --sitemap ->
build_botlinks -> build_search_index), each as a subprocess so a failing
step aborts the build before anything downstream runs (no partial
deploy). --check verifies non-destructively; --images also runs
optimize_images. This is what CI will run in U5.
Fix a build_content idempotence bug found via U4: the hero and
article-header regions were emitted with leading indentation while the
regex left the template's pre-tag whitespace in place, so indentation
compounded (+16 spaces) on every rebuild. First line is now unindented;
the full build is verified idempotent (identical git tree across two
runs). Added a fixpoint test (re-rendering a page's own output is
byte-identical) so this can't regress.
docs/local-preview.md: build + serve on :8000 in all three languages.
tools/README.md: document build_site + build_content.
26 tests pass.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replace the hand-maintained ARTICLE_DATES registry as the source of
truth with dates read from content/ front-matter, so publishing a post
via the CMS needs no manual registry edit (plan R4).
- tools/content_index.py: load_dates() reads content/ front-matter into
{source_folder/slug: {published|modified: iso}}. Standalone (os+yaml),
no import cycle with seo_inject.
- tools/seo_inject.py: ARTICLE_DATES is now {**static_fallback,
**content_dates}; content wins. The static table is kept only as a
fallback for entries not present in content/.
- migration now preserves the published-vs-modified distinction
(date_type), so the switch is byte-identical: content-derived dates
equal the old static table exactly (17/17, incl. the one 'modified').
Verified: full `seo_inject --all --sitemap` + build_listings +
build_botlinks run produces ZERO changes to any site page. build_content
--check still reproduces all 51 pages.
build_botlinks.py needs no change (it scrapes built pages, which are a
pure function of content/). Listings get content-driven dates for free
via seo_inject.ARTICLE_DATES.
tools/tests/test_content_index.py: 5 tests (dates == legacy registry,
merged registry unchanged, published/modified distinction, events key on
projects/<slug>, one ISO date per entry).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Add tools/build_content.py: renders every content/ entry into its blog /
news / event detail page by self-templating the page's own chrome and
replacing only content-derived regions (title, article header, hero,
body, and — for new pages only — the scrollspy).
Faithful reproduction, proven against the live site:
- body_format: html fragments injected verbatim, so guide-boxes, inline
figures and event galleries render unchanged.
- scrollspy sidebars are preserved (labels are hand-curated, shorter than
the H2s); only generated when a page has none.
- <title> suffix derived per page, so EN/SI " - MSOS", MK " - МСОС", and
suffix-less event titles all reproduce exactly.
- hero /images/… rewritten to the page's ../../../ depth.
- ASCII-only console output (no cp1252 crash on Windows/CI).
tools/tests/test_build_content.py: 8 characterization tests; the core one
renders all 51 pages and asserts each is semantically equivalent
(whitespace/attr-order-insensitive) to the current page — the R5
no-regression guarantee.
Does NOT overwrite the committed pages; whether generated pages are
committed vs built fresh in CI is decided in U4/U5. Run:
python tools/build_content.py --check # verify equivalence
python tools/build_content.py --write # emit pages (for U4)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Establish content/ as the source of truth for editable content and
back-fill it from the existing detail pages, per U1 of the no-code CMS
plan.
- docs/content-model.md: front-matter schema (EN/MK/SI), body_format
html-vs-markdown, content/events -> projects/ output mapping.
- tools/migrate_pages_to_content.py: bs4/yaml extractor; pulls title,
subtitle, category, ISO date (from ARTICLE_DATES), hero image+caption;
preserves the article body PLUS trailing sibling sections (event
galleries) verbatim as body_format: html. Reports unparseable pages
instead of emitting half-formed files.
- content/**: 51 files (17 entries x EN/MK/SI), 0 skips.
- tools/tests/test_migrate_roundtrip.py: 12 characterization tests
(structured fields, news-has-no-subtitle, event gallery/lightbox
preserved, Cyrillic UTF-8 round-trip, every page parses).
- .gitignore: ignore __pycache__/ and *.pyc.
U1 is unblocked and standalone. U2 (render content -> pages) is next;
U5/U6 remain blocked on Open Question Q1 (Gitea<->GitHub topology).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>