Engineering note

How PageScript Compiles .page Files to HTML

PageScript started with a practical complaint: a small generated site could require thousands of tokens of HTML, CSS, and JavaScript. The browser needs that output. I did not want to keep all of it in the authored source.

I designed PageScript to put a compiler between the source and the browser. The language and architecture are my work; implementation was AI-assisted. A compact .page file is parsed, validated, lowered into a typed intermediate representation, and rendered as standalone HTML.

I later designed a second compiler path for source-cited system and data-lineage diagrams. It keeps the reviewed facts in an Evidence Bundle and the presentation choices in a separate Explainer Spec. A diagram should not make an uncertain relationship look verified.

In both paths, the input describes the artifact and the compiler controls the browser output.

Compiler pipeline

The regular page path is straightforward: parse the .page file, validate it, compile it into Page IR, then render it. The explainer path validates an Evidence Bundle and Explainer Spec, checks that the spec is pinned to the right evidence digest, and projects both into Explainer IR.

flowchart LR
    P[".page source"] --> PA["Parser"]
    PA --> PV["Validation"]
    PV --> PI["Page IR"]

    E["Evidence Bundle"] --> EV["Evidence validation<br/>and canonical digest"]
    S["Explainer Spec"] --> SV["Spec validation"]
    EV --> EI["Explainer IR"]
    SV --> EI

    PI --> R["Safe renderer"]
    EI --> R
    R --> H["Standalone HTML"]

The IR is the boundary I care about. By the time the renderer runs, it has normalized components, graph nodes, effects, state, layout, and evidence references. It does not need to infer structure from the original syntax.

The Rust API exposes the same stages as the CLI:

let document = parse_page_script(source);
let diagnostics = validate_document(&document, &resolver);
let ir = compile_page_ir(&document, None, &resolver)?;
let html = render_to_html(&document, None, &resolver)?;

PageScript became easier to work on once I treated it as a small compiler instead of a compact file format. Parsing, validation, IR, and rendering each have a separate job, so failures are easier to isolate in tests.

Why keep the source small?

HTML is still the output. I am not trying to replace it as a browser target. The point is to avoid authoring the same SVG structure, responsive CSS, event wiring, and runtime configuration every time I need a familiar page pattern.

For example, this PageScript source defines part of a system map:

::scene id=command-center layout=split title="Revenue map"
  ::panel id=system-map title="Signal to action graph"
    ::node id=signals label="Product Signals" status=active x=105 y=90
    ::/node
    ::node id=scoring label="Fit Scoring" status=ready x=475 y=90
    ::/node
    ::edge from=signals to=scoring effect=flow
    ::/edge
  ::/panel
::/scene

The renderer expands that into browser-ready graph markup and styling. State, node selection, and effects use a fixed runtime owned by the compiler rather than page-specific scripts.

I also designed a standard library for patterns that do not belong in the language core. Product sections, data views, documentation layouts, and grid systems are recipes that expand at compile time. If I can build something from existing primitives, I would rather add a recipe than another special case to the parser.

The project includes a reproducible stats command so the compactness claim has a number behind it. The revenue-map example contains 1,787 o200k_base tokens of authored PageScript. Its standalone HTML output contains 4,975 tokens. That is a 64.08% reduction in authored artifact tokens.

The scope of that number matters. It compares the checked-in .page source with the generated HTML after normalizing line endings. It does not include prompts, tool calls, repair attempts, or earlier context. Those values depend on the surrounding workflow. What I can reproduce is that the authored artifact is substantially smaller than the browser artifact it produces.

Validation

The compiler validates input before anything reaches the renderer.

PageScript does not accept source-authored JavaScript in its deterministic core. Validation rejects raw and script escape hatches, executable URL schemes, unsafe elements and attributes, and content that could terminate the generated <style> block. Imports resolve inside an explicit canonical root, and recipe expansion cycles fail before rendering.

Interactive output is still possible. The source declares state changes, node selection, toggles, filters, and effects. Those declarations compile into typed configuration for the fixed runtime.

The conformance suite includes hostile URLs, style-tag termination, invalid import paths, unsafe tags, missing recipes, and recursive recipe expansion. If validation fails, rendering stops. That is easier to test than trying to sanitize arbitrary output at the last step.

Source-cited diagrams

I designed the explainer path because clean diagrams often hide where their claims came from. I wanted every entity and connection to retain its source location.

An Evidence Bundle contains sources, entities, relationships, and provenance. Sources carry SHA-256 digests. Extracted and inferred claims carry citations, and inferred claims also require a rationale.

The Explainer Spec contains the presentation choices: views, grouping, callouts, and approved design tokens. It references evidence IDs rather than copying the facts, and it pins the canonical digest of the Evidence Bundle.

If I change the evidence, an old spec fails validation against the new digest. It cannot silently render a report against a different fact set. The generated explainer keeps local path:line citations and makes no external requests.

The current alpha has a deliberate limitation here. Evidence bundles are reviewed JSON inputs. Repository and dbt extraction adapters are planned, but they are not yet a supported public workflow. PageScript can validate and render cited evidence today, but it does not currently scan an arbitrary codebase and produce that evidence on its own.

Why the active compiler moved to Rust

PageScript started with a TypeScript reference implementation. I later chose Rust for the active compiler because the project had grown beyond syntax experiments. I wanted explicit types across the compiler stages, exhaustive matching in the validator and renderer, and a native CLI that could ship as a single binary.

The Rust crate exposes the parser, validator, resolver, IR compiler, renderer, evidence types, digest functions, and token measurement as a library. The CLI uses the same code for validation, AST and IR inspection, rendering, statistics, evidence validation, and explainers.

CI runs formatting, Clippy with warnings denied, tests, release builds, conformance fixtures, JSON Schema checks, package installation smoke tests, documentation generation, and dependency auditing. Tagged releases produce checksummed binaries for macOS, Linux, and Windows.

None of that changes the syntax, but it makes the alpha installable and testable outside my machine.

Current alpha

PageScript is currently at v1.1.0-alpha.1 and implements Draft 0.7 of the language.

The current release can compile compact pages, expand standard-library recipes, produce token reports, validate Evidence Bundles, bind Explainer Specs to evidence digests, and render offline source-cited HTML. The next major piece is deterministic evidence extraction.

For repository adapters, I only want to emit structural facts that can be cited to an exact source location. For dbt, the adapter should use manifest and catalog artifacts rather than trying to reconstruct lineage from presentation code. Inferred relationships can still be useful, but the report needs to show that they are inferred.

Right now, most of my PageScript work is in the validator, fixtures, and evidence path rather than the syntax. The next release depends more on proving those compiler guarantees than adding new language features.