# Semantic HTML package 1.2.1 — extraction protocol 1.2.0

## Evaluation correction — 21 September 2026

The extraction protocol remains v1.2.0; the evaluator schema is now 2 (revision 2.0.0). No extractor was rerun for this correction. All original fixture/output bytes, dates, repeatability witnesses and the unfavorable newspaper4k result are retained. The original aggregate is archived as [results-original-v1.2.0.json](results-original-v1.2.0.json). Its evaluation fields, and those embedded in raw outputs, are the legacy interpretation; use the revised [results.json](results.json) for current verdicts.

Every case now reports all seven tools. For tables, row and column relations are separate; exact labels and values must occur in the same table. A caption marked XML role=head is not counted as column headers. Trafilatura preserves explicit column headers only in the th variant, and no explicit row header in either variant. Markdown header rows are reported as markdown_header_row, with the source semantics separately recorded. html2text joins the caption to its first column label here, so the exact association is not validated. A no_output status is not an evaluated negative relation. Each verdict points to the exact raw field, output SHA-256 and excerpt.

Caveat presence is evaluated; attachment to the limited claim is **not evaluated**. Heading level is evaluated; section membership is not independently proved by these checks.

The old four evaluator counter-test entries were fixed declarations, not executed checks. The revised runner executes four mutations per language: changed value, missing caveat, swapped values and matching text outside a table. Each starts with an accepted positive control. The eight actual verdicts, altered output hashes and excerpts are in evaluator_counter_tests. These are evaluator checks on archived output, not eight new extractor runs.

To recalculate the aggregate from archived outputs after installing the pinned Node dependencies in reproduce/:

```sh
node reproduce/refresh-relations.mjs "$PWD"
```


## Licence

Original Edikka-authored material in this package is available under
[CC BY 4.0](LICENSE.md). Third-party software, dependency notices and material
outside Edikka's rights remain under their respective terms.

This additive replay keeps v1.0 and v1.1 intact and runs the same 12 controlled FR/EN fixtures through seven tools: Mozilla Readability 0.6.0, Trafilatura 2.2.0, readability-lxml 0.8.4.1, newspaper4k 0.9.3.1, jusText 3.0.2, html2text 2025.4.15, and markdownify 1.2.2.

The first five are extraction/readability tools with different output models. The last two are HTML-to-Markdown converters, not content selectors. Their results must not be merged into one league table. `expected.json` defines the information and relationships under test independently from the generated fixtures.

Run from the repository root with an isolated Python environment:

```sh
SEMANTIC_PROTOCOL_VERSION=1.2.0 \
SEMANTIC_OUTPUT_ROOT=/tmp/semantic-html-replay-1.2.1 \
TRAFILATURA_PYTHON=/path/to/python \
MULTITOOL_PYTHON=/path/to/python \
node scripts/semantic-html-extraction/run.mjs
```

Every tool is executed twice on identical fixture bytes. The outputs measure preservation, flattening or removal in these synthetic cases—not accessibility conformance, indexing, ranking, retrieval, model comprehension or citation.

## Input actually supplied to the tools

The runner does not remove metadata elements or rewrite fixtures before extraction. The archived HTML file in `fixtures/` is passed unchanged to every adapter. For Mozilla Readability, JSDOM uses `runScripts: "outside-only"` and no resource loader: scripts are not executed and remote resources are not fetched. Preventing execution or loading is not the same operation as deleting document information.

Every raw output records the input hash, the DOM witness captured before extraction, the tool configuration and its native output. This separation makes a tool loss distinguishable from a preprocessing change—which does not occur in this version.

## Reproduce outside the repository

The `reproduce/` directory is the standalone bundle: runner, Python adapters, pinned requirements, `package.json`, lockfile and instructions. Copy that directory elsewhere, run `npm ci`, create the Python environment and use the command in `reproduce/README.md`. `SEMANTIC_OUTPUT_ROOT` makes the replay write fixtures and results into the copy.

`agent-interaction-results.json` is separate evidence: one real keyboard interaction in Chrome against a local fixture. It archives source HTML, DOM before/after JavaScript and accessibility nodes before/after. It is not one of the twelve extraction documents and is not an autonomous-agent benchmark.

## Real-page addendum

`real-pages/` contains eight complete Edikka production-page snapshots archived on 11 September 2026: two articles, contact and library pages in FR/EN. All seven tools run twice on every file, without executing scripts or fetching remote resources. This convenience sample only checks that the program runs beyond synthetic fixtures; it is not Web-representative, causal or a tool ranking.

The retained replay exposed newspaper4k instability on `article-ai-en`; both differing outputs are archived. The adapter now determines language from `html[lang]` rather than from the filename.
