# Observations v1.2.0

## Evaluation correction — 21 September 2026

The extraction protocol remains v1.2.0; the evaluator schema is now 2 (revision 2.0.0). No extractor was rerun for this correction. All original fixture/output bytes, dates, repeatability witnesses and the unfavorable newspaper4k result are retained. The original aggregate is archived as [results-original-v1.2.0.json](results-original-v1.2.0.json). Its evaluation fields, and those embedded in raw outputs, are the legacy interpretation; use the revised [results.json](results.json) for current verdicts.

Every case now reports all seven tools. For tables, row and column relations are separate; exact labels and values must occur in the same table. A caption marked XML role=head is not counted as column headers. Trafilatura preserves explicit column headers only in the th variant, and no explicit row header in either variant. Markdown header rows are reported as markdown_header_row, with the source semantics separately recorded. html2text joins the caption to its first column label here, so the exact association is not validated. A no_output status is not an evaluated negative relation. Each verdict points to the exact raw field, output SHA-256 and excerpt.

Caveat presence is evaluated; attachment to the limited claim is **not evaluated**. Heading level is evaluated; section membership is not independently proved by these checks.

The old four evaluator counter-test entries were fixed declarations, not executed checks. The revised runner executes four mutations per language: changed value, missing caveat, swapped values and matching text outside a table. Each starts with an accepted positive control. The eight actual verdicts, altered output hashes and excerpts are in evaluator_counter_tests. These are evaluator checks on archived output, not eight new extractor runs.

To recalculate the aggregate from archived outputs after installing the pinned Node dependencies in reproduce/:

```sh
node reproduce/refresh-relations.mjs "$PWD"
```


Scope: 12 synthetic fixtures, six paired comparisons, French and English. Translations are not independent observations.

## Seven tools, two tool classes

- Content/readability extraction: Mozilla Readability, Trafilatura, readability-lxml, newspaper4k and jusText.
- HTML-to-Markdown conversion: html2text and markdownify.

The two classes answer different questions and are not ranked together.

## Results in these fixtures

- Heading: Mozilla Readability, Trafilatura, readability-lxml, newspaper4k, html2text and markdownify distinguish the native heading from the generic element in their structured output. jusText classifies every short synthetic fixture as boilerplate and returns no selected text.
- Table: Mozilla Readability and readability-lxml preserve `th` and `scope`; Trafilatura represents column headers but loses the explicit row-header relationship. newspaper4k flattens the table. Both Markdown converters produce a Markdown table for both variants because both sources remain HTML `table` elements; they cannot recover header semantics that were absent in the `td` variant.
- Caveat: Mozilla Readability and Trafilatura drop the caveat in the `aside`. newspaper4k does too. readability-lxml and both Markdown converters preserve its text in both variants. jusText returns no selected text for either.

These observations do not reproduce RedactionSEO’s inputs: its non-semantic table variant uses `div` elements, while Edikka’s controlled pair keeps the `table` element and changes `th` to `td`. The two protocols therefore complement rather than duplicate one another.

All seven tools were executed twice on identical input bytes. The 12 combined output hashes match their replays.

## Real-page addendum

Eight complete first-party Edikka production snapshots archived on 11 September 2026 were replayed separately: two article templates, contact and library pages in French and English. This bounded convenience sample is not representative and supports no causal conclusion. Its purpose is narrower: check that the seven adapters execute on full pages beyond the controlled fixtures and retain failures or unstable results.

All seven tools produced a recorded status on all eight pages. A retained double run of `article-ai-en` differed only for newspaper4k; the first and replay outputs are both stored in `real-pages/outputs/article-ai-en.json`. The other component replays were stable in that run. This unfavorable result is evidence about one input, version and run—not a general reliability verdict.

During this addendum, the multi-tool adapter was corrected to read the language from `html[lang]`. Filename inference had selected the English jusText stoplist for some French real-page filenames. The synthetic fixtures were unaffected because their names already contained `-fr-`, but both datasets were replayed after the correction.

## JavaScript interaction addendum

`agent-interaction-results.json` records one local Chrome interaction with scripts enabled: source HTML hash, DOM before/after, accessibility nodes before/after, keyboard focus, Enter activation, `aria-expanded`, controlled-content visibility and live-status text. It verifies a deterministic browser interaction in the named environment; it is not an autonomous-agent benchmark or a screen-reader user test.
