Konabayev
DataWeb ScrapingContent ExtractionQuality Assurance

Article Extraction Test Dataset: 24 HTML Acceptance Cases

·7 min read
Last updated on
HTML input, expected article blocks, and an acceptance test report

Direct Answer

This article extraction test dataset contains 24 original synthetic HTML documents and explicitly authored expected results. It helps you check whether an extraction workflow follows a documented text policy before you connect it to a research archive, content migration, or retrieval system.

The cases distinguish missing content from unwanted content. They also distinguish an empty article region from a document that explicitly requests a rendered observation. No third-party extractor has been run or ranked using this release. A successful result means agreement with these fixtures, not proven accuracy across the web.

Download: JSON dataset, CSV, JSONL, and canonical generator and comparison function.

Release fieldScope
Version1.0.0; downloadable source released alongside the data
Inputs24 authored synthetic cases; no representative sample
ResultsExpected labels only; no vendor performance measured
ComparatorExact pass, mismatch or missing; diagnosis is manual

Download and use these files locally for your own acceptance tests, and cite the dataset and version. No open redistribution license is asserted for this fixture release. Keep the versioned source download with the generated files when reproducing this release. Before running it, compare the downloaded file’s SHA-256 with reproduction.source_sha256 in the version 1.0.0 JSON dataset. For example, shasum -a 256 generator.mjs prints the local file hash on macOS. Stop if the hashes differ, and retain the matching source and dataset together.

What the Dataset Tests

Each case isolates a small decision that an extraction contract needs to make. The HTML is written for this dataset. It is not copied from a publisher or sampled from a customer site.

CasesDecision being checkedExample expected behavior
EXTRACT-001 to 005Article scope and surrounding contentExclude navigation, aside and footer under the declared policy
EXTRACT-006 to 010Inline text, spaces and entitiesKeep link text and decode &
EXTRACT-011 to 014Hidden content and non-prose nodesExclude hidden descendants, scripts, styles and comments
EXTRACT-015 to 018Structural boundariesPreserve headings, list items, table cells and quoted paragraphs
EXTRACT-019 to 022Special text formsRetain code indentation, captions, details text and Unicode
EXTRACT-023 to 024Incomplete outcomesSeparate empty from render_required

Twenty-four is the size of this authored release, not a statistically representative sample. The value is the inspectable distinction between cases. Adding twenty near-identical paragraphs would increase the row count without testing an additional decision.

Define the Text Contract First

The expected output is ordered static article text, not every visible pixel or every DOM text node. Version 1 selects one main element and applies the rules embedded in the JSON file.

Headings, paragraphs, list items, table rows, preformatted blocks, captions and summaries become ordered blocks. Heading text stays in source order; the output does not retain heading-level metadata such as H1 versus H2. Ordinary whitespace is collapsed inside prose. Preformatted text retains its line breaks and indentation. Table cells are joined with a literal | so that “Plan” and “Units” do not silently become an indistinguishable string.

The policy excludes navigation, asides, footers, scripts, styles, comments, hidden descendants and aria-hidden="true" descendants. It keeps link text but does not emit link destinations or image alt text. Static details content is included even when the element is initially closed.

These are deliberate application choices. A migration that needs image alternatives or editorial sidebars should adopt a different contract and update the expected results before testing. Do not mark a tool incorrect solely because it follows a different, documented output specification.

Why Plain DOM Text Is Not the Same Test

An unfiltered textContent call does not implement this article policy. MDN’s documentation for Node.textContent explains that it includes the text of script and style elements, while innerText considers styling and differs in its treatment of hidden elements.

That distinction explains why the dataset specifies exclusions rather than promising that one browser property returns “the article.” The fixtures do not include external CSS, a rendering engine, shadow DOM or network-loaded content. A hidden class without a hidden attribute is therefore not evaluated as a computed-visibility rule.

The source was read on October 4, 2026 for these DOM distinctions. MDN does not supply or endorse our expected labels. Our choice to omit an aside, preserve a caption, or join cells with vertical bars remains an authored acceptance rule.

For a different evaluation scope, the Scrapinghub article extraction benchmark publishes ground-truth article bodies, system outputs and evaluation scripts. Its README, read on the same date, describes precision, recall, F1 and accuracy evaluation. Our fixtures isolate small authored contract cases and include no vendor predictions. Use them for explainable regressions; use an independently selected article corpus when testing whether an extractor generalizes. We have not reproduced that repository’s reported results.

For the broader collection workflow, see the article extractor guide and the web scraping platform guide. Product selection and fixture acceptance are separate decisions.

Read and Reproduce the Files

JSON is the complete contract; CSV and JSONL carry the same 24 case records. Each record contains a stable fixture_id, the input HTML and policy name, an expected result, rule IDs and a rationale.

The CSV stores input, expected and rule_ids as JSON inside quoted cells. Use a CSV parser first, then parse those three cells as JSON. Splitting a CSV row on commas will corrupt both HTML and nested values. The generator’s tests round-trip these fields, including quotes and newlines.

Download the version 1.0.0 generator linked above, save the text file as generator.mjs, and run these commands with Node.js 22 or later. No repository access or package installation is required:

node ./generator.mjs --out ./fixture-data
node ./generator.mjs --out ./fixture-data --check

That check compares local files with this source contract. To verify the published downloads instead, put all nine dataset files in a separate ./downloaded folder and run node ./generator.mjs --out ./downloaded --check. Expect “Verified 9 files”; this checks published byte parity, not an extractor’s accuracy.

This reproduces all three related fixture datasets. The script uses the standard library and makes no network requests. It serializes authored inputs and oracles; it is not an extraction implementation. Retain the downloaded version 1.0.0 source with the generated files so a future release does not silently change your baseline.

Connect an Extractor Without Hiding Failures

Write a small adapter that turns your actual output into the declared fields. For extraction, submit status, blocks and text under the matching fixture ID. The expected must_not_include list is an oracle constraint, not a field your extractor is supposed to generate.

The exported compareResults(dataset, results) function compares submitted outputs with the authored contract. Missing cases remain missing. Unknown or duplicate IDs reject the supplied run. An extra success row cannot compensate for an omitted difficult case.

Keep the raw extractor response alongside the adapter output. Record the extractor version, options, runtime, adapter revision and fixture dataset version. If an adapter normalizes whitespace, document that transformation. It must not copy the expected text or delete unexpected content solely to achieve a passing score.

For example, EXTRACT-017 expects two table blocks: Plan | Units and Example | 12. Outputting Plan Units Example 12 retains words but loses the contract’s cell and row structure. EXTRACT-019 checks a different issue: removing indentation from its preformatted second line fails even if the prose normalizer normally collapses spaces.

Interpret the Results by Failure Type

Use failures to identify a decision or defect, not to manufacture a product leaderboard. Manually inspect content loss, contamination, structural loss and unresolved-state handling separately. compareResults emits only pass, mismatch or missing; it does not automatically diagnose those failure categories.

EXTRACT-002 tests whether a site footer leaks into article text. EXTRACT-020 tests whether the caption survives while the image alternative remains outside the selected contract. EXTRACT-024 contains an explicit synthetic data-render-required marker. It does not demonstrate that a product can recognize every JavaScript loading state.

If you change the rules, publish a new version and preserve old case IDs where their meaning remains unchanged. Retire an ID rather than quietly reusing it for an unrelated test. A regression report should identify the changed cases and their consequences for the downstream task.

When extracted text feeds research, apply the citable data asset workflow separately. Passing these tests does not prove that a claim is true, a source is current, or its reuse is permitted.

FAQ

Is this a benchmark of commercial article extractors?

No. It is an original acceptance fixture dataset. We have not published measured vendor results for it. A reader can run an explicitly documented comparison, but the resulting claim must stay scoped to these cases, the chosen policy and the tested versions.

Does passing all 24 cases prove production reliability?

No. The corpus omits malformed real-world layouts, pagination, authentication, browser execution and many language or layout combinations. Add representative permitted documents from your own workload and independently review their expected outputs.

Can I use another definition of article text?

Yes. Version your rule changes and expected outputs together. If captions, links, image alternatives or sidebars matter to your downstream workflow, include them explicitly. Do not compare scores from different definitions as though they measured the same outcome.

Last checked: October 4, 2026. Fixture rules and deterministic export checks reviewed; MDN read for DOM methodology. No third-party extraction performance was measured.

Web Scraping Automation

Apify scrapers, data extraction pipelines, and scheduled monitoring workflows.

View service details

Google Preferred Sources

See more of my research in Google

Add Konabayev.com as a preferred source to find more fresh marketing and AI research in Google Search.

Have a relevant product or documented use case? View sponsorship formats and editorial requirements.