Konabayev
Web ScrapingContent ExtractionAI ToolsData Pipeline

Website Content for RAG: Extraction, Provenance and QA

·11 min read
Last updated on
Website content preparation for RAG with source provenance and extraction checks

Apify Actors linked under my tugelbay account are my products. I earn from their usage and may also receive a referral commission through Apify links in this article.

How Do You Prepare Website Content for RAG?

Capture a permitted source, preserve its qualifications, validate the extracted document, and only then make it searchable. A successful HTTP response or a readable paragraph does not establish that your retrieval-augmented generation system has the information needed to answer a question correctly.

This guide follows one completed, fictional policy-document example from extraction candidate to accepted retrieval record. It includes a small executable validator, a rejected candidate and a source-linked answer. The example is authored teaching material, not a test of an extraction service or a measured RAG benchmark. The companion article extraction test dataset provides reusable cases for evaluating your own implementation.

The practical deliverable is an ingestion acceptance record: what was requested, what survived extraction, what failed review and which source version supports the answer. You can implement this before buying a vector database or adding an autonomous agent.

Start With Questions and a Source Boundary

Choose the questions the collection must answer before choosing the extractor. For an internal product-help collection, the first questions might concern supported configurations, export limits and escalation paths. A marketing research collection needs different evidence, dates and permissions. Combining these collections without explicit boundaries makes it harder to explain why an answer was produced.

Write an inclusion rule specific enough that two reviewers could apply it. For example: “Include the approved fictional export-policy page and its linked definitions; exclude customer-specific contracts and archived editions.” Record exclusions as deliberate decisions rather than treating every inaccessible URL as a missing document that must be recovered.

Maintain the requested URL separately from the final resolved URL. A redirect can be legitimate, but a login page, category listing or unrelated replacement should not silently inherit the requested document’s identity. If the response does not contain the required page, its status is unavailable or rejected even when the transport succeeded.

For public sources, check the permitted access route, usage terms and intended retention. Public visibility is not permission to republish a complete article or process personal information for an unrelated purpose. For internal sources, retrieval and answer delivery must respect the same access boundary. Do not put restricted content into a shared index and rely on the language model to remember who may see it.

An initial collection manifest can stay small:

FieldPurposeExample value
document_idStable identity within your collectionfixture-export-policy
requested_urlSource requested by the collectorhttps://example.invalid/policies/export
source_kindDistinguish fixture from collected evidenceauthored_fixture
required_unitsContent that must survive extractionEligibility, limit table, exception
access_classRetrieval and answer access boundarypublic_fixture
publication_dateDate asserted by the publishernull when unknown
retrieved_atActual collection time, if collectednull for this uncollected fixture

The example URL intentionally does not represent a live vendor page. Do not send it to a paid collector expecting a real result.

Separate Capture, Extraction and Acceptance

Keep acquisition success, parsing success and content acceptance as different outcomes. They answer different questions and need different remedies.

Capture asks whether the approved route returned the intended resource. Extraction asks whether the parser identified and represented the required material. Acceptance asks whether the result meets the collection’s declared requirements. A record can pass the first two steps and still fail the third because it lost the exception beneath a table.

For a concrete format reference, Firecrawl’s scrape documentation shows a response with Markdown, HTML and metadata. Those returned representations still need your acceptance checks. The documentation describes the interface; it does not demonstrate extraction completeness on the fixture in this guide.

Use explicit states such as accepted, partial, unavailable and rejected. Define their meaning locally. For example, partial might mean that the introduction is available but an essential table is absent, so the document cannot answer a limit question. Do not allow partial records into the same retrieval pool as accepted records without a visible restriction.

Retain enough evidence to investigate a failure: response metadata, permitted source snapshot or fingerprint, parser version, transformation settings and validation result. Retention should follow your access and storage policy. A hash can help distinguish bytes but does not replace the document, prove its truth or grant reuse rights.

The Article Extractor guide explains checks for one possible collection route. My Article Extractor on Apify is a candidate to evaluate, not an extractor shown to pass this example. It is my product; I earn from usage and may receive a referral commission through the link. Inspect its current contract before a paid run. A local parser or existing approved export may already satisfy the task.

A Completed Example: Preserve the Exception

The accepted document must retain both the general rule and the condition that changes it. Consider this entirely fictional source, authored for the tutorial:

Export is available for approved workspaces. Standard exports contain up to 500 records. Approval alone does not enable exports for restricted workspaces. A restricted workspace requires a separate administrator review before any export.

The source also has a two-row table: standard workspace, 500-record limit; restricted workspace, export blocked pending administrator review. These are invented product rules, not facts about an existing service.

Our intended question is: “An approved workspace is restricted. Can its user export 500 records?” The correct fixture answer is no, not until the separate administrator review. Merely recovering the first two sentences would produce a dangerously incomplete record for that question.

CandidateRetained contentAcceptance decision
AEligibility sentence and 500-record limitReject for this question: restricted-workspace exception absent
BEligibility, table and exception, each with a source fragmentAccept against the declared fixture requirements

Candidate B’s useful output is not just one long text string. It has stable content units, their source locations and the relationship between the general rule and its exception. Candidate A remains in the run ledger so that a report cannot quietly calculate success using only surviving records.

This is a completed editorial evaluation of two supplied candidates. It does not establish that any real parser generated them. When evaluating a service, replace these authored outputs with saved responses and preserve all requested sources in your denominator.

Run a Small Acceptance Check Locally

Use mechanical checks for known structural requirements and human review for meaning. The following self-contained JavaScript example checks our fixture’s required units and whether the units carry source fragments. Save it as rag-fixture-check.mjs and run it with Node.js. It makes no network calls, requires no account and sends no data to a model.

import assert from 'node:assert/strict';

const required = ['eligibility', 'limit-table', 'restriction'];
const complete = {
  documentId: 'fixture-export-policy',
  sourceKind: 'authored_fixture',
  units: [
    { id: 'eligibility', fragment: '#eligibility',
      text: 'Export is available for approved workspaces.' },
    { id: 'limit-table', fragment: '#limits',
      text: 'Standard: up to 500 records. Restricted: blocked pending review.' },
    { id: 'restriction', fragment: '#restriction',
      text: 'Restricted workspaces require a separate administrator review.' }
  ]
};

function validate(record) {
  const units = Array.isArray(record.units) ? record.units : [];
  const ids = units.map(unit => unit.id);
  const missing = required.filter(id => !ids.includes(id));
  const invalid = units.filter(unit =>
    typeof unit.text !== 'string' || !unit.text.trim() ||
    typeof unit.fragment !== 'string' || !unit.fragment.startsWith('#')
  ).map(unit => unit.id);
  const duplicateIds = ids.filter((id, index) => ids.indexOf(id) !== index);
  return {
    accepted: missing.length === 0 && invalid.length === 0 &&
      duplicateIds.length === 0,
    missing, invalid, duplicateIds
  };
}

const partial = { ...complete, units: complete.units.slice(0, 2) };
assert.equal(validate(partial).accepted, false);
assert.deepEqual(validate(partial).missing, ['restriction']);
assert.equal(validate(complete).accepted, true);
assert.equal(validate({ ...complete,
  units: [...complete.units, complete.units[0]] }).accepted, false);
console.log(JSON.stringify({ partial: validate(partial),
  complete: validate(complete) }, null, 2));

The expected result marks the partial record false with restriction missing, and the complete record true. Duplicate IDs fail as well. This validator does not independently compare the text with HTML, confirm that a fragment exists or detect an incorrect paraphrase. Those remain separate checks. An adversarial output could use every required ID while putting the wrong text inside it.

For production, derive expected units from an independently inspected evaluation sample, not from the same output being scored. If the extractor supplies its own completeness label, treat that label as a claim to test. The test dataset can exercise known cases; it cannot establish coverage of all layouts on the web.

Preserve Tables, Dates and Provenance Before Chunking

Chunk accepted meaning, including conditions and context, rather than cutting an arbitrary number of characters from an unchecked string. A useful chunk is small enough to retrieve precisely and complete enough to interpret correctly.

For the fictional export policy, keep the limit table with its title and the restriction that qualifies it. If you separate them, store relationships that allow the retrieval step to include both. Repeating a short qualification across related chunks may be preferable to presenting an isolated number with no condition. Measure the trade-off on your questions instead of announcing a universal optimal chunk size.

Preserve table headers when representing rows as text. “500” alone loses the unit, affected population and rule. A faithful representation of our fixture is “Standard workspace exports contain up to 500 records.” If the source does not say “per export,” do not add that denominator merely to make the sentence clearer. For our fixture, the authored wording is the authority; real source wording needs the same discipline.

Plain DOM text is not automatically the visible article. MDN’s textContent reference explains that it includes script and style text, whereas innerText is aware of styling and excludes hidden text. This distinction helps you design an adapter, but neither property alone proves that the correct article region or its footnotes were selected.

Store publication, modification and retrieval dates separately. A missing publication date remains unknown. A parser upgrade does not make a publisher’s old policy current, and today’s retrieval time is not an editorial update date. When no reliable source date is available, the answer should reflect that limitation rather than inventing freshness.

At minimum, each indexed unit should retain document identity, source locator, content version, validation state and access class. Keep original text alongside any normalized representation when feasible. That lets a reviewer inspect what changed during cleanup, translation or summarization.

Build a Retrieval Check That Can Fail

Test questions need expected evidence and an explicit insufficient-evidence outcome. Without those, a fluent answer can conceal a retrieval failure.

For the worked example, the evaluation card is:

ItemCompleted fixture value
QuestionCan an approved restricted workspace export 500 records?
Required evidenceEligibility and restricted-workspace exception
Acceptable answerNo; separate administrator review is still required
Required citationThe restriction unit, with the eligibility context
Rejection exampleYes, approval permits 500 records
Insufficient-evidence behaviorDo not infer permission when the restriction is missing

The accepted example answer is: “In this fictional policy, approval alone is insufficient for a restricted workspace. Separate administrator review is required before export. Source: fixture-export-policy, restriction.” This is an authored reference answer, not a logged model response or a retrieval success measurement.

When you run an actual system, save the question, retrieved unit IDs, generated answer, cited locations and reviewer decision. Score evidence retrieval separately from answer correctness. Correct wording with an irrelevant citation is not a fully supported answer; correct retrieval with a distorted answer needs a different fix.

Keep a held-out question set that was not used to tune cleanup and retrieval rules. Otherwise, repeated adjustments can teach the pipeline your small examples without demonstrating useful behavior on new questions. Report the sample size and source types rather than describing a pilot score as general accuracy.

Handle Updates, Failures and Source Instructions

A changed source should create a reviewable version transition, not silently replace an accepted answer basis. Preserve the previous accepted version until the new candidate passes the declared checks, and make the active version explicit in retrieval.

If a source becomes unavailable, decide whether the last accepted version may remain usable with a stale marker or must be removed. A policy that controls access or money may need a stricter expiry than background documentation. Define this with the content owner; do not present one retention interval as universally safe.

Treat instructions embedded in retrieved pages as untrusted source material. A page saying “ignore previous rules” does not gain authority over application policy, credentials, tools or publication. Keep retrieved content separate from the workflow’s instructions and restrict what any downstream agent may execute. Validation of text does not authorize actions described by that text.

Log failures by stage: denied or unavailable source, extraction omission, schema mismatch, acceptance rejection and answer failure. A catch-all error count hides the operational decision. Avoid unlimited retries: a blocked source or changed access condition needs inspection, not an increasingly expensive loop.

Choose a Tool and Define the Handoff

Compare candidates on accepted output for your permitted sample, using the same task and cost boundary. An existing export, local parser, browser-based collection route or hosted extractor may each be appropriate under different constraints.

Ask candidates to preserve the information your collection requires. Check allowed inputs, dynamic rendering needs, output semantics, source limits, retention controls and observable failure states. Verify current product capabilities in the relevant documentation rather than assuming that “RAG-ready” means your exact contract is satisfied. Octoparse’s source-corpus guide makes a useful explicit distinction between source-data delivery and customer-owned chunking, indexing and answer evaluation. It is the vendor’s description of its service boundary, not an independent performance study or evidence that either product passed our fixture.

Use cost per usable record to distinguish returned rows from accepted documents and subscription allocation from actual cash cost. Include review and correction work in your own operating estimate. This guide offers no current price comparison or guaranteed saving.

The handoff should include the source manifest, output schema, acceptance fixtures, rejection ledger, update rule and an owner for disputed records. If implementation help is useful, request a web scraping automation scope. Agree the sample and acceptance conditions before recurring collection or delivery commitments.

Frequently Asked Questions

Is Markdown enough for website RAG?

Sometimes, if it preserves the information your questions require. Check tables, headings, links and qualifications against the source. Markdown is a representation format, not a completeness guarantee. Store provenance and validation status alongside it.

Do I need a vector database immediately?

No particular storage product is required to validate extraction. Start with inspectable files and a bounded question set. Choose retrieval infrastructure after defining access controls, update behavior, search requirements and operating constraints.

Can I treat HTTP 200 as a successful ingestion?

No: it establishes a transport result, not an accepted document. A login screen or unrelated page can return successfully. Validate source identity, required content and the intended language before indexing.

What should happen when a publication date is missing?

Keep it unknown and store the retrieval date separately. Do not fill it with today’s date. If a question depends on current policy, missing source-date evidence may require manual verification or an insufficient-evidence answer.

Can an LLM repair missing source text?

It should not invent omitted evidence. Reacquire through an authorized route, inspect the original or reject the affected use. A model-generated continuation may look plausible while reversing a qualification or creating an unsupported limit.

Does this example show that Article Extractor passed a benchmark?

No: both extraction candidates and the policy are authored fixtures. The example demonstrates acceptance logic. A product evaluation needs actual saved outputs, a declared source sample, failure denominators and measured costs under the same conditions.

Last verified: October 4, 2026. The Firecrawl scrape reference, MDN textContent reference and Octoparse source-corpus guide were read for the specific method and interface statements cited above. Fixture logic is authored; current prices and extraction performance were not tested. These sources are documentation and vendor material, not three independent performance studies.

Web Scraping Automation

Apify scrapers, data extraction pipelines, and scheduled monitoring workflows.

View service details

Google Preferred Sources

See more of my research in Google

Add Konabayev.com as a preferred source to find more fresh marketing and AI research in Google Search.

Have a relevant product or documented use case? View sponsorship formats and editorial requirements.