Build a source-linked RAG dataset from web content

Retrieval & agentsBy abotapi4 min read

Retrieval-augmented generation starts with a trustworthy corpus. Collect permitted content, preserve its source and date, and test whether your system retrieves evidence that actually answers the question.

Layered cobalt documents with a highlighted lime page and matching source marker.

Choose retrieval when facts need to stay current

RAG retrieves relevant material at answer time and supplies it to the model as context. It is useful for changing documentation, product information or market knowledge that should remain connected to a source. Fine-tuning can teach response behavior, but it does not provide a dependable update mechanism for fresh facts.

Define the questions the system should answer and the sources allowed to support them. Start with a narrow corpus and a small evaluation set. A large scrape of unrelated content can make retrieval noisier and harder to diagnose.

Collect content and preserve provenance

Use an abotapi actor whose documented output includes the text and source fields you need. Inspect a sample for empty bodies, duplicate pages, truncated content and navigation text. For an uncovered source or a specialized document shape, agree on the schema and quality criteria for a custom collection.

Keep a stable document ID, canonical source URL, title, collection timestamp and any source-provided publication date. Collection time and publication time describe different events; do not substitute one for the other. Review permissions for collection and reuse, and omit personal or confidential material that does not belong in the corpus.

Treat scraped text as untrusted input. A page can contain instructions that attempt to redirect your assistant. Keep source content separate from system instructions, enforce the application’s tool permissions, and test that retrieved text cannot authorize actions.

JSON
{
  "document_id": "guide-returns-policy",
  "source_url": "https://example.org/help/returns",
  "title": "Returns policy",
  "collected_at": "2026-10-08T00:00:00Z",
  "published_at": null,
  "text": "Reviewed document text goes here."
}
Illustrative document envelope; timestamps and values are examples.

Clean and chunk by meaning

Remove repeated menus, cookie notices and empty text, then deduplicate documents. Preserve headings and useful lists or tables. Keep an original copy alongside the cleaned version so you can investigate information lost during processing.

Split at logical boundaries such as sections before resorting to fixed-size windows. There is no universally correct chunk length: experiment with your document types, embedding model and evaluation questions. Each chunk should retain the document ID, source URL and a section reference so the final answer can cite its evidence.

Avoid placing unrelated facts together just to fill a token budget. Small chunks can lose context, while large chunks can dilute the match. Evaluate both retrieval accuracy and whether the retrieved text contains enough context for a correct answer.

Python
import hashlib
import json
from pathlib import Path

with Path("documents.jsonl").open() as source, \
     Path("chunks.jsonl").open("w", encoding="utf-8") as output:
    for line in source:
        doc = json.loads(line)
        for paragraph in doc["text"].split("\n\n"):
            text = paragraph.strip()
            if not text:
                continue
            digest = hashlib.sha256(text.encode()).hexdigest()[:16]
            chunk = {
                "chunk_id": doc["document_id"] + ":" + digest,
                "document_id": doc["document_id"],
                "source_url": doc["source_url"],
                "collected_at": doc["collected_at"],
                "text": text,
            }
            output.write(json.dumps(chunk, ensure_ascii=False) + "\n")
Create source-linked paragraph chunks from documents.jsonl. This simple character-based baseline needs evaluation; production token limits depend on your model.

Index, retrieve and refresh deliberately

Sentence Transformers provides embedding and retrieval tooling; pgvector supports vector similarity search in Postgres. These are possible building blocks, not mandatory dependencies. Choose based on your deployment, model requirements and measured retrieval quality.

Record the embedding model and version, normalization rules and chunking settings. Reindex consistently when those settings change. Use stable IDs so refreshes replace old chunks instead of appending duplicates, and delete obsolete chunks when a document is confirmed removed.

Where incremental collection is supported, it can help identify changed documents. However, a failed or partial crawl should not delete an entire source from your index. Promote a new corpus version only after coverage and quality checks pass, and preserve the last known good version for recovery.

Evaluate grounded answers and abstention

Create questions with known supporting documents, including questions the corpus cannot answer. Measure whether the relevant evidence appears among retrieved results before judging the language model’s response. Inspect wrong answers to distinguish a collection problem, a retrieval miss and an unsupported generation.

Check citation accuracy: the linked page must support the statement, not merely discuss a similar topic. Test expired facts, conflicting documents and malicious instructions embedded in source text. A useful assistant should say when the evidence is insufficient instead of inventing an answer.

Keep the collection and index versions with your evaluation results. This makes regressions reproducible and helps you decide whether to improve sources, chunking, retrieval or prompts. Begin with a modest corpus whose quality you can inspect, then expand when evidence supports it.

Put the guide into practice

Choose a scraper for your source, inspect a small dataset, and agree on the checks before scaling.

Find your scraper ← Back to all articles

Keep reading