Choose sources for the questions you need to answer
Start with your users’ questions. Identify the pages or records that contain the evidence needed to answer them, then choose a source-specific actor or discuss a custom collection.
Preserve context with the content
Keep the source URL, relevant title and available publication or collection timestamps. Retain meaningful structure such as headings, product attributes or listing fields.
Avoid assuming that every page should become one large text block. Structured fields are useful for filters and exact comparisons; longer text can support retrieval and summarization.
Plan freshness and deduplication
Decide how often the source changes and how quickly your application needs updates. Retain stable identifiers where possible so repeated collections can be reconciled.
Embedding, chunking, indexing and retrieval are downstream design decisions. Define their requirements before settling on the scraper’s output contract.
Check quality before connecting the model
Inspect representative records for missing content, unrelated navigation text and lost context. Test your target questions against the collected information.
Keep links to the original sources in your application so people can inspect the evidence behind an answer.
Common questions
Does scraping automatically create a knowledge base?
No. Collection produces source data. Chunking, embedding, indexing and retrieval need to be designed for your application.
Can existing actors supply data for AI?
Yes, when their source coverage and output match your requirements. Review the documented fields and inspect a sample run first.
What should I send for a custom AI data feed?
Send source URLs, the questions your application should answer, required fields and your freshness requirements.