AI pre-training datasets and web corpora

Collect web corpora for AI pre-training and continued pre-training. abotapi helps you scope source collection for language models, from broad text coverage to a focused domain corpus. Define the source mix, languages, quality criteria and delivery format before expanding collection.

Describe your project ↗Telegram @abotapi ↗Browse existing actors

Or email abotapi@proton.me with your source URLs and required fields.

Define your pre-training corpus

Pre-training uses a corpus to learn patterns in language and content, typically through a self-supervised objective such as next-token prediction. Continued pre-training applies further corpus-based training to an existing model, often to extend its exposure to a domain or language.

Describe whether your project needs broad coverage or a domain-specific corpus. Include target languages, source types, date ranges and intended use. Collection feasibility and delivery volume are assessed against your brief; a ready-made foundation-model corpus is not assumed.

Choose a source mix and preserve document structure

Articles, reference pages, documentation and other supported long-form sources can contribute text to a corpus. Source-specific scrapers may also supply descriptions and structured records, but their usefulness depends on the corpus design and the amount of meaningful text available.

Define which domains and content types belong in the collection. Preserve document boundaries, headings and meaningful text order, and distinguish article content from navigation, boilerplate and repeated page templates. Inspect representative output before committing to a source.

Keep provenance and corpus versions

Retain source URLs, source identifiers, collection times, language and available publication dates alongside the text. Record the extraction and filtering versions so the team can trace a document through preparation.

For each corpus release, agree on a manifest that describes sources, collection dates, accepted and rejected record counts, and applied transformations. Keep raw and processed material separately where the project permits it.

Scope text quality filters and deduplication

Review extraction failures, empty or truncated documents, repeated boilerplate, language mismatches and low-information text. Define acceptance rules and inspect both accepted and rejected samples so filtering does not silently remove useful domain material.

Exact duplicate removal and near-duplicate detection address different kinds of repetition. Deduplication across sources and corpus releases helps identify copied or repeatedly collected documents. Agree on the unit of deduplication and validate its effects before scaling.

Measure usable text and corpus coverage

Page count alone does not describe a pre-training corpus. Track retained documents, text size, language distribution and source composition after cleaning. Token counts depend on the tokenizer: name it when specifying a token target and distinguish estimated raw volume from retained tokens.

Define the desired mixture across domains and languages. Review samples for overrepresented sources and coverage gaps. Custom token counting, mixture analysis and filtering should be included explicitly in the preparation scope.

Plan held-out data and contamination checks

Keep validation documents separate from the training corpus and account for duplicate or near-duplicate content across the split. Decide how related source material and later corpus refreshes should be assigned.

Define exclusions and overlap checks for the evaluation material your team uses. Corpus-level deduplication does not by itself prove that benchmarks are absent. Record the checks performed and their limits with each release.

Specify document formats and delivery

A document-level JSONL format can keep text and provenance together, one record per line. Specify field names, encoding, document boundaries and any custom sharding or compression requirements in the brief. Confirm which outputs the chosen scraper provides and which conversions need custom work.

Tokenization, sequence packing and training-loader integration depend on your model pipeline. Define whether your team needs source records, cleaned documents or additional preparation, and agree on acceptance samples before a larger delivery.

Send a pre-training data brief

Share example sources, languages, domains, desired document or token volume, freshness needs and an example output schema. Include your filtering, deduplication, provenance and evaluation-exclusion requirements.

Identify the intended use and applicable source permissions, exclusions and retention requirements. Public visibility alone does not establish training rights. Source collection, custom preparation and model training are separate deliverables, with feasibility and scope agreed for each project.

Common questions

What is a pre-training dataset?

A pre-training dataset is a corpus used to learn patterns through a training objective, commonly self-supervised learning on text. For language models, it does not usually require a human-written target answer for each document.

Can you collect data for continued pre-training?

Share your domain, languages, representative sources and corpus requirements. We can assess collection using existing scrapers or a custom pipeline, then agree which preparation steps are included.

How is pre-training data different from supervised fine-tuning data?

Pre-training uses a corpus of content. Supervised fine-tuning uses curated task examples with desired responses or labels. They need different schemas, quality checks and evaluation plans.

Are raw scraped pages ready for pre-training?

They are source material. Text extraction, filtering, deduplication, corpus review and pipeline-specific tokenization or packing may still be needed. Specify the preparation boundary in your brief.

Can you guarantee a particular token volume?

Volume depends on source coverage and retained content after filtering. Token counts also depend on the tokenizer. A representative collection is needed to assess feasibility before confirming a delivery target.

How much does a pre-training corpus cost?

Cost depends on the sources, collection volume, text extraction, preparation requirements and refresh schedule. Send a source list and acceptance criteria to discuss scope and a quote.