Build an Unsloth fine-tuning dataset with abotapi scrapers

AI datasetsBy abotapi4 min read

A scraped page is source material, not automatically a training example. Use abotapi to collect the records, then build reviewed task examples with a consistent schema and a separate evaluation set.

Two separated groups of cobalt document tiles with lime tiles representing curated examples.

Define what the model should learn

Fine-tuning changes a model’s behavior. Start with a task such as classifying product attributes, extracting information from a description, or answering in a specific format. Define the expected input and the correct response before gathering a large dataset.

If the goal is to answer questions about changing prices or fresh listings, retrieval may be a better starting point. Fine-tuning is not a dependable database of current facts. Use it when examples can teach a repeatable behavior, and keep rapidly changing information in a retrieval system.

Collect usable source material

Choose an abotapi actor that exposes the fields your task needs. Validate a small dataset first: check completeness, stable IDs, source links and whether detail collection is necessary. Export the raw records and preserve the run configuration so the material can be reproduced.

Before adding content to a training corpus, review the source’s terms, permissions, content license and your intended use. Public visibility alone does not establish permission to train. Remove personal information and material you cannot use. Keep a provenance manifest recording the source, collection date, permission basis and exclusions; keep that manifest separate from the model’s training text.

Create labels that teach the intended behavior

For supervised fine-tuning, build examples with an input and a reviewed target response. A product description paired with a verified category can teach classification; a document paired with a checked structured answer can teach extraction. Copying raw text into a chat-shaped file does not create correct supervision.

Write an annotation guide with allowed labels, ambiguity rules and rejection criteria. Review a sample across sources and categories before scaling. If a model drafts the labels, check them against the original content and have a qualified reviewer correct mistakes. Otherwise the new model may simply learn the drafting model’s errors.

The JSONL example below uses a messages-style record for a synthetic classification task. It illustrates a data shape, not a universal Unsloth recipe. Select a supported model and training notebook, then match its dataset formatting and chat template.

JSONL
{"messages":[{"role":"user","content":"Classify this product: stainless steel insulated bottle."},{"role":"assistant","content":"drinkware"}]}
Synthetic classification example; one complete JSON object per line.
Python
import json
from pathlib import Path

seen = set()
with Path("reviewed_examples.jsonl").open() as source, \
     Path("train.jsonl").open("w", encoding="utf-8") as output:
    for line in source:
        row = json.loads(line)
        text = row.get("source_text")
        label = row.get("label")
        if (row.get("reviewed") is not True or
            not isinstance(text, str) or not isinstance(label, str)):
            continue
        text, label = text.strip(), label.strip()
        if not text or not label:
            continue
        if text in seen:
            continue
        seen.add(text)
        example = {"messages": [
            {"role": "user", "content": "Classify this product: " + text},
            {"role": "assistant", "content": label},
        ]}
        output.write(json.dumps(example, ensure_ascii=False) + "\n")
Convert human-reviewed examples to messages JSONL. Input keys: source_text, label, reviewed. Split related records before running this on each split.

Split before formatting and training

Remove duplicates and near-duplicates before selecting your train, validation and test sets. Keep related records together: variants of a product, copies of an article or pages from the same document should not leak across the split. If you are measuring generalization to new suppliers, group by supplier rather than randomly splitting rows.

Keep a final test set that is not used to tune prompts, choose checkpoints or repair labels. Audit the distribution of classes, sources, languages and input lengths in each split. Check empty responses, malformed JSON, contradictory labels and inputs that exceed the selected model’s context budget.

Hugging Face Datasets can load JSONL files, while Unsloth’s documentation explains supported formatting paths. Use the current documentation and a maintained notebook for the model you choose; installation details, trainer arguments and chat templates are version-specific.

Evaluate the result against the original model

Run the same held-out examples through the base model and the fine-tuned model. For classification, inspect per-class precision and recall, not only overall accuracy. For extraction, measure schema validity and field-level correctness. Also test ambiguous inputs and cases that should produce an “unknown” answer.

Document the dataset version, preprocessing rules, model version and evaluation conditions. Report measured outcomes only after evaluation; more rows do not guarantee a better model. Start with a small reviewed batch, learn which errors remain, and add examples that address those errors.

Put the guide into practice

Choose a scraper for your source, inspect a small dataset, and agree on the checks before scaling.

Find your scraper ← Back to all articles

Keep reading