Define the behavior you want to teach
Write down the input your model will receive and the output you want: a product category, a set of extracted attributes, a structured answer or a response in a consistent style. Include difficult cases such as incomplete source information and ambiguous labels.
Create a small reference set before scaling collection. A few reviewed examples can reveal unclear instructions, inconsistent labels or output fields that the source material cannot support.
Choose an example structure
For classification, pair input text with a label from a defined taxonomy. For extraction, pair source text with the expected structured fields. Instruction-response examples include the task and a reviewed target answer; conversational examples may include several turns.
Keep source metadata alongside your examples for traceability, and map only the fields accepted by your training system into its upload format. JSONL is a container format, not a universal fine-tuning schema: confirm the exact message roles, fields and size limits required by your chosen system.
Ground target answers in the collected evidence
Scraped pages supply material; they do not automatically supply correct target answers. Define how answers or labels are produced, who reviews them and how disagreements are resolved. If an answer cannot be supported by the record, decide whether to reject the example or teach an explicit missing-information response.
For a product extraction task, the source might provide a title, price and currency. The target should include only supported attributes and follow the agreed rules for absent values. Inventing a missing brand or specification creates a misleading training example.
Review consistency before increasing volume
Use a written labeling guide with examples and edge cases. Check that the same input patterns receive compatible labels and that response formatting is consistent. Separate source collection from annotation and human review when estimating scope.
If synthetic examples are part of your process, track their origin and verify their targets. Automatically generated responses can repeat errors or reduce variety; volume should follow an accepted quality sample.
Keep evaluation examples independent
Deduplicate before assigning examples to training, validation and test sets. Group related products, documents or conversation variants where appropriate so the evaluation set does not simply repeat training content.
Evaluate the existing model on held-out examples first, then compare the fine-tuned model on the same task criteria. Track dataset versions, labeling changes and failure cases. A successful upload or training job does not by itself demonstrate better behavior.
Choose fine-tuning, RAG or both
Fine-tuning targets learned behavior such as task execution, response structure or style. Retrieval-augmented generation (RAG) supplies external material when a question is answered, which can be useful for changing facts and source-linked responses.
A product assistant may retrieve current prices while using a fine-tuned model for a consistent extraction format. Decide whether your bottleneck is missing information, inconsistent behavior or both before selecting the dataset workflow.
Send a fine-tuning data brief
Include your task, intended training system, source URLs, input and target-output examples, languages and desired volume. Specify the output schema, labeling guide, review expectations and train/validation/test requirements.
Collection, conversion, labeling and model training are distinct pieces of work. We assess source feasibility and agree the deliverables before confirming scope; a published scraper run does not include a trained model or an automatic annotation service.
Common questions
Can raw web text be used directly for supervised fine-tuning?
Supervised fine-tuning needs task-appropriate inputs and target outputs. Raw text can be source material, but usually needs example construction and review before it becomes a supervised dataset.
Can I request JSONL for an LLM fine-tuning dataset?
Include JSONL and a valid sample of your target schema in the brief. Confirm any custom conversion requirements; different training systems accept different fields and example structures.
How many fine-tuning examples do I need?
There is no single useful number for every task and model. Start with reviewed examples and a held-out evaluation set, then use measured errors and coverage gaps to guide expansion.
Does this include training or hosting a model?
The website offers data collection and scoped custom data work. Model training and hosting are not included by default; state any additional requirements when describing your project.
Should I choose fine-tuning or a retrieval dataset?
Choose based on the problem. Retrieval supplies relevant external information at answer time. Fine-tuning is a way to adjust learned behavior. Some applications use both and need separate datasets for each.