How to monitor your market with abotapi and Python

Marketing intelligenceBy abotapi4 min read

Turn scattered product pages, reviews and search results into a repeatable market report. Start with one question, collect a consistent sample, and compare like with like before deciding what changed.

Cobalt observation tiles arranged across time with lime tiles highlighting a pattern.

Start with a decision, not a scrape

A useful marketing report answers a question someone can act on. Are competitors discounting the products you sell? Are customers mentioning a recurring delivery problem? Is a brand appearing more often in the search results you track? Choose one question before choosing a scraper.

Write down the market, location, language, competitor set and reporting cadence. Keep these settings stable between runs. A different query or country can change your sample more than any real movement in the market. Treat the first collection as a baseline, not a trend.

Choose sources and a small collection scope

Browse the abotapi actor directory for the sources your audience uses. Ecommerce pages can support price and availability comparisons; review sources can reveal recurring customer concerns; social and search sources can help you observe published content. Read each actor’s current input and output documentation before designing your report.

Run a small collection and inspect the records. Check that the identifiers, source URLs, dates and fields required for your question are present. Some fields need a detail or review option, and some sources do not publish them at all. If the documented output does not support your question, change the source or request a custom collection.

Keep a dated snapshot of every successful run

Use an Apify schedule for recurring collection after the initial run is validated. Record the run ID, dataset ID, input scope and collection timestamp alongside each export. These are the coordinates of your analysis: without them, a chart is difficult to reproduce.

Where the actor documents incremental mode, use it to observe new and changed records. Keep a full baseline too. A changed row is not necessarily a business event: a rotating image URL, collection timestamp or moving review count can introduce noise. Check the actor’s supported change controls and decide which fields matter to your report.

Fetch records using the official Python client. The following example reads an existing dataset; it does not launch a paid actor run. Keep your token in an environment variable and store exports outside version control.

Python
import os
import json
from pathlib import Path
from apify_client import ApifyClient

client = ApifyClient(os.environ["APIFY_TOKEN"])
items = client.dataset(os.environ["APIFY_DATASET_ID"]).iterate_items()
with Path("snapshot.jsonl").open("w", encoding="utf-8") as output:
    for item in items:
        output.write(json.dumps(item, ensure_ascii=False) + "\n")
Read an existing dataset. Set APIFY_TOKEN and APIFY_DATASET_ID in your environment.

Separate coverage from performance

Normalize identifiers before joining snapshots. A product URL with tracking parameters may represent the same product as yesterday’s clean URL. Resolve duplicate records, standardize currencies and time zones, and count missing values. Keep an untouched raw export so transformation mistakes can be traced.

Report coverage alongside the metric: how many sources completed, how many comparable records appeared, and how many observations were missing. A lower average price can mean discounts, or simply that the expensive items disappeared from your sample. A missing product should become “not observed” until you have evidence that it was removed.

Use pandas for grouping, joins and summaries. For review analysis, keep the quoted text and source URL behind each theme; manually check a sample of classifications. Counts of visible reviews or posts describe your collected sample, not the entire market.

Python
import json
from collections import Counter
from pathlib import Path

rows = [json.loads(line) for line in
        Path("observations.jsonl").read_text().splitlines() if line.strip()]
coverage = Counter(row.get("source", "unknown") for row in rows)
report = {
    "observations": len(rows),
    "sources": dict(coverage),
    "missing_timestamps": sum(not row.get("observed_at") for row in rows),
}
Path("coverage-report.json").write_text(json.dumps(report, indent=2))
print(json.dumps(report, indent=2))
Summarize normalized observations.jsonl. Each record needs source and observed_at; optional review_count is not summed.

Make the report useful every week

Build a compact report with the original question, collection window, sample coverage, notable changes and recommended next steps. Link observations back to source records. Prefer a few well-supported findings over a dashboard full of numbers that nobody uses.

Review the workflow after several successful collections. Adjust alert thresholds from observed variability and team needs, not an invented universal percentage. Verify the source when a large change appears. Marketing monitoring works best as a repeatable evidence trail that informs a human decision.

Put the guide into practice

Choose a scraper for your source, inspect a small dataset, and agree on the checks before scaling.

Find your scraper ← Back to all articles

Keep reading