Run a web scraper on a schedule, built by your AI agent

Have your AI agent build a web scraper that runs on a timetable, saves what it finds to a database and warns you when it breaks, running 24/7 on aidrop.it.

A web scraper is a small program that visits web pages and copies out the facts you care about: prices, job listings, new posts. You can have one that runs on a timetable, keeps what it finds in a database and tells you when something breaks. Your AI agent, such as Claude Code, Codex or Cursor, writes it and puts it online on aidrop.it from one paragraph you paste. Plan on about 30 minutes.

What you will have at the end

  • A scraper that checks your chosen pages on schedule, around the clock, with your laptop closed.
  • Everything it collects, saved in a database with the time and the page it came from. Ask your agent for it as a spreadsheet file whenever you like.
  • A message to you when a run fails, or when a page suddenly yields nothing.
  • The code in a repository, so your agent can change what it collects later.

Words you will meet

  • Scraper — a program that reads web pages for you and writes down what it finds.
  • Headless browser — a real web browser running without a screen, for pages that only show their content after their scripts run.
  • robots.txt — a public file where a website says which pages automated visitors should leave alone.
  • Database — an organised store for the results. aidrop.it runs one next to your scraper.
  • Secret — a private value, like an API key, saved in your aidrop.it project and handed to the scraper when it runs.

Two ways to read a page

Most pages can be read the simple way: the scraper downloads the page and picks out the parts it needs. This is fast and fits the smallest size aidrop.it offers.

Some pages are nearly empty until the browser runs their scripts, as on many shops and dashboards. For those, the scraper drives a headless browser. It works, but it is slower and needs much more memory: plan for the 1 GB size, which starts at the Pro plan. Your agent tries the simple way first and switches only if the page comes back empty. Playwright’s own team advises not to point its browser image at untrusted sites, so keep the browser for sites you know.

The timetable works as on the cron job page: the scraper is one always-on app with its own clock. A run due during a restart is caught up when the app starts again.

Scrape politely

  • Read the site’s robots.txt and its terms. If the site offers an official API or data export, use that instead.
  • Go slow. Wait a few seconds between pages and fetch only the pages you need.
  • Say who you are. The scraper should name you and a contact address in every request.
  • Skip personal data you do not need. Privacy laws may apply to it.
  • If a site blocks you or shows a “prove you are human” check, stop. Do not try to get around it.

aidrop.it logs where each of your app’s outgoing requests goes.

What you need

  • An aidrop.it account with your agent connected. The Quickstart shows how.
  • A plan with a database: Trial, or Pro and above. Starter has none. Pages that need a headless browser need Pro or above for the larger size. Prices are on aidrop.it.
  • The pages to watch, what to collect from each, and how often.
  • Somewhere for alerts: an email service such as Resend, or a Slack incoming webhook.

Say this to your agent

Say this to your agent
Build me a web scraper that runs on a schedule on aidrop.it.
Use my existing aidrop.it Project, or create one called
"Scraper" if I have none.
What to collect: <pages and fields, e.g. "name and price of
each product on https://example.com/sale">.
How often: <e.g. "every day at 7 am London time">.
- Python. Plain HTTP first (httpx with selectolax or
BeautifulSoup). Use Playwright with headless Chromium only if
the data is missing from the raw HTML. Then use the official
Playwright image pinned to the library's version, run it as
its non-root user, and pass the HTTPS_PROXY value to
Playwright's proxy launch option.
- Read aidrop.it's runtime rules first and follow them. One
always-on app with an in-process scheduler (APScheduler) and
a small HTTP health endpoint. Schedule in my time zone.
- Be polite: obey robots.txt, a few seconds between requests,
a User-Agent with my name and a contact address. Stop on 403,
429 or a CAPTCHA and alert me; never work around them.
- Save results in the Project's PostgreSQL with the time and
source URL, plus one row per run (status, items, error).
On startup, run once if a scheduled run was missed.
- Alert me over HTTPS (email API or Slack webhook) when a run
fails or finds nothing. Keys are Project secrets: tell me the
names to save and wait until I confirm.
Build it, check it runs, do one test run, then show me the
first results and when the next run is due.

What your agent does

  1. Picks your project, checks for a database, and asks you before creating one.
  2. Looks at your pages and picks the simple way or the headless browser.
  3. Tells you which secrets to save, and waits.
  4. Writes the scraper, its timetable, the results table and a small “I’m alive” page that aidrop.it checks.
  5. Builds it on aidrop.it and checks it runs. If not, it reads the error, fixes it and builds again.
  6. Does one test run and shows you the first results.

Check that it works

  • The first results match what you see on the site yourself.
  • The next run time your agent reports matches what you asked for.
  • After the next scheduled run, ask your agent “show me the last runs”. Each says how many items it found.
  • Ask for the results as a spreadsheet file. It opens in Excel or Google Sheets.

When something goes wrong

What you see What to tell your agent
Runs finish, but find nothing the page clearly shows “This page fills in with scripts. Use the headless browser for it.”
The site answers “Access denied” or asks you to prove you are human The site does not want automated visitors. Ask your agent whether it has an official API; if not, stop.
The values are wrong after the site’s redesign “The page layout changed. Update what the scraper looks for.”
The browser crashes or runs out of memory “Move to the next size up and open one page at a time.”
aidrop.it says the app is not running “Add a small ‘I’m alive’ page that answers quickly, and build again.”

FAQ

It depends on where you are, what you collect and the site’s terms. Collecting public facts such as prices is common. Trouble usually starts with copying copyrighted content, collecting personal data, or getting around logins and blocks. Read the site’s terms first, prefer an official API, and ask a lawyer when the stakes are real.

What is the best language for a web scraper?

Python is the most common choice. It has simple tools for picking parts out of a page, such as Beautiful Soup, and Playwright for pages that need a real browser. JavaScript works as well. When an agent writes the code, the language matters less than scraping politely.

Can AI build a web scraper for me?

Yes. An AI coding agent can write the scraper, test it on your pages and fix it when a site changes its layout. On aidrop.it the same agent also puts it online, sets up the database next to it and checks it runs, so the scraper keeps working after you close the chat.

Search the docs