Run · Data pipelines

Data pipelines that read messy public sources and turn them into pages that rank.

Scheduled jobs that read government feeds, portals, spreadsheets and PDFs, normalise them into a database with provenance kept on every record, and publish pages and reports from it every night. GrantWindow is built this way, and so is the publishing behind our media brands.

GrantWindow, built in this lane
Directories, registers, catalogues, daily publishingGood for
A source at a time, live within the first milestoneShape
GrantWindow, Pets Trend Store, GaraMasalaFMProof
Your server or ours, your database, your repositoryHandover
How it is built

The shape of every data pipeline project we run.

Public sourcesFeeds, portals, PDFs, sheetsYour own dropsA folder, a form, an inboxScheduled readersNightly, retried, loggedNormaliseOne shape, provenance keptDatabaseEvery record, every versionPages and APIPublished every nightDiffWhat changed since yesterdayAlerts and reportOnly when it matters
The stackScheduled timersPython and Node readersSQLite and PostgresProvenance on every rowStatic publishingDiff alertsHetznerCloudflare
What lands in your hands

Six things you get, shown on products we already run.

[01]   Data pipelinesReaders for sources that were never meant to be read

Portals without an API, PDFs, spreadsheets that change shape. Each one read on a schedule and retried when it fails.

[02]   Data pipelinesA database with a memory

Every record carries where it came from, when it was read and what it looked like last time.

[03]   Data pipelinesPages published from the data

Directories, detail pages and boards regenerated every night, thousands at a time if the data warrants it.

[04]   Data pipelinesDiffs, not dumps

What changed since yesterday, as a list a person can act on.

[05]   Data pipelinesAlerts when a source moves

A feed that changes shape or goes quiet is caught the same night, before the pages go stale.

[06]   Data pipelinesA report that says what happened

Rows read, rows changed, pages built, anything that needs a decision.

01 / 06Scroll sideways or use the arrows
Also in this lane

Where this work usually goes next.

[01]Publishing pipelines

Drafting, assembly and scheduled publishing across a site and social channels, with a person approving before anything goes out. GaraMasalaFM runs on one.

[02]Catalogues and directories

Products, stations, grants, venues: a source read nightly and a page for every item.

[03]Reporting

A morning email with the numbers that changed and nothing else.

Questions

Asked before hiring us for Data pipelines.

Where does it run?

On a server we manage, or on yours. Scheduled timers, logs kept, and a dead-man check that notices when a job did not run.

What if a source blocks automated reading?

We read from where it is allowed and say clearly where it is not. Some sources need to run from a server rather than a laptop, and we know which.

Can it feed my existing site?

Yes. The output can be pages, a JSON API or rows in your database.

Get started

Tell us what you need built.

+ One paragraph from you+ Reply within one business day, Sydney time+ Written scope and a fixed price, or an honest no