Docs · Concepts

Recipe reference

A recipe is the small, declarative description of how to read a source: which repeated thing is a record, and which values to take from each one. The console's builder writes it for you from a live preview; this page is for checking what it wrote, or for writing one by hand for the API. Recipes are versioned: changing one creates a new version, and every run keeps the exact recipe it used.

Recipes are deliberately limited: no code, no logins, no form submissions or request headers. That keeps every run reproducible and keeps SourceFinch to public data (see rights).

Keys

KeyWhat it does
formatWhat the source returns: html (default), json, xml or csv. A run whose evidence turns out to be another type fails rather than guessing.
itemsSelects one node per record: a CSS selector for HTML and XML, a JSON path such as results[*] for JSON, or * for every CSV data row.
fieldsOutput field name → selector or path, relative to each item. 1 to 12 fields; names start with a letter and use letters, digits and underscores (max 40).
requiredFields that must be filled in every record. A run with one missing is marked degraded instead of becoming your current data.
typesOutput type per field: string, integer, number, boolean, date or datetime. Values that do not convert become null.
paginationFetch 2–5 pages by query parameter (kind: param) or by following a next-page link on the same site (kind: next). HTTP only.
csvDelimiter (comma, tab, semicolon or pipe), whether the first row is a header, and the text encoding. CSV recipes only.
modehttp (default) fetches the URL. browser renders it with scripts enabled, for pages built in the browser; Pro and above.
waitForBrowser mode only: a CSS selector to wait for before the page is captured.

HTML pages

items is a CSS selector; each field is a CSS selector inside that item. Text is read by default; add @attribute to read an attribute instead, for example a@href for a link.

{
  "format": "html",
  "items": "table.permits tbody tr",
  "fields": {
    "permit_id": "td:nth-child(1)",
    "status": "td:nth-child(2)",
    "address": "td:nth-child(3)",
    "detail_url": "td:nth-child(1) a@href"
  },
  "required": [
    "permit_id",
    "status"
  ]
}

JSON APIs

items is a path to the array of records, with [*] where an array fans out. Fields are paths inside each record: title, address.street, agencies[0].name. A leading $. is optional.

{
  "format": "json",
  "items": "$.results[*]",
  "fields": {
    "id": "$.id",
    "title": "$.title",
    "published": "$.published_at"
  },
  "types": {
    "published": "datetime"
  }
}

XML and RSS/Atom feeds

Same as HTML: CSS selectors over the XML elements, for example items: "item" and title: "title", link: "link" for RSS.

CSV and TSV files

Every data row is a record (items: "*"). Fields name a header column, or use #n (1-based) for files without a header row. Large files are streamed; one run can hold up to 10,000 records.

{
  "format": "csv",
  "items": "*",
  "fields": {
    "licence": "Licence Number",
    "holder": "Holder Name",
    "expires": "Expiry Date"
  },
  "csv": {
    "delimiter": ",",
    "header": true,
    "encoding": "utf-8"
  },
  "types": {
    "expires": "date"
  }
}

Several pages

{
  "format": "html",
  "items": "ul.notices > li",
  "fields": {
    "title": "h3",
    "posted": "time@datetime",
    "link": "a@href"
  },
  "pagination": {
    "kind": "param",
    "param": "page",
    "start": 1,
    "step": 1,
    "maxPages": 3
  }
}

param pagination sets a query parameter (or fills a {page} placeholder in the URL) for each page; next follows a next-page link matched by a selector, on the same site.

Pages built in the browser

Some pages arrive empty and fill themselves with scripts. Browser mode renders them in an isolated, network-restricted Chromium before extracting (Pro and above). Try HTTP first: most pages, feeds and APIs do not need it.

{
  "mode": "browser",
  "format": "html",
  "items": "[data-row]",
  "fields": {
    "name": "[data-name]",
    "price": "[data-price]"
  },
  "waitFor": "[data-row]"
}

Record identity and changes

Choose a key field (the record's ID, such as a permit or filing number) when you create a source. With one, an edited record shows as a field-level change (status: Under review → Approved); without one, records are compared by their whole content, so an edit appears as one removed and one added record.

When a source changes shape

If a run validates but looks broken compared with recent healthy runs (far fewer records, a usually-filled field now mostly empty, or a required field missing), it is marked degraded. Its evidence is kept, but it does not replace your current records or count as a change until an editor accepts it as correct.