Docs · Concepts
Recipe reference
A recipe is the small, declarative description of how to read a source: which repeated thing is a record, and which values to take from each one. The console's builder writes it for you from a live preview; this page is for checking what it wrote, or for writing one by hand for the API. Recipes are versioned: changing one creates a new version, and every run keeps the exact recipe it used.
Recipes are deliberately limited: no code, no logins, no form submissions or request headers. That keeps every run reproducible and keeps SourceFinch to public data (see rights).
Keys
| Key | What it does |
|---|---|
format | What the source returns: html (default), json, xml or csv. A run whose evidence turns out to be another type fails rather than guessing. |
items | Selects one node per record: a CSS selector for HTML and XML, a JSON path such as results[*] for JSON, or * for every CSV data row. |
fields | Output field name → selector or path, relative to each item. 1 to 12 fields; names start with a letter and use letters, digits and underscores (max 40). |
required | Fields that must be filled in every record. A run with one missing is marked degraded instead of becoming your current data. |
types | Output type per field: string, integer, number, boolean, date or datetime. Values that do not convert become null. |
pagination | Fetch 2–5 pages by query parameter (kind: param) or by following a next-page link on the same site (kind: next). HTTP only. |
csv | Delimiter (comma, tab, semicolon or pipe), whether the first row is a header, and the text encoding. CSV recipes only. |
mode | http (default) fetches the URL. browser renders it with scripts enabled, for pages built in the browser; Pro and above. |
waitFor | Browser mode only: a CSS selector to wait for before the page is captured. |
HTML pages
items is a CSS selector; each field is a CSS selector inside that item. Text is read by default; add @attribute to read an attribute instead, for example a@href for a link.
{
"format": "html",
"items": "table.permits tbody tr",
"fields": {
"permit_id": "td:nth-child(1)",
"status": "td:nth-child(2)",
"address": "td:nth-child(3)",
"detail_url": "td:nth-child(1) a@href"
},
"required": [
"permit_id",
"status"
]
}JSON APIs
items is a path to the array of records, with [*] where an array fans out. Fields are paths inside each record: title, address.street, agencies[0].name. A leading $. is optional.
{
"format": "json",
"items": "$.results[*]",
"fields": {
"id": "$.id",
"title": "$.title",
"published": "$.published_at"
},
"types": {
"published": "datetime"
}
}XML and RSS/Atom feeds
Same as HTML: CSS selectors over the XML elements, for example items: "item" and title: "title", link: "link" for RSS.
CSV and TSV files
Every data row is a record (items: "*"). Fields name a header column, or use #n (1-based) for files without a header row. Large files are streamed; one run can hold up to 10,000 records.
{
"format": "csv",
"items": "*",
"fields": {
"licence": "Licence Number",
"holder": "Holder Name",
"expires": "Expiry Date"
},
"csv": {
"delimiter": ",",
"header": true,
"encoding": "utf-8"
},
"types": {
"expires": "date"
}
}Several pages
{
"format": "html",
"items": "ul.notices > li",
"fields": {
"title": "h3",
"posted": "time@datetime",
"link": "a@href"
},
"pagination": {
"kind": "param",
"param": "page",
"start": 1,
"step": 1,
"maxPages": 3
}
}param pagination sets a query parameter (or fills a {page} placeholder in the URL) for each page; next follows a next-page link matched by a selector, on the same site.
Pages built in the browser
Some pages arrive empty and fill themselves with scripts. Browser mode renders them in an isolated, network-restricted Chromium before extracting (Pro and above). Try HTTP first: most pages, feeds and APIs do not need it.
{
"mode": "browser",
"format": "html",
"items": "[data-row]",
"fields": {
"name": "[data-name]",
"price": "[data-price]"
},
"waitFor": "[data-row]"
}Record identity and changes
Choose a key field (the record's ID, such as a permit or filing number) when you create a source. With one, an edited record shows as a field-level change (status: Under review → Approved); without one, records are compared by their whole content, so an edit appears as one removed and one added record.
When a source changes shape
If a run validates but looks broken compared with recent healthy runs (far fewer records, a usually-filled field now mostly empty, or a required field missing), it is marked degraded. Its evidence is kept, but it does not replace your current records or count as a change until an editor accepts it as correct.