Documentation

Get data in and out

Import, connections and exports.

Audience: consultant. Prerequisite: 02 — Build the app.

A delivered app is only useful with the client’s real data in it. This page covers getting data in, keeping it clean, and getting information back out.


Connections — author the source once

Build ▸ Connections (/engagements/:id/build/connections).

A connection is a saved, credentialed data source. Author it once and reference it from anywhere, instead of embedding a URL and a password in each pipeline.

ProtocolFor
RESTa JSON HTTP API — its base address
GraphQLan endpoint plus a query
SQLPostgres / MySQL / SQL Server
SFTPa remote CSV or JSON feed
Webhookan outbound POST/PUT

Credentials are stored in a vault, not in the configuration, and are resolved at run time. A connection’s secret never comes back out of the API — when you edit a connection, the secret field shows as already-saved rather than being re-displayed.

Test the connection before building a pipeline on it. The wizard has a test step; a connection that has never been tested is a guess.

A REST connection is the system’s base address — the part every resource sits under, such as https://api.client.example/v2. A base usually has nothing at its root and answers 404, so the test step takes a path to test with (/titles) and where the rows are (data.items). They are kept with the connection, so every Test — in the wizard and on the Connections list — asks the same question, and a scheduled import or a page that picks this connection starts from them rather than asking again (change either there and the connection is left alone).

Where the rows are is chosen, not typed. Press Test and the answer is read for lists: each one is offered with how many rows it holds and what is in them (data.items — 6 rows · ref, pet, when). Type a path yourself only for a list the sample could not show — an endpoint that happens to be empty today.

Scheduled imports

Build ▸ Scheduled Imports (/engagements/:id/build/etl/pipelines). A scheduled import brings records in from a connected system, again and again. (Underneath it is an ETL pipeline; you will meet the word only under More steps.)

New scheduled import is four steps, in the order the job has — Import a Spreadsheet’s shape, done repeatedly:

  1. What and from where — which records it brings in, and which connection they come from. A scheduled import reads a saved connection, so the system’s address and credential live in one place. It does not take a file: a one-off spreadsheet is Import a Spreadsheet’s job. (The wizard opens on the kind of source you are adding, so a scheduled import of a file re-reads the same stored file on every run.) A name is optional: left empty it is VetBook — appointments → Appointment. For a REST connection, where to read is filled in from the connection and folded away; open it to change:
    • Path — the resource under the connection’s base (/titles). Filled in from the connection’s own test settings when it has them.
    • Where the rows are — the key the list sits under, dots between levels (data.items). Empty when the response is the list. Also filled in from the connection.
    • How it arrives — all in one response, or in pages (limit and offset, a next-page cursor, or the Link header), with the rows per page. A run reads up to 50 pages; if the source still has rows after that, the run loads nothing and says so — a part of the rows would look like all of them. Raise the page size.
  2. Match the columns — the connection’s own columns, read from one sample row, each beside the field it fills: their column → your field. A column is matched for you only when its name is a field’s name or key — the same matcher Import a Spreadsheet uses; anything else says Don’t bring in until you choose. Everything else a pipeline can do (the steps below) is under More steps.
  3. Check it — runs on arrival, against real sample data, without loading or saving anything: how many rows are ready, how many need a look, how many cannot be brought in, and a sample of rows by field name. Looking never saves the import; only Save does.
  4. How often — how to recognise a row it has seen before (below), then the schedule, then Save or Save & run now.

The transform steps

StepDoes
Column mapsource column → target field
Type coercetext → number / date / boolean
Value transformrewrite a field’s value
Derived fieldcompute a new field
Filter rowskeep or drop rows matching a condition
Fuzzy matchmatch messy values against existing records
Dedupecollapse duplicates, exact or fuzzy
Validateper-field checks with a severity
Enrichlook up a value on another entity and copy fields across
Entity coercecoerce every mapped cell to the target entity’s declared type
Resolve referencematch a reference by a business key (“the Account whose Name is this”)

Entity coerce and Resolve reference are what let a client import their own spreadsheet rather than an export of yours.

  • Entity coerce reads the target entity’s schema and turns each cell into the declared type — "$129,000" into a number, "North" into the north enum value, TRUE into a boolean. Without it a row lands holding raw strings, looks imported, and is refused the next time anyone edits it because it violates its own schema. It uses the same coercion the builder’s own Import a Spreadsheet panel uses, so the two paths cannot drift.
  • Resolve reference turns Account: Talbot Holdings into that account’s record id. Matching is exact after trimming and case-folding — never fuzzy, and an ambiguous match refuses the row rather than picking one. If two accounts share a name, you get told; use Fuzzy match when you deliberately want near-misses routed to review.

Dates are asked about, not assumed: 03/04/2026 is March 4th to an American and 3 April to everyone else, so the importer offers a day-first / month-first choice rather than guessing.

Every step is configured with typed controls — field references are pickers over the target entity’s fields and the source columns detected from your preview. Raw JSON survives only behind an explicit “Advanced” disclosure for long-tail options.

Recognise a row it has seen before — do not skip this

Choose the fields that identify a record (a booking reference, an invoice number). A row that arrives matching a record already in the app on those fields updates it instead of creating a second one. (Engineers call this the de-duplication, or upsert, key.)

Scheduling needs it. Without it every run would add the whole dataset again, so the schedule stays at Manual only until you choose. That refusal is a feature.

Review queue

Rows the pipeline is not confident about land in a review queue with a confidence score and the specific issues found. You accept, correct, or reject them. Nothing ambiguous is silently loaded.

Editing a scheduled import

An existing import is fully editable: open it and use Edit, which reopens the same builder with it loaded, on Match the columns. Changing a mapping, a connector or the target entity never requires deleting and recreating it.

The import’s Steps tab shows the same typed controls that authored the step, in a read-only form — so what a step does reads the same way whether you are inspecting it or editing it.

Scheduled imports run themselves

A scheduled import runs autonomously in the deployed app. No one clicks Run.

What the reader is told about where the data came from

Every page in a deployed app is live: a write anywhere — a form, a webhook, a workflow, an import — reaches the screens that show it within a few seconds, and the page says Live · updated just now under its title. Data that arrives in batches says so as well:

  • An import through the deployed app’s Import button is remembered with its file name, who loaded it and when. (The studio’s Import a Spreadsheet does the same for the studio’s own preview data — records do not travel from the studio to a client’s app.) An exhibit built from that data reads As of the upload on 12 Sep 09:14 by J. Smith (hauliers.csv). The import audit keeps the file name too.
  • A pipeline run is remembered with the pipeline’s name and its schedule in words. The exhibit reads Synced 14:00 · hourly.

⚙️ The line is there only while it is true. The moment anything live writes the same entity — a technician’s form, a webhook — the figure stops being “as of” the file, and the line disappears. So an entity seeded once by import and then maintained in the app reads as live, which it is.

⚠️ “Live” describes the page, not your source. It means the page will show a change the moment one is recorded. A webhook-fed entity whose sender has stopped sending still reads Live — nothing arrived, so nothing changed. That is the gap What this does not tell you (below) is about, and the anomaly watcher on freshness is still the way to close it.

Inbound webhooks

Build ▸ Inbound Webhooks (/engagements/:id/build/webhooks). Author an endpoint the client’s systems can POST to, mapped onto an entity. Use this when the source system pushes rather than waits to be polled — including when the work happens somewhere else entirely and only the result belongs here.

An endpoint carries its own mapping: which part of the payload holds the rows, which incoming field lands in which entity field, which fields identify a row for upsert, and — for any field that points at another record — which business key to match on. Posting the same row twice with a changed value updates it in place rather than creating a second one, so a job that re-sends its whole output is safe to re-run.

Receive first, map second. You do not type the sender’s keys up front. Make the endpoint for a kind of record — its path, which records, and how it is signed — and give the sender the address and secret. Its first deliveries are kept, and nothing is written; the sender is told so in words (“Not matched to records yet — the delivery is kept and nothing was written.”) and nobody is alerted, because nothing failed. The endpoint then shows Match what arrived: every key the sender actually sent, how many deliveries carried it and one value it carried, already matched where a key is named like a field (the same matcher Import a Spreadsheet uses). Choose the rest, say what finds a linked record and what recognises a row it has seen before, and save; the next delivery is written with it. A delivery kept before the matching stays kept — ask the sender to send again.

⚠️ The studio lists what the studio received. Match from deliveries sent to the address the endpoint shows in the studio. A sender posting to the client’s deployed app is kept there too, and nothing is written, but the studio cannot read that app’s deliveries yet, so the endpoint there would still read Waiting for the first delivery.

If the sender later spells a key differently, the key lands empty and the deliveries’ check says so — “pet_name arrived in 1 of 2 deliveries and nothing reads it”. Change the matching on the endpoint lists the new spelling beside the old one; switch it and save. Before this the only fix was to delete the endpoint and give the sender a new address.

⚙️ A row that changes nothing writes nothing. A matched row identical to the stored record is left alone: no update, no lineage event, no workflow run. The import preview counts these separately, as will leave unchanged, and the commit reports unchanged. Before this, an hourly feed re-wrote every row every hour: 450 readings and 3,600 lineage events per identical run on the Ashgrove app.

⚙️ A feed may send only what it knows. A row that matches a stored record is judged on the stored record with the row’s fields laid over it, so a back-office feed sending a reference and one changed date updates that date. Before this, every such row was refused for the required fields it did not send. A partial row whose key matches nothing is still refused: a new record needs every required field.

Rows are validated the same way any other write is. A payload missing a required reference is refused, and the response says how many rows were rejected rather than accepting the request and quietly storing nothing:

{ "received": true, "recordsLoaded": 0,
  "loadError": "3 row(s) rejected, none loaded — no Unit matches externalRef \"U-9999\"" }

The message names the cause, not just the count. A count on its own tells you something is wrong and nothing about what.

⚠️ A pushing system knows its own keys, not yours

The system posting to you knows its asset tag — AST-1187 — and cannot know the id your app generated for that asset. There is no point in a push integration where a person could look one up.

So a reference field is matched by a business key, exactly as an import does: say “the Asset whose Reference is this”, and send the tag. A key that matches nothing is refused — never stored as an unresolved link, because a record that looks connected and points at nothing is worse than one that was honestly rejected.

Every delivery is recorded — the payload, the source IP, whether the signature verified, whether a linked rule fired, and how many records landed. When something did not arrive, that list is where you look first.

⚙️ Where to look: Build ▸ Inbound Webhooks, under the endpoint — “Deliveries”. Each one answers three questions: what arrived, what it wrote, and what it dropped. A delivery that refused rows says so on its own line — “1 row(s) rejected, none loaded — Pet=“Biscuit” references a nonexistent pet” — and opening it names each refused row and shows the payload as sent.

⚙️ The log has a screen. Every delivery is readable where it lands rather than only in the page that rendered deliveries is firm-level and cannot see an app’s endpoints, so from the tab where you configure an endpoint a misconfigured sender looked healthy for ever. It was also recording a COUNT and not an outcome — a delivery whose rows were all refused read as “0 records loaded” with no reason, although every refused row had carried its reason all along. That is why the earlier finding about a dropped booking had to be reproduced by posting a controlled payload at a running app: there was nothing to look at.

⚙️ The record is bounded: each endpoint keeps its newest 1,000 deliveries, and none older than 30 days. A daily sweep removes the rest, so a busy feed does not accumulate its whole history — a feed posting every ten seconds adds about 8,600 a day — and the database, every backup of it and every restore check grew at the rate of the busiest sender. The records a delivery LOADED are unaffected: this bounds the log of arrivals, not your data.

⚙️ In the deployed app, Integrations ▸ What sends to this app lists each endpoint with its latest delivery — when it arrived, how many records it loaded, and whether it was refused or failed to load, for every seat. An administrator’s Sender details on a feed hold the address the sender posts to and, on request, the secret it signs with (every reveal audited). A deployed app generates its own secret on first start, so what a vendor needs comes from the app, not the studio. The full delivery list, payloads included, is still the studio’s.

Set upWhat it is for
SignatureHMAC over the raw body, so a forged POST is rejected. Leave it off only for a source that cannot sign.
MappingPayload field → entity field, plus the key that makes a repeat an update.
Linked ruleFire one specific rule on arrival — notify someone, start a workflow, escalate.
Linked workflowStart one published workflow on each verified delivery. It reads what was sent as {{payload.<field>}}. The sender’s answer says workflowTriggered: true with the run’s id only when a run started; otherwise it’s false with the reason (for example, the workflow is still a draft).

Your app reacts to what arrives, however it arrives

There are two ways a rule runs when data comes in, and it is worth knowing which is which:

  • The linked rule above is attached to the endpoint. It fires on delivery — before anything is mapped — and is the right place for “tell me whenever this system calls us”.
  • Your entity rules — the ones under Build ▸ Behavior, “when a ticket reaches tipped, notify the supervisor” — now fire on the records an ingest creates or updates, exactly as they do when a person edits the record by hand. That covers an inbound webhook, a scheduled import and a spreadsheet upload alike.

⚙️ Every ingest path fires both rules and workflows — an inbound webhook, a scheduled import, a spreadsheet commit, the JSON importer and a reviewed ETL load all behave as a typed-in record does. A weighbridge posting a ticket alerts the supervisor exactly as the same change typed by a person would.

⚙️ A trigger scoped to fields fires only when one of those fields changed, so a scheduled re-pull that changes nothing starts no runs.

⚙️ A bulk import does not fan out without limit. Rules fire for up to 200 records per ingest; beyond that the records still all land, the rule fan-out stops, and the response and the log both say how many were skipped. A ten-thousand-row migration is history being loaded, not ten thousand things that just happened — and one rule with an outbound webhook would otherwise make ten thousand calls at somebody else’s system.

⚠️ The endpoint is public by design — that is the point of a webhook — so the signature is what stands between it and anyone who learns the URL. Set one whenever the sending system can.

Modelling elsewhere, landing it here

A practice that already models in Excel or Python does not have to stop. The webhook above is how the output gets here, and it is a legitimate architecture rather than a workaround — but only if you know which half does what.

The modelling stays where the analyst already works. A scenario, a forecast, a scored list, a monthly re-run of something with real statistics in it — that belongs in the tool built for it. What lands here is the output: the rows a reader needs to see.

What the platform adds is the part a spreadsheet is bad at. Durable records rather than a file with a date in its name. A governed audience rather than whoever was on the email. An audit trail. A screen the client can open themselves, that looks the same next month.

The pipe is the one above. Point the job at an inbound endpoint, map its fields onto an entity, and set the upsert key so a job that re-sends its whole output updates rows in place instead of doubling them. Sign the request. Read the delivery list when something is missing. Nothing in this section is new — that is the point.

Landed rows are ordinary records

⚠️ This is the part practices learn from a client. A row that arrived by webhook is not a special kind of data with its own visibility rules. It is a record, and every rule about who may see a record applies to it unchanged — including the two that surprise people: an entity grants nothing until you say so, and walking your own app as an admin proves nothing about what a client’s staff will see, because an admin never touches the permission matrix at all. The decision is drawn in full in 02 — Build the app; do not take a green screen of your own as evidence.

⚙️ A model’s output usually wants read for the client’s staff and write for nobody. The writer is the webhook, not a person. The instinct is to grant the role that “owns” the numbers permission to change them; that role almost never needs it, and granting it means a hand edit can silently disagree with the model that produced the row.

What this does not tell you

⚠️ A delivery record proves what arrived. Nothing here proves what did not. If the job dies on a Tuesday, no request is made, no delivery is recorded, and the app goes on showing Monday’s numbers — correctly, and looking entirely healthy. Silence and success are the same shape from this side of the pipe.

The mitigation is an anomaly watcher on freshness: alert when the newest record for the entity is older than the cadence the job is supposed to run at. It is the difference between finding out yourself and being told by the client.

⚙️ Worked through end to end, with the real numbers and the parts that went wrong, in Landing a model’s output — re-pointing references between instances, an upsert key that turned out not to be one, and the seat check that found a defect the row count could not.

Documents are uploaded, stored, and OCR’d where applicable, then searchable alongside records. Attach them to records so the context travels with the data.

File storage is pluggable: local on disk by default, S3-compatible object storage (AWS S3, Cloudflare R2, DigitalOcean Spaces, MinIO) by configuration only — no code change to move a client to their own bucket.

Reports out

End-user reports are templates the client’s staff can run themselves in the delivered app: pick an entity, choose fields, group, aggregate, and scope to a role. Author the templates in Build; the client runs them in the app.

Records can also be exported from the app’s data tables directly.


What “done” looks like for data

  • Every connection has been tested, not just saved.
  • Every pipeline has a de-duplication key (and therefore can be scheduled).
  • You ran a preview and looked at what landed in the review bucket.
  • A re-run of the same import updates rather than duplicates — verify it once.
  • The client can get their own answers out without asking you.