The provenance graph
How every result's origin is recorded, badged, and made tamper-evident.
Every result in Dalea can answer two questions: where did this come from? and what was made from this? The provenance graph is the record behind those answers — the chain raw file → import → query → code → report, captured by the server at the moment each change is enforced, never reconstructed from what a person or an AI says they did.
Runs and nodes
The graph has two kinds of things in it:
- A run is one recorded activity — an Import, a Data pull, a Code execution, a Document edit, an Export, an External run, a Data release, a Package export.
- A node is an immutable version of something — file version 3, document snapshot 12, one version of a data object, pinned by its version key. Where the bytes were actually observed the node also carries a content hash, which is what keeps it matchable later. The graph never points at "the file"; it points at the exact version a run read or wrote.
On the fail-closed surfaces (imports, code execution, document edits, staging applies, exports, external-run deposits, file uploads, inventory operations) the run is written on the same transaction as the change itself: mutation and evidence commit or roll back together, so there is no evidence-free write. A smaller set of surfaces (record create and update, table create, saved-query edit, file restore, addon build, marketplace install, Benchling sync) emits its run immediately after the substrate write has committed. There the edge is repairable from the substrate rather than guaranteed atomically, and every dropped emit is counted as a service metric instead of passing unnoticed.
Read and write edges only ever connect runs to nodes: a run read these versions and wrote those. Beside them the graph carries node-to-node derivation edges (wasDerivedFrom), each attributed to the run that made the claim, because reading X and writing Y does not on its own prove Y came from X. Walking edges backward from an artefact is its origin; walking forward from a source is its impact: the retraction question, "what do we have to re-examine if this input was wrong?"
The cell sealed with its code, its pulls, and its outputs. The dashed edge is coarse: reference_ranges.csv was open in the session, so it is recorded as a candidate source — over-approximated on purpose.
Who did it
Every run records its actor twice over: the kind of actor — human, AI agent, or service — and the accountable person. An AI-performed run always renders as "AI agent — on behalf of you": the machine executed it, a named human answers for it. The two are never collapsed, which is what makes the review gate on exports possible (more below).
How honest is each link
Dalea refuses to pretend all evidence is equal. Every edge carries a quality grade; derivation edges carry a precision grade as well.
| Axis | Value | Meaning |
|---|---|---|
| Quality | verified | A server observed this directly. |
| declared | Recorded from a producer's own statement, not independently observed. | |
| inferred | Derived by the platform from surrounding evidence. | |
| Precision | exact | This exact input was observed feeding this output. |
| coarse | Over-approximated: a candidate source, not a confirmed one. Rendered dashed and badged. |
A plain read or write edge is graded on quality only. Precision belongs to the derivation claim, which is where over-approximation actually happens.
Coarse edges deliberately over-approximate — if a file was merely open in a code session, it is recorded as a possible source. That is the safe direction for a recall: an impact walk may flag an artefact that was not really affected, but it will not miss one that was.
The same honesty applies to AI work. What an assistant claims its lineage was is stored separately from what the server observed — the two are compared, not merged, so a hallucinated citation can never masquerade as evidence.
What survives deletion
Provenance endures deletion. When a record is removed, its trail remains and the node stays matchable by content hash — the badge simply reads deleted. An absent trail, on the other hand, is not evidence that nothing happened: capture begins when the first run touches an entity.
Can the record itself be trusted
A lineage record is only as good as its resistance to tampering, so the graph carries its own evidence in two layers:
- Structural checks re-derive every hash, chain link, and plan reference from the raw records — do the records agree with themselves?
- Tamper evidence: every workspace's history is committed into an append-only transparency log (the same construction Certificate Transparency uses) with signed checkpoints, and those checkpoints are anchored outside Dalea's own control.
Structural does not equal truth — a consistent record could still have been
consistently rewritten before anchoring. That is why every verification verdict
carries an honest anchored flag: when it is false, no external
attestation covers the segment yet, and the verdict says so rather than
overclaiming.
A plate reader exports plate_reader.csv for study DLA-7. Dr Okafor
imports it — 247 rows land, each carrying an origin pointer to that exact file
version. Later she asks the assistant for a PK summary plot; the sandbox pulls
the rows (query plan and result hash recorded), the cell seals, and
pk-summary.png appears. Opening the plot's provenance shows the
full chain back to the CSV — every hop verified, the one speculative input
marked coarse — and an impact walk from the CSV would find the plot if the
instrument's calibration were ever questioned.