Data Provenance
Transparency Trail's Data Provenance layer answers the questions that come after the analysis: How many observations were actually analyzed? Were these results produced with the same data as the other figures? Which version of the software generated this output? All of it is captured automatically and travels with every export.
What is Data Provenance?
Data provenance — the documented history of where data came from and how it was processed — is a prerequisite for computational reproducibility. In practice, most publications provide only a partial picture: the sample size is reported, but not the number excluded; the software is named, but not the version; the analysis is described, but its relationship to other analyses on the same dataset is not declared.
Licklider's Data Provenance system closes these gaps across three dimensions: observation-level tracking (N Disclosure), dataset-level tracking (Analysis Family Ledger), and software-level tracking (Software & Version Provenance).