Dataset discovery
Large objects are found by querying run metadata rather than by directory convention.
- Vector
- JSON / documents
Instruments, samples and results, linked.
Research produces heterogeneous data at every size, from kilobytes of annotations to terabytes of instrument output, and the metadata is what makes it reusable.
A dataset found by a query against its metadata, not by remembering a directory.
Datasets, runs and derived products at large sizes. It is the first of 4 workloads running in Scientific Research.
Large objects are found by querying run metadata rather than by directory convention.
Scientific Research data moves through 4 stages — Hypothesis → Experiment → Compute → Publish. The shapes in play are Objects, JSON / documents, Time series, SQL, Vector, and answering one question means reading across all of them.
Research data is defined by its metadata. These are the parts that keep metadata first class.
Simulations and instruments write large objects with structured run metadata, while notes, papers and reference material are documents. PLOMID holds the objects beside the metadata that describes them, so a dataset is found by a query rather than by remembering a directory.
5 shapes carry this domain. Choose a stage to read the operation, or a shape to see every stage that handles it.
Hypothesis
Notes, references and related work
Vector · JSON / documents
Walk the path the data takes, from the environment that produces it to the questions it answers. Select a station, or a shape, to read each step.
A dataset found by a query against its metadata, not by remembering a directory.
The instrument
Simulations and instruments writing large objects with structured run metadata.
Objects · JSON / documents
Each one keeps the measurement with what qualifies it, read from the same layer rather than reconciled later.
Large objects are found by querying run metadata rather than by directory convention.
Parameters, versions and calibration records stay attached to the output they produced.
Related work and internal notes are retrieved alongside the data they explain.
One layer serves several groups without each maintaining a store.
4 workload families over one set of shapes. Choose one to see what it moves and where it lands.
Datasets, runs and derived products at large sizes.
Parameters, versions, calibrations and provenance.
Laboratory notes, drafts, publications and reviews.
Time-ordered observations and monitoring series.
The work, the shapes it names and the path a request takes — with the protocol beside the measurement.
What runs against this data.
The shapes those workloads read and write.
One path from a request to the data it names.
How the work reaches the layer.
Institutions run their own infrastructure, often on a shared cluster rather than a managed service.
Deployment, residency and control