Scientific Research

Instruments, samples and results, linked.

Research produces heterogeneous data at every size, from kilobytes of annotations to terabytes of instrument output, and the metadata is what makes it reusable.

A dataset found by a query against its metadata, not by remembering a directory.

The story

A dataset found by a query against its metadata, not by remembering a directory.

Where it starts

Instrument and simulation output

Datasets, runs and derived products at large sizes. It is the first of 4 workloads running in Scientific Research.

The question it raises

Dataset discovery

Large objects are found by querying run metadata rather than by directory convention.

Why one question is hard

From Hypothesis to Publish

Scientific Research data moves through 4 stages — Hypothesis → Experiment → Compute → Publish. The shapes in play are Objects, JSON / documents, Time series, SQL, Vector, and answering one question means reading across all of them.

What PLOMID contributes

Research data is defined by its metadata. These are the parts that keep metadata first class.

The environment

Instrument output, simulation results, reference material and datasets.

Simulations and instruments write large objects with structured run metadata, while notes, papers and reference material are documents. PLOMID holds the objects beside the metadata that describes them, so a dataset is found by a query rather than by remembering a directory.

Data models in play

The shapes, in one layer.

5 shapes carry this domain. Choose a stage to read the operation, or a shape to see every stage that handles it.

PLOMID · Scientific Research hypothesis · experiment · compute · publish Select a stage
Stage

Hypothesis

Notes, references and related work

Vector · JSON / documents

Workload map Scientific Research workload map. Every shape on it is a surface of the layer, and each stage names the part of the operation it carries.
  • Objects Large assets with queryable metadata
  • JSON / documents Documents and nested objects
  • Time series Measurements and events in time order
  • SQL Records, keys and joins
  • Vector Similarity and semantic retrieval
The data journey

How scientific research data reaches one layer.

Walk the path the data takes, from the environment that produces it to the questions it answers. Select a station, or a shape, to read each step.

A dataset found by a query against its metadata, not by remembering a directory.

Environment

The instrument

Simulations and instruments writing large objects with structured run metadata.

Objects · JSON / documents

Where the data goes to work

Questions evidence has to answer.

Each one keeps the measurement with what qualifies it, read from the same layer rather than reconciled later.

Dataset discovery

Large objects are found by querying run metadata rather than by directory convention.

  • Vector
  • JSON / documents

Reproducible runs

Parameters, versions and calibration records stay attached to the output they produced.

  • Time series
  • JSON / documents

Evidence retrieval

Related work and internal notes are retrieved alongside the data they explain.

  • Objects
  • SQL

Shared infrastructure

One layer serves several groups without each maintaining a store.

  • JSON / documents
  • Objects
One environment · many workloads

What runs against scientific research data.

4 workload families over one set of shapes. Choose one to see what it moves and where it lands.

Datasets, runs and derived products at large sizes.

  • Planned once against the layer, not once per store
  • Read beside the records it shares a key with
  • Persisted under one storage contract
Workload architecture

From measurement to evidence.

The work, the shapes it names and the path a request takes — with the protocol beside the measurement.

Scientific Research · workload architecture
Workloads

What runs against this data.

  • Instrument and simulation output
  • Run and dataset metadata
  • Notes and papers
  • Measurements
Data models

The shapes those workloads read and write.

  • Objects
  • JSON / documents
  • Time series
  • SQL
  • Vector
The layer

One path from a request to the data it names.

  • Planning Predicates narrow the work before it runs
  • Execution Records, fields and windows answered together
  • Transactions Readers and writers do not block each other
Surfaces

How the work reaches the layer.

  • SQL surface The query language the layer is documented in
  • Applications Services and jobs writing and reading as they run
  • Analytics & AI clients The same layer, the same access path
Deployment & residency

Where this data is allowed to run.

Institutions run their own infrastructure, often on a shared cluster rather than a managed service.

Deployment, residency and control
What you build next

Data Infrastructure

One layer for records, documents and time-ordered data, instead of one system per shape.

If Dataset discovery is your question, start here.