Storage systems are designed to retain bits, but rarely to retain context across data lifetimes. Data sans context are just bits. Bits swiftly become waste, and a rather expensive liability.
Enterprise and healthcare data are domains for another day, each its own minefield to explore and defuse. Research data is the domain de jour…
Research data comes in more flavours than gelato. From a quiet filing cabinet filled with a life's collection of qualitative notes on fifth century script transformation to supercomputing centres stacked with 700 PB of radio astronomy data annually and everything in between, the volume, velocity, variety, veracity and value of research are all over the map. And yet, most research institutions are given singular, cookie-cutter mandates to "manage" it all.
The data management meal typically starts with an entree of storage provisioning –– How much stuff are we going to need to hang on to for how long? With a side helping of security and backup –– Where's the perimeter? Who's allowed in and how do we keep everything safe? Topped off with a sweet reward of governance and provenance allocations, if you're lucky and didn't fill up on bread before being served. But all too often, budgets get bloated and no one has time to linger on attaching meaning to all that carefully stored data. The data's progenitor is responsible for serving and downing the digestif on their own.
Keeping data in context was "manageable" in the pre-big data era, at least hypothetically. In practice, plenty of context was lost under stacks of papers, in drawers of poorly labelled or degraded samples, in lab fires/floods/earthquakes, departmental mergers/moves/restructures, erased discs and hard drives, graduations, human errors, firings and deaths. The law of entropy doesn't particularly care about our best laid plans and keeping data interpretable requires steadfast order. The difference today is the magnitude of the problem and the ramifications thereof.
To avoid becoming curators of chaos, systems must indelibly link bits with context. The context needs to go wherever, whenever, however the bits do. Without fail, without overhead, without a small army of data-herds to shuffle them around from field to field. Most systems are the blinded Cyclops erroneously relying on touch to count Odysseus’ men under sheep skins, presuming to know what all we have, even when all is lost.
Meaningful context is more than just when was the file created/modified by whom; more than when did it get migrated or backed up from that bucket to this one. When it comes to research data, it is instrument operating conditions and operator, place in sequence, governance policy and grant number. It is the answers to: Who all has touched this? When did they do it? and What did they do to it? Alas, we may only wonder Why?!
Research data management systems require a glue to bind context to bits. Everything any researcher or data manager wants or needs to know about each datum must remain attached, keeping data interpretable for life. Without interpretability, ‘data’ is merely debris. For the sake of progress, our future selves, and future generations, we need the data we keep to stay interpretable.
Mediaflux® from Arcitecta makes it possible to bake in context; to adhere the context to the bits. Fully extensible metadata schemata enable researchers, lab/instrument technicians, collaborators, data managers, reviewers, secondary users, etc. to attach their sticky notes of context to data. Teams using Mediaflux DAMS can further boost their collaborative efforts with in-collection chat functions, which tie research discussions directly to their object. With the XML database imbedded directly into the file system and every state change traced, evolution gets indelibly recorded without overwriting what went before.
Check out the new Mediaflux DAMS for Research Environments: https://www.arcitecta.com/mediaflux/dams-research/.
