A practical guide for institutions charting a path from storage-first to data-first thinking.
Starting with the Wrong Question
Most research institutions begin data management platform evaluations with the same question: “How much storage is required?” It sounds reasonable, but it’s feeding the wrong end of the horse. This storage-first mindset has pushed many universities and research organisations into expensive and unsustainable positions.
The right question centres on manageability, not capacity.
Will you be able to locate your own data five years from now? Can a researcher in another department discover a dataset relevant to their work? Will your institution meet funding mandates, HIPAA and FERPA requirements, or FAIR principles without heroic manual effort?
Whether you’re evaluating vendors, building a business case, or trying to navigate an increasingly crowded technology landscape, the data management criteria truly matter.
The Core Problem: Silos Cost More Than You Think
Most research data is siloed and fragmented, making it difficult to gain visibility into a critical asset.
A common scenario looks like this: a botany lab maintains its own datasets, a social science department uses an external cloud service, the library maintains its own archive, and every other research group, school, or faculty has organically grown its own idiosyncratic data management ‘system’. IT is tasked with overseeing the institutional infrastructure, but no one has a complete picture of what data exists, how much is duplicated, where it resides, or whether it is still actively used.
This is not simply an IT problem. It’s an institutional juggernaut that no single stakeholder can unpick. And the consequences are significant:
Researchers lose significant time or duplicate work because existing datasets cannot be readily located or discovered.
Inactive data accumulates on expensive tiers exponentially without automated lifecycle management.
Manual and time-consuming compliance audits create major risk exposure.
Cross-department collaboration is hindered by inconsistent access controls.
Storage costs spiral without a sustainable long-term strategy.
Buying more storage is not a real solution to these issues, that just kicks the can down the road tacking on unnecessary additional TCO. To overcome this hurdle, institutions need a platform able to treat research data as a shared institutional asset rather than a commodity, like a chair or a paper shredder, isolated within individual projects or departments and abandoned like a stack of outdated textbooks when the researcher moves on.
Getting It Right: Princeton University's TigerData
Princeton University faced many of the same challenges confronting large research institutions today: Exponential data growth and fragmented, incompatible, and inconsistently governed storage systems without a sustainable long-term strategy to guide them out of their predicament.
Rather than asking, "How do we solve our storage problem?" they turned the problem on its head and instead considered:
"How will we manage our research data for the next 100 years?"
That question fundamentally changed the conversation.
Planning a century-scale data management strategy raised a critical challenge: how often might underlying storage technologies be supplanted, and how could data remain accessible and governed through repeated technology refresh cycles?
The answer emerged as TigerData, a metadata-driven research data management platform powered by Arcitecta's Mediaflux.
Rather than continuously increasing storage capacity, Princeton implemented a lifecycle-aware architecture capable of operating across multiple storage technologies, including:
Dell PowerScale
Dell ECS
IBM GPFS
IBM Diamondback tape libraries
By introducing Mediaflux as a persistent data management layer above the underlying storage infrastructure, Princeton gained unified access, metadata-driven automation, lifecycle governance, and the flexibility to evolve storage technologies without disrupting researchers.
Today, TigerData:
Manages 200 petabytes of research data
Tracks nearly 497 million digital assets
Unifies more than 70 PB of heterogeneous storage under one platform
Supports research disciplines ranging from genomics to linguistics
TigerData upgraded Princeton from reactive storage provisioning to proactive data stewardship.
Their lesson is more than a matter of scale, it is a story of institutional transformation, shifting from a “storage-first” model to a “data management-first” approach.
The Secret Ingredient: Metadata Management
When evaluating a data management platform, the focus tends to be on storage capacity, cost, and cloud integration. Metadata management is frequently overlooked, making it a significant reason many implementations fail to deliver long-term value.
Metadata is the connective layer enabling large-scale data management capabilities.
It answers essential questions:
What is this dataset?
Who created it?
When was it last accessed?
Is it still relevant?
Does it meet preservation and compliance requirements?
Without a robust metadata-driven architecture, a large-scale storage environment is little more than a resource-hungry burial ground. Data may be retained, but it is more dead weight than viable asset.
When evaluating platforms, institutions should ask:
Can the platform support custom metadata schemas by discipline or project type?
Is metadata captured automatically during ingest?
Is metadata searchable and queryable across the full environment?
Does the platform support FAIR principles, making data findable, accessible, interoperable, and reusable?
With Mediaflux, institutions can evolve from simply storing data to organising, governing, and operationalising it, flipping the risk-benefit scales in their favour.
At Princeton, metadata-driven organisation creates a dynamic data ecosystem, wherein millions of digital assets remain discoverable, retrievable, and reusable in perpetuity, rather than simply sucking up dollars and sense.
The Seven Things That Really Matter
Based on real-world institutional deployments, these criteria separate the platforms built for long-term research support from those that only overcome today’s storage problems.
1. Built for Decades of Growth
Research data does not shrink over time. A platform that works today may become strained five or ten years from now. Institutions should ask vendors for evidence of deployments operating at institutional scale, not just raw storage benchmarks, but proven management of hundreds of millions of assets over time.
2. Supports Heterogeneous Storage
Most institutions already have substantial storage investments. A data management platform should act as a unifying layer across vendors, rather than require a disruptive rip-and-replace migration to upgrade hardware. Platforms such as Mediaflux unify on-premises, cloud, and tape storage into a single operational environment, allowing institutions to retain infrastructure investments while modernising data management practices.
3. Manages the Full Research Data Lifecycle
Research data has a reasonably predictable lifecycle. It is generated, actively used for a relatively short intense period, published, archived, and potentially reused later.
An effective platform automates movement across storage tiers according to institutional policies and access patterns, reducing costs and administrative overhead while keeping data accessible when needed.
4. Checks the Compliance and Security Tick-Boxes
Research institutions face increasingly complex compliance obligations:
FERPA for student records
IRB requirements for human subject research
NIH and NSF data sharing mandates
HIPAA for healthcare-related research
Compliance cannot be an afterthought. Platforms should provide governance, auditing, policy enforcement, and security controls as foundational capabilities. As these obligations apply across all versions and backup copies, a sensible platform links assets to mandates via metadata, facilitating automated governance adjustments aligned to retrospective policy updates.
5. Supports FAIR Data Principles
Funding agencies and publishers increasingly require research data to be findable, accessible, interoperable, and reuseable.
A strong metadata architecture enables FAIR principles operationally, rather than relying on manual efforts down the track.
6. Unifies Access Regardless of Storage Location
Researchers should not need to keep track of whether their data resides on disk, tape, object storage, or cloud.
An effective data management platform provides a unified access layer through web interfaces, shared network folders, APIs, and automation workflows, making data location an abstracted infrastructure concern, rather than a researcher headache.
A persistent global namespace decouples logical data access from physical storage infrastructure, allowing institutions to evolve storage technologies over time without disrupting users or workflows.
7. Makes Costs Transparent and Automates Storage Tiering
Unchecked data growth is more than a technical problem; it is a significant financial management issue.
Platforms that provide visibility into storage consumption and automate movement of inactive data to lower-cost tiers allow institutions to align storage spending with actual usage, guided by policy.
The Mistake That Derails Most Evaluations
The most common mistake organisations make when selecting a data management platform is prioritising storage capacity instead of data management.
Storage capacity is easy to compare. It fits neatly into spreadsheets and procurement exercises.
Manageability is harder to quantify but ultimately determines whether a platform remains effective as data volumes grow by ≥10x.
When organisations focus exclusively on storage metrics, they optimise for:
lowest cost per terabyte
maximum throughput
fastest deployment
These are important considerations, but they are no longer differentiators.
More importantly, once your data is in the platform, what can you actually do with it?
Is it discoverable? Can it be governed consistently? Can it move intelligently across storage tiers? Can the institution evolve infrastructure without disrupting users?
That is what underpins long-term sustainability.
Your Data Management Platform Evaluation Checklist
Use this checklist when evaluating platforms or reviewing your current data management infrastructure.
The “Ask Vendors” column highlights key questions that distinguish proven platform capabilities from marketing claims.
| Criteria | Ask Vendors |
|---|---|
| Metadata-driven architecture | How does the platform manage metadata at institutional scale? |
| FAIR data principles | Is data findable, accessible, interoperable, and reusable? |
| Tiered storage & lifecycle management | Can data move flexibly and automatically across storage tiers? |
| Heterogeneous storage integration | Does it work with existing storage vendors? |
| Scalability (decades of growth) | What is the proven upper limit of data/assets managed? |
| Unified access layer | Can users access all data through one interface? |
| Compliance & security | Does it support HIPAA, FERPA, IRB, or funding mandates? |
| Cost transparency & storage tiering | Are storage costs visible and manageable by tier? |
| Governance & accountability | Who owns each dataset and how are permissions managed? |
| Researcher/end-user usability | Has the platform been tested with real research users? |
It’s critical to recognise from the outset that successful platform implementation necessarily involves researchers and end users in evaluations, not just infrastructure teams. The platform needs to make data management feel natural to scientists and scholars in research environments, not just tick administrative boxes, to be truly effective.
Think in Generations, Not Refresh Cycles
Princeton’s 100-year strategy may sound ambitious, but it reflects a reality every research institution will eventually confront: The data created today must remain accessible, understandable, and reusable by researchers who may not yet be born. Technology will change many times over. Storage vendors will come and go. The data is what persists.
The right data management platform is not the one with the largest storage pool or the lowest cost per terabyte. It is the platform that treats data as a long-term institutional asset, supported by metadata architecture, lifecycle governance, integration flexibility, and operational scalability capable of supporting research for decades to come.
The checklist above is a practical place to begin. The next step is to shift the conversation from storage to stewardship. Rather than asking, "Where will we store our data?", consider: "How will we manage, govern, and sustain research activities for the long term?"
