Skip to content

Data Catalog and Discovery

Unified data cataloging, discovery, and governance powered by OpenMetadata and AWS data services.

Overview

RosettaHub integrates with OpenMetadata and AWS data services to provide a complete data management layer alongside compute orchestration. Researchers can discover datasets, understand data quality and lineage, and launch pre-configured environments to analyze data -- all from within the platform.

OpenMetadata Integration

OpenMetadata is an end-to-end open source metadata platform providing data catalog, discovery, data quality, governance, collaboration, and lineage capabilities.

Integration Architecture

RosettaHub provides deep integration with OpenMetadata:

Capability Description
SSO via Keycloak Researchers sign in to OpenMetadata with their institutional identity -- no separate credentials
Role propagation RosettaHub organization roles (admin, manager, user) map to OpenMetadata roles automatically
Datasets as cloud model artifacts Datasets cataloged in OpenMetadata appear as first-class artifacts in the RosettaHub cloud model
Access control Dataset access permissions are managed from within RosettaHub and enforced in OpenMetadata

What OpenMetadata Provides

  • Data catalog -- browse and search all datasets across the organization
  • Data discovery -- find relevant datasets by topic, tag, owner, or schema
  • Data quality -- define and monitor quality rules, track data freshness
  • Data lineage -- visualize how data flows from source to analysis
  • Collaboration -- comment on datasets, request access, share knowledge
  • Governance -- classify data sensitivity, enforce retention policies

Timeline

OpenMetadata integration is production-ready in early April 2026.

AWS Data Integrations

RosettaHub integrates with AWS data services to bring external datasets into the platform:

AWS Data Exchange

Deep integration with AWS Data Exchange enables:

Capability Description
Mirror public datasets All public datasets available through ADX are mirrored to a dedicated RosettaHub marketplace
Auto-map acquired datasets Datasets acquired via ADX are automatically mapped to RosettaHub artifacts
One-click environments Launch Formations pre-configured to visualize and analyze specific datasets

Additional AWS Data Services

Service Integration
Amazon DataZone Metadata domains and data products integrated with RosettaHub catalog
Amazon SageMaker Catalog ML model and dataset catalog accessible from RosettaHub
AWS Lake Formation Fine-grained data access controls and data lineage
AWS Registry of Open Data Public research datasets (Sentinel-2, Landsat, ERA5, NOAA) available as pre-mounted storages

Data + Compute Bundles

One of RosettaHub's unique capabilities is bundling data access with compute environments:

  1. Catalog a dataset in OpenMetadata (or acquire via AWS Data Exchange)
  2. Create a Formation that pre-mounts the dataset as a Storage
  3. Publish to the marketplace -- researchers browse and launch with one click
  4. Access controls apply -- only researchers with the right project membership and data tier can access the formation

This eliminates the gap between "finding data" and "analyzing data" that exists in most research environments.

Private Marketplace

Institutions can build private marketplaces of data-configured environments:

  • Institutional data catalogs -- curated datasets relevant to the organization
  • Pre-mounted storages -- S3 buckets, EFS file systems, or cross-cloud storage attached to formations
  • URL-based sharing -- share data-configured environments via link
  • Folder-level S3 mapping -- different teams see scoped views of the same underlying bucket