Database Architecture and Data Location Expectations

Database Architecture and Data Location Expectations

Standards and Governance Basis

This guidance is intended to support the role of the Data Management and Coordinating Center (DMCC) in promoting consistent, fit-for-purpose, and standards-informed data management practices across studies in the Rare Disease Clinical Research Network (RDCRN). Database design decision-making is expected to be proportionate to the study type, participant risk, data criticality (a reflection of how vital data are to the mission of the RDCRN), and downstream use of the data.

This approach is informed by several overlapping expectations (Table 1).

Table 1. Rationale for database decision-making approach

Source / framework

Relevance to database design

NIH data management and sharing expectations

Supports planning for scientific data management, sharing, reuse, validation, and replication. This reinforces the need for research databases to be understandable, reusable, and appropriately protected.

ICH E6(R3) / Good Clinical Practice

Supports quality by design, proportionality, fit-for-purpose systems and processes, participant protection, and reliability of results. This reinforces the need to apply controls based on data importance and study risk.

GCDMP / clinical data management best practices

Supports clear planning, data collection design, eCRF guidance, data management plans, EDC implementation, study conduct, data quality, maintenance, and closeout.

HIPAA / privacy and minimum necessary principles

Supports appropriate access to PHI/PII and reduction of unnecessary use, disclosure, export, or sharing of identifiable information.

Network harmonization and trial-readiness expectations

Supports consistent data structures across studies and consortia to enable reporting, pooling, reuse, and future clinical trial readiness where applicable.

 Why This Matters

RDCRN database architecture is not only a REDCap build decision, as it affects participant privacy, access control, export risk, data quality, change management, harmonization, reporting, pooling, reuse, and future trial-readiness.

Separating operational, eConsent, and research data helps ensure that each type of data is collected, accessed, maintained, and shared according to its purpose and sensitivity (Table 2).

Table 2. Data Type / Location / Rationale Matrix

Data type

Location

General rationale

Network-required core variables*

Research database

Supports network-level reporting, pooling, reuse, and trial-readiness

Demographics*

Research database

Supports harmonization and cross-study comparability

Eligibility / enrollment*

Research database

Defines screened/enrolled populations and supports participant accountability

Consent status/date*

Research database

Needed for participant accountability; source consent remains separate

Study Disposition*

Research database

Needed for participant accountability and status

Diagnosis / disease classification

Research database

Enables consortium- and network-level analysis

Protocol-defined outcomes

Research database

Analysis-bound data; standards should apply where practical

Validated instruments / surveys

Research database

Analysis-bound and often harmonization-relevant

Source eConsent documents

eConsent database

Contains identifiable consent documentation and should remain separate

Name, email, phone, address, DOB

Logistical/operational or eConsent database

Needed for communication or study conduct, not generally for analysis; increases PHI/PII access and export risk. May be included if needed for analysis (e.g. geocoding) or for participant survey completion

Recruitment/contact attempts

Logistical/operational database

Used to manage study conduct, not analyze study outcomes

Scheduling/reminder tracking

Logistical/operational database

Workflow data that can create unnecessary cleaning and change-control burden in the research database

Shipping/device logistics

Logistical/operational database, or limited research status field if needed

Include in research database only when needed for accountability, missingness, or analysis interpretation

Internal coordinator notes

Logistical/operational database or restricted location

Reduces free-text/privacy risk and avoids cluttering the research record

External data receipt status

Research database or operational database, depending on use

Location depends on whether the field is needed for accountability, missingness, or analysis completeness

Derived variables

Analysis dataset/specification

Should generally be specification-driven rather than redundantly collected as raw eCRF fields

*Required elements

Hybrid Database Designs

Hybrid database designs where limited operational data are collected in the research database are not automatically prohibited, but they should be intentional and documented.

A hybrid design may be acceptable when the study is small or simple, PHI/PII is limited and carefully controlled, the same user group legitimately needs access to both operational and research data, and operational fields are clearly identified as non-analysis data.

A hybrid design should not be used if it creates unnecessary access, export, privacy, data cleaning, change management, or quality-control risks.

References / Supporting Standards