Database Architecture and Data Location Expectations
Standards and Governance Basis
This guidance is intended to support the role of the Data Management and Coordinating Center (DMCC) in promoting consistent, fit-for-purpose, and standards-informed data management practices across studies in the Rare Disease Clinical Research Network (RDCRN). Database design decision-making is expected to be proportionate to the study type, participant risk, data criticality (a reflection of how vital data are to the mission of the RDCRN), and downstream use of the data.
This approach is informed by several overlapping expectations (Table 1).
Table 1. Rationale for database decision-making approach
Source / framework | Relevance to database design |
NIH data management and sharing expectations | Supports planning for scientific data management, sharing, reuse, validation, and replication. This reinforces the need for research databases to be understandable, reusable, and appropriately protected. |
ICH E6(R3) / Good Clinical Practice | Supports quality by design, proportionality, fit-for-purpose systems and processes, participant protection, and reliability of results. This reinforces the need to apply controls based on data importance and study risk. |
GCDMP / clinical data management best practices | Supports clear planning, data collection design, eCRF guidance, data management plans, EDC implementation, study conduct, data quality, maintenance, and closeout. |
HIPAA / privacy and minimum necessary principles | Supports appropriate access to PHI/PII and reduction of unnecessary use, disclosure, export, or sharing of identifiable information. |
Network harmonization and trial-readiness expectations | Supports consistent data structures across studies and consortia to enable reporting, pooling, reuse, and future clinical trial readiness where applicable. |
Why This Matters
RDCRN database architecture is not only a REDCap build decision, as it affects participant privacy, access control, export risk, data quality, change management, harmonization, reporting, pooling, reuse, and future trial-readiness.
Separating operational, eConsent, and research data helps ensure that each type of data is collected, accessed, maintained, and shared according to its purpose and sensitivity (Table 2).
Table 2. Data Type / Location / Rationale Matrix
Data type | Location | General rationale |
Network-required core variables* | Research database | Supports network-level reporting, pooling, reuse, and trial-readiness |
Demographics* | Research database | Supports harmonization and cross-study comparability |
Eligibility / enrollment* | Research database | Defines screened/enrolled populations and supports participant accountability |
Consent status/date* | Research database | Needed for participant accountability; source consent remains separate |
Study Disposition* | Research database | Needed for participant accountability and status |
Diagnosis / disease classification | Research database | Enables consortium- and network-level analysis |
Protocol-defined outcomes | Research database | Analysis-bound data; standards should apply where practical |
Validated instruments / surveys | Research database | Analysis-bound and often harmonization-relevant |
Source eConsent documents | eConsent database | Contains identifiable consent documentation and should remain separate |
Name, email, phone, address, DOB | Logistical/operational or eConsent database | Needed for communication or study conduct, not generally for analysis; increases PHI/PII access and export risk. May be included if needed for analysis (e.g. geocoding) or for participant survey completion |
Recruitment/contact attempts | Logistical/operational database | Used to manage study conduct, not analyze study outcomes |
Scheduling/reminder tracking | Logistical/operational database | Workflow data that can create unnecessary cleaning and change-control burden in the research database |
Shipping/device logistics | Logistical/operational database, or limited research status field if needed | Include in research database only when needed for accountability, missingness, or analysis interpretation |
Internal coordinator notes | Logistical/operational database or restricted location | Reduces free-text/privacy risk and avoids cluttering the research record |
External data receipt status | Research database or operational database, depending on use | Location depends on whether the field is needed for accountability, missingness, or analysis completeness |
Derived variables | Analysis dataset/specification | Should generally be specification-driven rather than redundantly collected as raw eCRF fields |
*Required elements
Hybrid Database Designs
Hybrid database designs where limited operational data are collected in the research database are not automatically prohibited, but they should be intentional and documented.
A hybrid design may be acceptable when the study is small or simple, PHI/PII is limited and carefully controlled, the same user group legitimately needs access to both operational and research data, and operational fields are clearly identified as non-analysis data.
A hybrid design should not be used if it creates unnecessary access, export, privacy, data cleaning, change management, or quality-control risks.
References / Supporting Standards
NIH Data Management and Sharing Policy / NIH expectations for data sharing and reuse
NIH Common Data Elements resources, where applicable
RDCRN-specific network operations and data management resources
CDASH Standards