About Omics Data

The GS cohort is a rich, world-leading multi-omics resource. We have genome-wide genotype data and DNA methylation data from both recruitment waves of the cohort, as well as proteomic and metabolomic data from the original recruitment wave.

Original GS Recruitment (2006-2011)

DNA from over 20,000 participants has been analysed by high density genome-wide chip genotyping, Illumina OmniExpress SNP GWAS (700k) and exome chip (250K), with low failure and high call rates. The data was cleaned using quality score metrics and quality control analyses were performed, filtering for missingness, MAF, HWE and sex mismatches. Ancestry outliers were removed following principal component analysis of the data merged with the 1000 Genomes reference population. Genetic profiles have been imputed using three different reference panels: 1000 Genomes, Haplotype Reference Consortium and Trans-Omics for Precision Medicine.

Pedigrees were constructed using relationship information provided by study participants and validated with genetic kinship information following genotyping. The cohort contains 1,361 singletons (with no relatives in the study) and 5,501 families of at least two people, with a mean size of 4.1 family members. Sample identity was verified against recorded gender and pedigree and data checked for unknown relationships based on estimated identity-by-descent.

DNA methylation (DNAm) data for 18,858 participants has been generated using the Illumina HumanMethylationEPICv1 BeadChip array at >850,000 CpG sites, from blood samples.

Protein levels have been quantified in plasma samples from 1,065 participants using the SOMAscan V.4 array from SomaLogic. Liquid chromatography mass spectrometry (LC-MS) proteomics data, with measures for 325 proteins, is available for 18,826 participants.

Quantification of 54 urinary metabolite biomarkers in 2,743 GS participants has been conducted by Nightingale Health using nuclear magnetic resonance.

GS has contributed to multiple genome-wide association study (GWAS) meta-analyses. Summary statistics and polygenic scores with GS data excluded ("Leave-One-Out") have been calculated for select phenotypes (major depression and schizophrenia from the Psychiatric Genomics Consortium). 

Omics dataN (%)Sample typeMeasurement tool
Genotype20,019 (83%)Blood

Illumina HumanOmniExpressExome8V.1-2_A array

Illumina HumanOmniExpressExome-8V.1_A array

Beadstudio-Gencall V.3 genotype caller

Methylation18,858 (79%)BloodIllumina HumanMethylationEPIC BeadChip array
Proteomics

1,065 (4%)

18,826 (78%)

Blood

Blood

SomaScan SOMAscan V.4 array 

Liquid chromatography mass spectrometry

Metabolomics2,743 (11%)UrineNuclear magnetic resonance spectroscopy

NextGenScot Recruitment (2022-2025)

DNA from saliva samples of over 10,000 participants has been genotyped using the Illumina GSA-MD-48v4.0_A2 array. The data was cleaned using quality score metrics and quality control analyses were performed, filtering for missingness, MAF, HWE and sex mismatches. Ancestry outliers were removed following principal component analysis of the data merged with the 1000 Genomes reference population. The full multi-ancestry dataset has been imputed using the 1000 Genomes reference panel and the European only subset has been imputed with the Haplotype Reference Consortium reference panel.

DNA methylation data for over 10,000 participants has been generated using the Illumina MSA-48v1.0_A1 array at >200,000 CpG sites. Quality control processing was carried out using the SeSAMe R package.

Omics dataN (%)Sample typeMeasurement tool
Genotype 10,832SalivaIllumina GSA-MD-48v4.0_A2 array
Methylation 10,799SalivaIllumina MSA-48v1.0_A1 array