Overview
GOcontext provides tools for constructing restricted Gene Ontology (GO) graphs prior to enrichment analysis.
In standard GO enrichment workflows, all GO terms within an ontology (Biological Process, Molecular Function, or Cellular Component) are tested simultaneously. However, for many biological questions, large parts of the ontology are irrelevant to the experimental system under study. Testing these terms unnecessarily increases the number of hypotheses and therefore the multiple testing burden.
GOcontext addresses this by enabling users to construct context-specific GO subgraphs before performing enrichment analysis. The package:
- loads the GO structure from
GO.db - attaches organism-specific GO annotations from an
OrgDb - restricts the ontology graph to mapped terms or selected branches
- exports the resulting mappings for enrichment tools
By reducing the number of tested GO terms while preserving the underlying directed acyclic graph (DAG) structure, enrichment analyses can focus on biologically relevant parts of the ontology while maintaining a transparent and reproducible GO universe.
The typical workflow is:
load_go() → attach_org() → filter_mapped() → subset_go() → enrichment
1. Load a GO ontology
GOcontext loads GO directly from the Bioconductor package GO.db. Only a single ontology (BP, MF, or CC) is loaded at a time.
go_bp <- GOcontext::load_go("BP")##
The returned object contains:
- GO term metadata
- parent–child edges defining the GO DAG
- adjacency lists for graph traversal
- the GO.db version used to construct the graph
2. Attach organism annotations
GO enrichment requires mappings between GO terms and genes. These are obtained from an organism annotation package (OrgDb).
go_bp <- GOcontext::attach_org(
go = go_bp,
OrgDb = org.EcK12.eg.db
)This attaches a GO-to-gene mapping table to the GO graph object.
3. Restrict the graph to mapped terms
Many GO terms have no gene annotations for a given organism. These terms cannot contribute to enrichment results and can therefore be removed before further analysis.
go_bp_mapped <- GOcontext::filter_mapped(go_bp)The returned object is a GOSubgraph containing only terms that have at least one gene mapping.
This step often reduces the ontology size substantially.
4. Restrict the graph to selected branches
GOcontext allows restricting the ontology graph to specific regions.
Two modes are supported:
- keep – retain selected terms and all their descendants
- exclude – remove selected terms and their descendants
Example: restrict the graph to the descendants of a particular GO term.
go_sub <- GOcontext::subset_go(
go = go_bp_mapped,
ids = "GO:0006740",
mode = "keep"
)The returned GOSubgraph represents the restricted GO graph.
The object contains:
- retained GO terms
- removed GO terms
- induced parent–child edges
- updated adjacency lists
- the attached GO-to-gene mapping
5. Export mappings for enrichment
After restricting the graph, GO annotations can be exported in formats commonly used by enrichment tools.
TERM2GENE mapping
term2gene <- GOcontext::as_term2gene(
go = go_sub,
minGSSize = 10,
maxGSSize = 500
)
head(term2gene)## term gene
## 1 GO:0006098 944748
## 2 GO:0006098 945865
## 3 GO:0006098 946370
## 4 GO:0006098 946398
## 5 GO:0006098 946554
## 6 GO:0006098 947006
TERM2NAME mapping
term2name <- GOcontext::as_term2name(
go = go_sub,
minGSSize = NULL,
maxGSSize = NULL
)
head(term2name)## term name
## 1 GO:0006098 pentose-phosphate shunt
## 2 GO:0006740 NADPH regeneration
## 3 GO:0009051 pentose-phosphate shunt, oxidative branch
## 4 GO:0009052 pentose-phosphate shunt, non-oxidative branch
These tables can be supplied directly to clusterProfiler::enricher().
6. Run enrichment analysis
Example workflow with clusterProfiler:
library(clusterProfiler)
enrichment_results <- clusterProfiler::enricher(
gene = gene_list,
TERM2GENE = term2gene,
TERM2NAME = term2name
)Because the GO graph was restricted beforehand, enrichment is performed on a smaller and biologically focused hypothesis space.
Interactive exploration
GOcontext also provides a console browser for exploring the GO hierarchy interactively.
res <- GOcontext::browse_go(go_bp)The browser allows navigation through the GO DAG and selection of terms of interest. The selected terms can then be used as input for subset_go().
Harmonizing gene identifiers
Gene lists used for enrichment analyses often contain mixed identifier types (e.g. gene symbols, Entrez IDs, aliases). Before running enrichment, it can be useful to check whether the identifiers match the annotation database used to define the GO universe.
GOcontext provides harmonize_gene_ids() to diagnose and harmonize such identifiers.
genes <- c(
"b0002", # SYMBOL
"thrA", # ALIAS
"944742", # ENTREZID
"not_a_gene",
NA_character_
)
res <- GOcontext::harmonize_gene_ids(
genes = genes,
OrgDb = org.EcK12.eg.db,
to_keytype = "ENTREZID",
from_keytypes = c("SYMBOL", "ALIAS")
)The result is a table aligned with the input vector showing whether each identifier matched the target key type, matched a fallback type, or could not be mapped.
head(res)## input matches_target from_match_status to_match_status harmonized
## 1 b0002 FALSE none <NA> <NA>
## 2 thrA FALSE ambiguous <NA> <NA>
## 3 944742 TRUE <NA> <NA> 944742
## 4 not_a_gene FALSE none <NA> <NA>
## 5 <NA> FALSE <NA> <NA> <NA>
If a unique mapping exists, the harmonized identifiers can be extracted for downstream analyses.
gene_ids <- res$harmonizedSummary
GOcontext enables reproducible construction of restricted GO graphs prior to enrichment analysis.
A typical workflow consists of:
- Loading GO using
load_go() - Attaching organism annotations with
attach_org() - Restricting the ontology to mapped terms using
filter_mapped() - Restricting the graph structure with
subset_go() - Exporting GO mappings with
as_term2gene()andas_term2name() - Running enrichment analysis
By explicitly defining the tested GO universe before enrichment, GOcontext reduces the multiple testing burden while preserving the ontology structure and ensuring transparent analysis workflows.
## R version 4.5.2 (2025-10-31)
## Platform: x86_64-pc-linux-gnu
## Running under: Ubuntu 22.04.5 LTS
##
## Matrix products: default
## BLAS: /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3
## LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/libopenblasp-r0.3.20.so; LAPACK version 3.10.0
##
## locale:
## [1] LC_CTYPE=C.UTF-8 LC_NUMERIC=C LC_TIME=C.UTF-8
## [4] LC_COLLATE=C.UTF-8 LC_MONETARY=C.UTF-8 LC_MESSAGES=C.UTF-8
## [7] LC_PAPER=C.UTF-8 LC_NAME=C LC_ADDRESS=C
## [10] LC_TELEPHONE=C LC_MEASUREMENT=C.UTF-8 LC_IDENTIFICATION=C
##
## time zone: UTC
## tzcode source: system (glibc)
##
## attached base packages:
## [1] stats4 stats graphics grDevices utils datasets methods
## [8] base
##
## other attached packages:
## [1] org.EcK12.eg.db_3.22.0 AnnotationDbi_1.72.0 IRanges_2.44.0
## [4] S4Vectors_0.48.0 Biobase_2.70.0 BiocGenerics_0.56.0
## [7] generics_0.1.4 knitr_1.51 GOcontext_0.1.0
## [10] BiocStyle_2.38.0
##
## loaded via a namespace (and not attached):
## [1] bit_4.6.0 jsonlite_2.0.0 crayon_1.5.3
## [4] compiler_4.5.2 BiocManager_1.30.27 blob_1.3.0
## [7] Biostrings_2.78.0 jquerylib_0.1.4 Seqinfo_1.0.0
## [10] png_0.1-8 systemfonts_1.3.2 textshaping_1.0.5
## [13] yaml_2.3.12 fastmap_1.2.0 XVector_0.50.0
## [16] R6_2.6.1 bookdown_0.46 desc_1.4.3
## [19] DBI_1.3.0 bslib_0.10.0 rlang_1.1.7
## [22] KEGGREST_1.50.0 cachem_1.1.0 xfun_0.56
## [25] fs_1.6.7 sass_0.4.10 bit64_4.6.0-1
## [28] RSQLite_2.4.6 memoise_2.0.1 cli_3.6.5
## [31] pkgdown_2.2.0 digest_0.6.39 GO.db_3.22.0
## [34] lifecycle_1.0.5 vctrs_0.7.1 evaluate_1.0.5
## [37] ragg_1.5.1 rmarkdown_2.30 httr_1.4.8
## [40] pkgconfig_2.0.3 tools_4.5.2 htmltools_0.5.9