Harmonize gene identifiers using an OrgDb annotation package
Source:R/harmonize_gene_ids.R
harmonize_gene_ids.RdHarmonizes a character vector of gene identifiers to a target identifier
type using an OrgDb object.
The function is intentionally strict and follows a "don't guess"
strategy. Input identifiers are kept unchanged when they already match
to_keytype. Otherwise, candidate source identifier types in
from_keytypes are checked.
Harmonization is only performed when exactly one source identifier type
matches and this yields exactly one target identifier. All non-unique or
unresolved cases are returned as NA in the harmonized
column.
Missing input values are accepted and returned unchanged in the output structure.
The returned data.frame is always row-aligned with genes,
including duplicated and missing values.
Usage
harmonize_gene_ids(
genes,
OrgDb,
to_keytype,
from_keytypes = c("ALIAS", "ENTREZID", "ENSEMBL")
)Arguments
- genes
A character vector of gene identifiers to harmonize.
- OrgDb
An
OrgDbobject providing organism-specific gene identifier mappings.- to_keytype
A non-empty character scalar giving the target identifier type.
- from_keytypes
A character vector of candidate source identifier types to try when an input value does not already match
to_keytype. Duplicates are removed internally, andto_keytypeis excluded if present.
Value
A data.frame with the following columns:
inputThe original input identifier.
matches_targetLogical indicator of whether the input already matched
to_keytype.from_match_statusStatus of matching against
from_keytypes. One of"none","unique","ambiguous", orNAwhen the input already matchedto_keytype.to_match_statusStatus of mapping from the uniquely matched source identifier type to
to_keytype. One of"none","unique","ambiguous", orNAwhen no target mapping was attempted.harmonizedThe harmonized identifier, or
NAwhen no unique harmonization was possible.
Examples
if (requireNamespace("org.EcK12.eg.db", quietly = TRUE)) {
genes <- c(
"b0002", # SYMBOL
"thrA", # ALIAS
"944742", # ENTREZID
"not_a_gene",
NA_character_
)
res <- harmonize_gene_ids(
genes = genes,
OrgDb = org.EcK12.eg.db::org.EcK12.eg.db,
to_keytype = "ENTREZID",
from_keytypes = c("SYMBOL", "ALIAS")
)
head(res)
# Extract harmonized identifiers
res$harmonized
}
#> [1] NA NA "944742" NA NA