Skip to contents

Harmonizes a character vector of gene identifiers to a target identifier type using an OrgDb object.

The function is intentionally strict and follows a "don't guess" strategy. Input identifiers are kept unchanged when they already match to_keytype. Otherwise, candidate source identifier types in from_keytypes are checked.

Harmonization is only performed when exactly one source identifier type matches and this yields exactly one target identifier. All non-unique or unresolved cases are returned as NA in the harmonized column.

Missing input values are accepted and returned unchanged in the output structure.

The returned data.frame is always row-aligned with genes, including duplicated and missing values.

Usage

harmonize_gene_ids(
  genes,
  OrgDb,
  to_keytype,
  from_keytypes = c("ALIAS", "ENTREZID", "ENSEMBL")
)

Arguments

genes

A character vector of gene identifiers to harmonize.

OrgDb

An OrgDb object providing organism-specific gene identifier mappings.

to_keytype

A non-empty character scalar giving the target identifier type.

from_keytypes

A character vector of candidate source identifier types to try when an input value does not already match to_keytype. Duplicates are removed internally, and to_keytype is excluded if present.

Value

A data.frame with the following columns:

input

The original input identifier.

matches_target

Logical indicator of whether the input already matched to_keytype.

from_match_status

Status of matching against from_keytypes. One of "none", "unique", "ambiguous", or NA when the input already matched to_keytype.

to_match_status

Status of mapping from the uniquely matched source identifier type to to_keytype. One of "none", "unique", "ambiguous", or NA when no target mapping was attempted.

harmonized

The harmonized identifier, or NA when no unique harmonization was possible.

Examples

if (requireNamespace("org.EcK12.eg.db", quietly = TRUE)) {
    genes <- c(
        "b0002",    # SYMBOL
        "thrA",     # ALIAS
        "944742",   # ENTREZID
        "not_a_gene",
        NA_character_
    )

    res <- harmonize_gene_ids(
        genes = genes,
        OrgDb = org.EcK12.eg.db::org.EcK12.eg.db,
        to_keytype = "ENTREZID",
        from_keytypes = c("SYMBOL", "ALIAS")
    )

    head(res)

    # Extract harmonized identifiers
    res$harmonized
}
#> [1] NA       NA       "944742" NA       NA