align_inputs ¶
align_inputs(
names: object, spec: LayeredSpec | AdjacencySpec
) -> np.ndarray
Return an integer index that orders features to spec inputs.
Feature-table labels rarely match the input-node order the
parsed edgelist fixes. This returns a 1-D int64 index
into the caller's feature axis so that axis can be gathered
into that order. It does not take the matrix, copy sample
rows, or densify. Apply the index on whatever holds X:
X[:, col] for numpy, scipy CSR/CSC, AnnData .X in
those formats, or a tensor; df.to_numpy()[:, col] for
a DataFrame. COO-style sparse layouts do not support
integer column indexing; convert first. For a
LayeredSpec, a node with width greater than 1 repeats
its column index across those units, so the length is
spec.layer_dims[0]. On CSR/CSC that repeat copies
those columns and stays sparse; it is not a view. For an
AdjacencySpec the same repeat uses node_widths, so
the length is len(spec.input_index).
Call it after parsing, once, instead of hand-ordering
columns. DataFrames, tensors, and matrices are rejected.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
sequence of labels
|
Feature-axis labels, in the order they currently sit on
the matrix: |
required |
|
LayeredSpec or AdjacencySpec
|
Parsed edgelist whose |
required |
Returns:
| Type | Description |
|---|---|
numpy.ndarray of dtype int64, shape (width,)
|
Positions into |
Raises:
| Type | Description |
|---|---|
Kpnn2Error
|
If |
See Also
parse_layered : Builds the LayeredSpec whose
input_nodes set this column order.
parse_adjacency : Builds the AdjacencySpec for the packed
layout, where the result needs scattering first.
gather_hop_inputs : Assembles a later hop's input; layer 0
comes from indexing with this result instead.
Notes
PyTorch never sees feature names. Passing a hand-stacked
array into the model can silently wire the wrong features if
the column order differs, which is what this guards against.
A tensor whose columns already follow spec.input_nodes
needs no alignment and goes straight to the model. AnnData,
numpy matrices, and scipy sparse matrices are not accepted
as names; pass the labels that sit on that axis and
apply the index on the matrix. scipy CSR/CSC and AnnData
.X in those formats stay sparse under X[:, col]
until the caller densifies one row block. COO and similar
layouts cannot be indexed that way. Width greater than 1
repeats indices; on CSR/CSC that copies the duplicated
columns.
The returned length for a LayeredSpec is
spec.layer_dims[0], which equals
len(spec.input_nodes) only when every input node has
width 1. hops[0] reads layer 0 alone, so a row block
indexed with this result feeds the first hop directly and
needs no gathering. For an AdjacencySpec it is not
the state width: to_mask() is
(state_dim, state_dim) over every node's units, while
the index covers the input units only
(len(spec.input_index)). Scatter the gathered columns
into the spec.state_dim-wide state vector with
state.index_copy(-1, input_index, x) before calling
MaskedLinear(spec.to_mask()). input_index is a
buffer, torch.as_tensor(spec.input_index).
Examples:
Extra names are ignored and remaining columns are reordered:
>>> import numpy as np
>>> import pandas as pd
>>> import kpnn2
>>> edgelist = pd.DataFrame(
... {
... "source": ["A", "H"],
... "target": ["H", "C"],
... }
... )
>>> spec = kpnn2.parse_layered(edgelist)
>>> spec.input_nodes
('A',)
>>> names = ["unused", "A"]
>>> col = kpnn2.align_inputs(names, spec)
>>> col.dtype
dtype('int64')
>>> col.tolist()
[1]
>>> values = np.array([[9.0, 0.5], [8.0, 1.5]])
>>> values[:, col].tolist()
[[0.5], [1.5]]
An AdjacencySpec works the same way, but the index is
len(spec.input_index) long and the gathered columns must be
scattered into the state vector before they reach
MaskedLinear(spec.to_mask()):
>>> import torch
>>> cyclic = pd.DataFrame(
... {
... "source": ["x", "a", "b", "a"],
... "target": ["a", "b", "a", "y"],
... }
... )
>>> state_spec = kpnn2.parse_adjacency(cyclic)
>>> col = kpnn2.align_inputs(["x"], state_spec)
>>> col.tolist(), tuple(state_spec.to_mask().shape)
([0], (4, 4))
>>> x = torch.tensor([[0.5], [1.5]])[:, col]
>>> state = torch.zeros(
... 2,
... state_spec.state_dim,
... )
>>> index = torch.as_tensor(state_spec.input_index)
>>> state = state.index_copy(
... -1,
... index,
... x,
... )
>>> state.tolist()
[[0.0, 0.0, 0.5, 0.0], [0.0, 0.0, 1.5, 0.0]]
A tensor is not accepted; pass the names that label it:
>>> t = torch.tensor([[0.5], [1.5]])
>>> kpnn2.align_inputs(t, spec)
Traceback (most recent call last):
...
Kpnn2Error: 'names' is a tensor; pass the feature names that ...