PackedLinear ¶
PackedLinear(
source_index: object,
target_index: object,
out_features: int,
in_features: int,
bias: bool = True,
*,
identity: str | None = None,
constraint: Module | None = None,
generator: Generator | None = None
)
Bases: Module
Affine map with one trainable scalar per live edge.
An nn.Linear-style layer (call layer(x); not a
subclass, not a full model) for a graph laid out by
parse_layered or parse_adjacency. On a hop it stores
one weight per live edge of that hop instead of the dense
(out, in) rectangle MaskedLinear(hop.to_mask())
would. On an AdjacencySpec it stores one weight per
graph edge instead of an (n, n) square. Reach for it
when that rectangle or square strains RAM. On an
AdjacencySpec, input nodes have no incoming edges, so
writing inputs into the state each step is the caller's job.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
torch.Tensor or sequence of int
|
1-D integer indices of length |
required |
|
torch.Tensor or sequence of int
|
1-D integer indices of the same length. Entry |
required |
|
int
|
Width of the output axis. Must be a positive int. Note
the order: |
required |
|
int
|
Width of the input axis. Must be a positive int. |
required |
|
bool
|
If |
True
|
|
str or None
|
Opaque checkpoint identity, typically
|
None
|
|
Module or None
|
Optional per-entry map on the packed |
None
|
|
Generator or None
|
Isolated RNG for |
None
|
Attributes:
| Name | Type | Description |
|---|---|---|
in_features |
int
|
Number of input columns. |
out_features |
int
|
Number of output columns. |
nnz |
int
|
Number of live edges, |
weight |
Parameter
|
Trainable packed weights of shape |
source_index |
Tensor
|
Int64 buffer of input columns, length |
target_index |
Tensor
|
Int64 buffer of output rows, length |
constraint |
Module or None
|
The constructor |
bias |
Parameter | None
|
Trainable bias, or |
identity |
str | None
|
The constructor |
Raises:
| Type | Description |
|---|---|
Kpnn2Error
|
If the indices are empty, not 1-D integers, mismatched in
length, out of range, or duplicated as
|
See Also
MaskedLinear : Dense (out_features, in_features) weight;
the GEMM hatch when that rectangle fits.
Hop : Supplies per-layer source_index / target_index
on a LayeredSpec.
AdjacencySpec : Supplies graph-wide source_index /
target_index; this layer takes those tuples, not the
spec object.
PackedMultiheadAttention : Attention over the same packed
pairs, when the update is a contraction rather than one
scalar per edge.
scatter_hop_outputs : Split a transposed hop's concatenated
output back onto source layers.
torch.nn.Linear : Dense equivalent, and the reference for
shapes, bias, calling the module, and training.
Notes
Forward gathers x[..., source_index], multiplies by
effective_weight() (constraint(weight) when
constraint is set, else weight), and index_adds
into zeros of shape (..., out_features), adding bias
when present. x is an ordinary dense activation tensor
whose last dimension must be in_features, as for
nn.Linear. The gather reads columns by position, so a
wider tensor would be read without complaint; forward
raises instead. A stale align_inputs index of the same
width after a reparse still cannot be detected here:
recompute the index on the new spec. Nothing scatters into
a dense (out, in) matrix, nothing imports
torch.sparse, and no tensor subclass is involved, so
torch.compile(layer, fullgraph=True) traces it. Index
buffers stay integer after .half() / bfloat16 /
.double(); weight and bias follow the module
floating dtype like nn.Linear. torch.autocast is
unsupported: forward disables it and casts x to
the parameter dtype so AMP cannot mix Half into
index_add. Cast the module with .to(dtype=...)
(or .half() / .double()) instead.
weight is the unconstrained tensor, so optimizer
weight_decay pulls it toward 0. Under softplus that
pulls the effective weight toward ln 2, not toward 0.
Give a constrained layer weight_decay=0 in its param
group and penalize effective_weight() in the loss
instead. A right_inverse on the constraint keeps the
degree-aware init; a constraint without one can initialize
weight itself from init_bound().
reset_parameters uses per-row packed degree as
fan_in, not full in_features. A row with
fan_in == 0 has no packed weights and its bias stays 0;
input nodes of an AdjacencySpec are exactly that case,
and this layer does not invent identity connections for them.
Optional generator isolates those draws from other torch
RNG consumers.
state_dict keys are weight, optional bias,
source_index, target_index, index_digest, and
identity when the constructor was given one.
index_digest is a 1-D CPU uint8 tensor of length 32:
the SHA-256 of the live index buffers' int64 C-contiguous
bytes plus out_features and in_features as
fixed-width integers, not a registered buffer. A missing
digest is not an error, even with strict=True. The
digest catches same-shape rewiring. A rename that leaves
the packed index pattern unchanged is caught by
identity when callers pass spec.fingerprint.
copy.deepcopy works.
transpose() is the tied-autoencoder helper: it swaps
the index buffers and the feature sizes so packed slot
i is still the same live edge, then shares or copies
weight and constraint together. Bias is never tied.
Do not reparse a reversed edgelist and assign
dec.weight = enc.weight: that permutes slots. On a hop
whose source axis is several layers, split the transposed
output with scatter_hop_outputs.
Examples:
An input feeding a two-node feedback core plus one output, stepped once over the shared state vector:
>>> import pandas as pd
>>> import torch
>>> import kpnn2
>>> edgelist = pd.DataFrame(
... {
... "source": ["x", "a", "b", "a"],
... "target": ["a", "b", "a", "y"],
... }
... )
>>> spec = kpnn2.parse_adjacency(edgelist)
>>> n = spec.state_dim
>>> core = kpnn2.PackedLinear(
... spec.source_index,
... spec.target_index,
... n,
... n,
... identity=spec.fingerprint,
... )
>>> core.in_features, core.out_features, core.nnz
(4, 4, 4)
>>> state = torch.zeros(2, n)
>>> state[:, spec.input_index] = torch.ones(2, 1)
>>> relu = torch.nn.ReLU()
>>> state = relu(core(state))
>>> tuple(state.shape)
(2, 4)
>>> tuple(state[:, spec.output_index].shape)
(2, 1)
A layered hop uses the same constructor on that hop's packed indices:
>>> layered = kpnn2.parse_layered(
... pd.DataFrame(
... {
... "source": ["A", "H", "A"],
... "target": ["H", "C", "C"],
... }
... )
... )
>>> hop = layered.hops[1]
>>> layer = kpnn2.PackedLinear(
... hop.source_index,
... hop.target_index,
... hop.out_features,
... hop.in_features,
... )
>>> layer.nnz
2
Methods:¶
effective_weight ¶
effective_weight() -> torch.Tensor
Return the (nnz,) packed weights forward uses.
constraint(weight) when constraint is set, else
weight itself. Recomputed on every call and
differentiable, so it is the tensor to read, export, or
penalize as the layer's live edge weights; index it with
edge_location slots. MaskedLinear has the same
method on its dense rectangle.
Returns:
| Type | Description |
|---|---|
Tensor
|
The effective packed weights, one per live edge. |
init_bound ¶
init_bound() -> torch.Tensor
Return the degree-aware init bound of every packed slot.
Slot i gets 1 / sqrt(fan_in) of its output row
target_index[i], the bound reset_parameters draws
that slot's effective weight from. Use it to initialize a
constraint that has no right_inverse yourself.
MaskedLinear has the same method on its rectangle.
Returns:
| Type | Description |
|---|---|
Tensor
|
Shape |
reset_parameters ¶
reset_parameters(
generator: Generator | None = None,
) -> None
Initialize from per-row packed degree, not full width.
For output row j, fan_in is the number of packed
edges with target_index == j (bincount,
minlength=out_features). Each live edge into that
row, and bias[j] if present, is drawn uniformly from
[-1 / sqrt(fan_in), 1 / sqrt(fan_in)]
(init_bound()). If fan_in == 0, that row has no
packed weights and bias[j] stays 0.
The draw is the degree-aware value for the effective
weight. Without constraint, or when constraint
has no right_inverse, weight stores the draw
itself. When constraint defines right_inverse
(the torch.nn.utils.parametrize convention),
weight stores constraint.right_inverse(draw), so
effective_weight() keeps the degree-aware scale. The
draws, and so the RNG stream, are the same either way.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
Generator or None
|
Isolated RNG for these draws. |
None
|
Raises:
| Type | Description |
|---|---|
Kpnn2Error
|
If |
transpose ¶
transpose(
bias: bool = True,
*,
tie: bool = True,
identity: str | None = None,
generator: Generator | None = None
) -> PackedLinear
Return a packed layer that applies the same edges backwards.
Packed slot i stays the same live edge: the 1-D
weight is not permuted. The new layer reads the
former output axis and writes the former input axis.
That is the packed analogue of enc.weight.T for a
tied autoencoder. With tie=True the whole live map
is shared: weight and the constraint module.
Bias is never shared.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
|
bool
|
If |
True
|
|
bool
|
If |
True
|
|
str or None
|
Checkpoint identity for the new layer, typically
|
None
|
|
Generator or None
|
Forwarded to the inner |
None
|
Returns:
| Type | Description |
|---|---|
PackedLinear
|
New module: |
Raises:
| Type | Description |
|---|---|
Kpnn2Error
|
If |
Notes
Do not parse a reversed edgelist and assign
dec.weight = enc.weight. Packed order is
lexicographic by (source name, target name) of
each spec, so those slots do not line up. This
method keeps slot i as the same edge.
On a layered hop, the transposed output is
hop.in_features wide. Split it with
scatter_hop_outputs when the hop reads several
source layers; add those pieces into the caller's
saved dict, because two reversed hops may write
the same earlier layer. An AdjacencySpec map is
already n-wide; no gather or scatter.
MaskedLinear has no packed slots: use
F.linear(h, layer.weight.T, dec_bias).
Examples:
Slot i is the same edge after a transpose, so
sharing weight is the tied map:
>>> import torch
>>> import kpnn2
>>> layer = kpnn2.PackedLinear(
... [0, 1],
... [0, 0],
... 1,
... 2,
... bias=False,
... )
>>> with torch.no_grad():
... layer.weight[:] = torch.tensor([2.0, 3.0])
>>> mirrored = layer.transpose(bias=False)
>>> mirrored.weight is layer.weight
True
>>> mirrored.in_features, mirrored.out_features
(1, 2)
>>> layer(torch.tensor([[1.0, 4.0]])).tolist()
[[14.0]]
>>> mirrored(torch.tensor([[1.0]])).tolist()
[[2.0, 3.0]]
forward ¶
forward(x: Tensor) -> torch.Tensor
Gather live inputs, scale by packed weights, index_add.
contrib = x[..., source_index] * effective_weight(),
then index_add into zeros of shape
(..., out_features). Adds bias when present.
Packed 1-D weights, one per live edge; not
torch.sparse; forward is index_add.
torch.autocast is unsupported: this path
disables it and casts x to the parameter dtype.
Raises Kpnn2Error when x is not a tensor or its
last dimension is not in_features.