Node2Vec embeddings for the edge_ML metabolomics graph

128-dimensional Node2Vec embeddings for the 18,494 nodes of an undirected metabolite co-response graph with 2,709,209 edges. Held-out link prediction reaches AUC 0.988, against 0.920 for a degree-only baseline.

The graph, node properties and full pipeline are in the companion dataset repository.

Files

File Contents Size
edge_ML_expected_ge5_n2v.pt {embedding: [18494, 128] float32, node_id: [18494], args: {...}} 9.7 MB
node2vec_model.py Model definition, training loop, embedding export β€”

Using it

import torch

ck = torch.load("edge_ML_expected_ge5_n2v.pt", weights_only=False)
z = ck["embedding"]                      # [18494, 128] float32
index = {nid: i for i, nid in enumerate(ck["node_id"])}
v = z[index["MTBLS1405_0002_00003332"]]  # one node's vector

node_id[i] is the original string ID for row i; the order is lexicographic over the union of the graph's two endpoint columns, matching the dataset's graph object. Scores were computed with cosine similarity, which is also the metric to use downstream.

Training

Parameter Value
embedding_dim 128
walk_length 20
context_size 10
walks_per_node 10
num_negative_samples 1
p, q 1.0, 1.0 (unbiased walks)
Batch size 128 seed nodes, 145 batches per epoch
Optimiser SparseAdam, lr 0.01
Epochs 20
Parameters 2,367,232 (18,494 Γ— 128)

Loss fell from 9.92 at initialisation to 0.880, flat from about epoch 14, at roughly 0.9 s/epoch on one H100. sparse=True on the model is what allows SparseAdam; changing either requires changing the other.

uv run python node2vec_model.py --epochs 20

Node2Vec requires pyg-lib >= 0.6.0 for its random-walk kernel, which is not on PyPI; the dataset repository's pyproject.toml pins pyg-lib 0.9.0+pt214cu130 from data.pyg.org.

Evaluation

200,000 sampled positive edges against 200,000 non-edges verified absent from the full edge set, scored by cosine similarity, AUC by the Mann-Whitney rank identity.

Model Scored edges Score AUC
90/10 retrain Held-out 10%, never seen cosine 0.9880
90/10 retrain Its own training edges cosine 0.9892
90/10 retrain Held-out 10%, never seen degree product d_u Γ— d_v 0.9201
Full graph (this release) Its own training edges cosine 0.9892
Full graph (this release) Its own training edges dot product 0.9868

The held-out row is the one that matters: a second model was trained from scratch on 90% of the edges and scored on the 10% it never saw. Held-out 0.9880 against in-sample 0.9892 is a gap of 0.001, so the model learns graph structure rather than memorising pairs. The degree baseline matters because the graph is dense (median degree 90) β€” a high AUC that merely reproduced the degree distribution would carry little information.

Other checks on the released embeddings:

  • Neighbourhood recovery β€” of each node's 10 nearest embeddings, 49.9% are true graph neighbours against 1.6% expected by chance (31.7Γ—); at top-50, 41.3% (26.2Γ—).
  • Embedding health β€” all finite; L2 norms 0.94 / 1.92 / 9.49 (min / median / max); per-dimension standard deviation 0.15–0.28, so no dead dimensions; mean cosine over 200,000 random pairs is 0.0038, ruling out collapse.
  • Species purity β€” 90.7% of all nodes have ten nearest embeddings sharing their species, rising above 98% for the three largest species and falling to 68–79% for species with a few hundred nodes.

Limitations

  • Topology only. The walks are unweighted, so neither the graph's edge_attr (OddsRatio_log2, ChiTestsPValue) nor its node features x influence these embeddings. Letting association strength steer the walks needs a weighted sampler or a pre-thresholded edge set; using the node features needs a message-passing model.
  • Transductive. Node2Vec learns one vector per node in a fixed graph. There is no way to embed a node that was not present at training time.
  • Species and study are entangled. Edges form mostly within a study and a study is normally one species, so the clean species separation partly reflects how the graph was assembled, not an independent biological signal.
  • In the 90/10 evaluation split, 97 low-degree nodes were left isolated in the training graph and their vectors stay near initialisation. That affects only the held-out experiment; the released full-graph model has no isolated nodes.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support