rafmacalaba's picture
Initial deploy of PRWP Dataset Extraction Dashboard
5567216 verified
|
Raw
History Blame Contribute Delete
2.16 kB

Neo4j Graph Database Import Guide

This folder contains Neo4j-import-ready CSV files containing:

  • Nodes: Papers, Authors, and deduplicated Datasets.
  • Relationships: Authorship (:AUTHORED) and dataset citations (:MENTIONS).

CSV File List

  1. nodes_papers.csv: Contains Policy Research Working Papers.
  2. nodes_datasets.csv: Contains canonicalized datasets.
  3. nodes_authors.csv: Contains unique authors.
  4. edges_authored.csv: Maps Authors to Papers.
  5. edges_mentions.csv: Maps Papers to Datasets with page numbers, context, and model confidence scores.

Import Cypher Queries

To import these files into your Neo4j instance, place the CSV files in your Neo4j project's import/ directory, open the Neo4j Browser, and execute the following queries:

1. Create Constraints

CREATE CONSTRAINT UNIQUE_paper FOR (p:Paper) REQUIRE p.id IS UNIQUE;
CREATE CONSTRAINT UNIQUE_dataset FOR (d:Dataset) REQUIRE d.id IS UNIQUE;
CREATE CONSTRAINT UNIQUE_author FOR (a:Author) REQUIRE a.id IS UNIQUE;

2. Load Nodes

// Load Papers
LOAD CSV WITH HEADERS FROM 'file:///nodes_papers.csv' AS row
MERGE (p:Paper {id: row.id})
SET p.title = row.title,
    p.year = toInteger(row.year);

// Load Datasets
LOAD CSV WITH HEADERS FROM 'file:///nodes_datasets.csv' AS row
MERGE (d:Dataset {id: row.id})
SET d.name = row.name,
    d.acronym = row.acronym,
    d.typology = row.typology,
    d.producer = row.producer,
    d.geography = row.geography;

// Load Authors
LOAD CSV WITH HEADERS FROM 'file:///nodes_authors.csv' AS row
MERGE (a:Author {id: row.id})
SET a.name = row.name;

3. Load Relationships

// Load AUTHORED
LOAD CSV WITH HEADERS FROM 'file:///edges_authored.csv' AS row
MATCH (a:Author {id: row.author_id})
MATCH (p:Paper {id: row.paper_id})
MERGE (a)-[:AUTHORED]->(p);

// Load MENTIONS
LOAD CSV WITH HEADERS FROM 'file:///edges_mentions.csv' AS row
MATCH (p:Paper {id: row.paper_id})
MATCH (d:Dataset {id: row.dataset_id})
MERGE (p)-[:MENTIONS {
    pages: split(row.pages, ';'),
    context: row.context,
    confidence: toFloat(row.confidence)
}]->(d);