Stage W0: private living prototype12 of 30 manuscripts drafted, 0 reviewed; 50 lexicon entries in draftWhat each later stage would need to show

Mathematical and computational foundations

Knowledge graphs

A graph is a way of writing down who is connected to what. It is not a way of knowing whether the connection is real.

MonographDraftNX-M0615 min read

Draft. This monograph is an unreviewed draft. Its sources have not been checked by a named person and no domain reviewer has approved it. Treat every claim as provisional.

Central question. How does a graph organize knowledge?

Definition and scope

A Knowledge graph is a representation of knowledge as a graph: nodes standing for entities, and edges standing for relationships between them, where each edge has a type and, usually, a direction. Nodes and edges carry identifiers so that the same entity can be referred to from many statements and the statements can be joined. The result can be traversed: from this molecule, to the target it binds, to the cell types that express the target, to the tissues in which those cells are found. The traversal is mechanical. Whether the path corresponds to anything in the world depends entirely on whether the edges were true when written and are still true now.

The term has no single authoritative definition. Hogan and colleagues, in a 2021 survey, settle on "a graph of data intended to accumulate and convey knowledge of the real world, whose nodes represent entities of interest and whose edges represent relations between these entities," and note that this covers a wide range of practice. Ehrlinger and Wöß, surveying earlier attempts, found definitions that required a reasoning engine, definitions that required integration of multiple sources, and definitions that required nothing more than a graph. This monograph uses the broad definition and treats the additional features (an ontology, a reasoner, provenance, integration) as things a knowledge graph may have, each of which should be stated rather than assumed.

The scope is the graph as a representation. The formal semantics that let a graph support inference belong to Description logics. The standards that encode graphs for exchange belong to Semantic web standards. The question of where each edge came from, raised here, is developed in Provenance and evidence.

Key distinctions

A node is not a string. The slogan under which Google introduced its Knowledge Graph in 2012 was "things, not strings." A string, "Agent A," is a label; it can be misspelled, translated, or shared by two different things. A node is an entity with an Identifier, to which labels in any number of languages are attached as data. Two statements about the same node are about the same thing. Two statements that happen to use the same string may not be.

An edge has a type and a direction. "Agent A, Target P" is a pair, not a statement. "Agent A binds Target P" is a statement, and it is a different statement from "Target P binds Agent A." Some relation types are symmetric; most are not. A graph that stores untyped or undirected edges stores associations, which is a legitimate thing to store but is not what this monograph means by knowledge.

A graph is a set of assertions, not a set of facts. Each edge is a claim that someone, or some process, made at some time on some basis. The graph records the claim. It does not, by recording it, make it true, and the absence of an edge does not make the corresponding claim false. A graph that lacks an edge from Agent A to Target Q has not asserted that Agent A does not bind Target Q; it has said nothing. This is the open world assumption, and it is the default in the RDF family of standards. Property graph systems and most relational databases make the opposite, closed world assumption, in which absence means falsity. Which assumption a graph makes must be stated, because queries mean different things under each.

A knowledge graph is not an ontology. A graph records what is related to what. An Ontology records what kinds of things there are and what relations are possible among them. A graph may be built according to an ontology, which then constrains what edges may be drawn and licenses inferences from them. Or it may be built without one, as a graph of whatever its authors chose to connect. Both are called knowledge graphs. The difference is visible in what the graph can be asked.

Historical development

Graphs as mathematical objects date from Euler's 1736 treatment of the bridges of Königsberg, in which the question of whether the seven bridges could be crossed exactly once was answered by abstracting the city to points and lines. Nothing about the water, the stone, or the city survived the abstraction except connectivity, and that was enough. Every knowledge graph inherits this: it is a structure of connectivity, and what it discards is everything else.

Graphs as representations of meaning begin in the late 1960s with Quillian's semantic memory, in which concepts are nodes, associations are links, and the meaning of a concept is the set of nodes reachable from it. The semantic networks of the 1970s generalized the idea, and Sowa's conceptual graphs (1984) gave it a logical foundation. The lesson of that period, learned painfully, was that a network of labeled links has no fixed meaning until the labels are given a semantics: a link labeled "is a" was used in early systems for class membership, for subclass, for identity, and for typical property, and the inferences licensed by each are different.

The web gave graphs a second life. The Resource Description Framework (RDF), first standardized by the World Wide Web Consortium in 1999 and revised in 2004 and 2014, represents all data as subject, predicate, object triples, in which the subject and predicate are identified by Internationalized Resource Identifiers (IRIs) and the object is an IRI or a literal value. Berners-Lee, Hendler, and Lassila's 2001 Scientific American article proposed a web of such data in which machines could follow links between statements as humans follow links between pages. RDF 1.1 Semantics gives the triples a model-theoretic interpretation, so that entailment between sets of triples is defined.

Property graphs developed in parallel from the database community, where the need was for fast traversal rather than formal semantics. Angles and Gutierrez's 2008 survey traces the lineage. In a property graph, nodes and edges both carry key-value properties and edges have a type; the model is informal but convenient, and query languages such as Cypher (Francis and colleagues, 2018) made it widely used. In 2024 the International Organization for Standardization published GQL (ISO/IEC 39075), the first standard query language for property graphs.

The phrase "knowledge graph" was popularized by Google's 2012 announcement, after which it was adopted for structures that had previously been called semantic networks, linked data, or graph databases. In biomedicine, integrative graphs such as Hetionet (Himmelstein and colleagues, 2017), which connected genes, compounds, diseases, anatomy, and other entity types from dozens of sources, showed both the power of the approach for hypothesis generation and the difficulty of knowing, for any given edge, how much to trust it.

Philosophical or technical account

The two models

In the RDF model, everything is a triple. A node is an IRI (or a blank node, a placeholder without a global identifier, or a literal such as a number or string). An edge is a triple whose predicate is itself an IRI, which means that the relation type is a first-class entity about which further triples can be written. There are no properties on edges; if one wants to say when or by whom an edge was asserted, one must either reify the edge (make a node that stands for the statement) or use a named graph (a set of triples with its own IRI, about which triples can be written). RDF 1.2, in development at the time of writing, adds triple terms so that a triple can be the subject or object of another triple directly; its status should be checked at source check.

In the property graph model, nodes and edges are both objects with an internal identity, a type, and a map of properties. "Agent A binds Target P, asserted by source S1 on 2026-10-10" is one edge with two properties. The convenience is obvious. The cost is that the properties are not themselves nodes: source S1 is a string on the edge, not an entity in the graph, unless the modeler makes it one. Nothing in the model forces a decision either way.

The two models can represent the same information. They differ in what they make easy, what they make explicit, and what formal guarantees they carry. RDF has a specified semantics and a standard query language (Semantic web standards covers SPARQL); property graphs now have a standard query language and are beginning to acquire a formal semantics. The choice should follow the use: formal entailment and web-scale identity favor RDF; traversal performance and edge metadata favor property graphs.

Identifiers

An identifier does three jobs. It distinguishes one entity from every other within a system. It allows statements made in different places to be joined. And, if it is an IRI, it allows the entity to be looked up. A knowledge graph that reuses identifiers from an external authority (a terminology, a chemical database, a protein registry) inherits that authority's identity decisions: what the authority counts as one thing, the graph counts as one thing. This is a benefit when the authority's decisions fit the use and a hazard when they do not. It is also why this publication never guesses an external identifier: a wrong identifier does not merely mislabel a node, it silently merges it with a different entity.

Provenance

Every edge was asserted. The assertion had a source (a publication, a database release, a human curator, a model), a time, a method, and a degree of confidence that is meaningful only relative to the method. Provenance is the record of these. The W3C PROV data model (PROV-DM) and its ontology (PROV-O) give a standard vocabulary: an entity was generated by an activity, associated with an agent, and derived from other entities. A knowledge graph that records provenance for each edge can answer "why does the graph say this?" One that does not can only answer "it does." Provenance is also what makes correction possible: if source S is retracted, the edges derived from S can be found and reviewed. Without it, the graph continues to assert what its source has withdrawn.

Completeness

A graph is complete with respect to a question if every edge relevant to the question is present. No graph of any interest is complete in this sense, and under the open world assumption the graph does not claim to be. A query that returns no path from A to B has found no path in the graph; it has not found that there is no path. Users forget this constantly, because the graph presents itself as a map, and a map's blank space is read as empty territory. The review gate for this monograph is the reminder: a graph does not guarantee truth, because its edges are assertions, and it does not guarantee completeness, because its silences are silences.

Biomedical relevance

Biomedical knowledge is distributed across sources built for different purposes with different identity decisions. A gene database identifies genes; a protein database identifies proteins; a chemical database identifies compounds; a disease terminology identifies diagnoses; an anatomy ontology identifies structures. Joining them requires an identifier for each entity and typed, directed edges between entities of different kinds, which is exactly what a knowledge graph provides. Hetionet is a documented case: its authors integrated 29 public resources into a graph of 11 entity types and 24 edge types and used path patterns to prioritize compounds for repurposing. The paper is also candid that the predictions depended on the quality of the source edges and that some edge types contributed little.

Two further lessons matter. First, the value of a biomedical knowledge graph lies in the joins, and the joins depend on identifier mappings whose correctness is itself a knowledge claim with its own provenance. A graph is as trustworthy as its weakest mapping. Second, biomedical edges are context-dependent in ways a plain triple cannot express. "Target P is expressed in cell type Q" may be true in one tissue, at one developmental stage, under one measurement method, and false otherwise. An edge that records the claim without the context has recorded something true only sometimes, and a traversal through it carries that qualification silently. Making context explicit (as a node, a named graph, or an edge property) is the difference between a graph that can be checked and one that can only be believed.

Nuclear medicine relevance

Everything in this section is a synthetic example. The agents, targets, cell types, and sources are invented for teaching. The IRIs use the reserved documentation domain example.org and do not resolve. Identifiers of the form NX-T-nnnn are local and correspond to no product, terminology code, or registry entry.

A department wants to answer one question from recorded knowledge rather than memory: for a given therapeutic agent, which diagnostic agent targets the same molecule, and where should uptake of that diagnostic agent be expected on the basis of target expression? Here is the smallest graph that can answer it, written first as RDF triples in Turtle syntax.

@prefix nxt:  <http://example.org/nx-teaching/> .
@prefix rdfs: <http://www.w3.org/2000/01/rdf-schema#> .
@prefix prov: <http://www.w3.org/ns/prov#> .

nxt:NX-T-0001  rdfs:label         "Agent A (synthetic)" ;
               nxt:labeledWith    nxt:fluorine-18 ;
               nxt:hasLigand      nxt:ligand-L1 .

nxt:NX-T-0002  rdfs:label         "Agent B (synthetic)" ;
               nxt:labeledWith    nxt:lutetium-177 ;
               nxt:hasLigand      nxt:ligand-L2 .

nxt:ligand-L1  nxt:binds          nxt:target-P .
nxt:ligand-L2  nxt:binds          nxt:target-P .

nxt:target-P   nxt:expressedIn    nxt:cell-type-Q .
nxt:cell-type-Q nxt:foundIn       nxt:tissue-R .

# Provenance for one edge, using a named graph in TriG style
nxt:g-0007 { nxt:target-P nxt:expressedIn nxt:cell-type-Q . }
nxt:g-0007 prov:wasDerivedFrom nxt:source-S1 ;
           prov:generatedAtTime "2026-10-10"^^<http://www.w3.org/2001/XMLSchema#date> ;
           prov:wasAttributedTo nxt:curator-C1 .

Read aloud, the triples are ordinary sentences. Agent A is labeled with fluorine-18 and has ligand L1. Ligand L1 binds target P. Target P is expressed in cell type Q, found in tissue R. The last block says that the expression claim came from source S1 and was recorded by curator C1 on a stated date.

The same graph as a figure follows.

Synthetic teaching knowledge graph for two radiopharmaceuticals Seven boxes connected by labeled arrows. Agent A and Agent B each point to a ligand (has ligand) and to a radionuclide (labeled with). Both ligands point to Target P (binds). Target P points to Cell type Q (expressed in), which points to Tissue R (found in). A dashed box labeled Source S1 points to the expressed-in arrow with the label provenance. Agent ANX-T-0001 Agent BNX-T-0002 Fluorine-18 Lutetium-177 Ligand L1 Target P Cell type Q Ligand L2 Source S1curator C1, 2026-10-10 labeled with labeled with has ligand has ligand binds binds expressed in provenance of "expressed in" found in: Tissue R (shown as a return path to keep the figure compact)
Figure 1. A synthetic teaching knowledge graph. Boxes are nodes (entities with local identifiers). Solid arrows are directed, typed edges. The dashed box and arrow show provenance attached to one edge: the claim that Target P is expressed in cell type Q was derived from Source S1 by curator C1 on a stated date. Tissue R is drawn as a dotted return path for compactness; in the graph it is an ordinary node reached from Cell type Q by a "found in" edge. Synthetic example: no real product, target, or source is represented. Readable without color.

Now ask the department's question as a traversal. Start at NX-T-0002. Follow has ligand to ligand L2. Follow binds to target P. Follow binds backwards (which a query language permits when asked) to every ligand that binds target P: L1 and L2. Follow has ligand backwards from L1 to NX-T-0001. The diagnostic agent sharing the target is Agent A. Then, from target P, follow expressed in to cell type Q and found in to tissue R. On the basis of this graph, uptake of Agent A is to be expected in tissue R.

Now ask what the answer is worth. It is worth exactly the edges it traversed. The claim that L1 binds target P came from somewhere; the graph does not say where, so the answer inherits an unstated warrant. The claim that target P is expressed in cell type Q has provenance: source S1, curator C1, a date. A reader can go to S1 and check. The claim that cell type Q is found in tissue R has no provenance and no context; it may be true in healthy tissue and not in tumor. The traversal carried that silence through to the conclusion without comment.

And ask what the graph does not say. It does not say that Agent A does not bind any other target; it has no edge, which under the open world assumption means nothing. It does not say that target P is not expressed elsewhere. If the department reads "uptake expected in tissue R" as "uptake expected only in tissue R," it has read completeness into silence. The quality of the answer is the quality of the edges plus the honesty of the reader about the gaps.

A property graph version would put the source, curator, and date as properties on the expressed in edge, and a Cypher or GQL query would match the same path. The traversal is the same; the representation of provenance differs; and property graph systems typically assume the closed world, reporting "no match" where RDF reports "unknown."

Disagreements and limitations

What counts as a knowledge graph. The field has not settled whether the term requires an ontology, a reasoner, provenance, or integration of multiple sources. This monograph uses the broad reading. A reviewer who holds the narrow reading should note that the synthetic example has provenance on one edge, no ontology, and no reasoner, and is a knowledge graph on the broad reading only.

RDF versus property graphs. Advocates of RDF point to its formal semantics, global identifiers, and standards; advocates of property graphs point to convenience, performance, and the naturalness of edge properties. The two communities are converging (RDF 1.2 on one side, GQL and formal semantics for property graphs on the other), but a choice between them still carries consequences for what a graph can promise. This draft takes no side beyond saying that the choice should be explicit.

Open world and closed world. Under the open world assumption, the graph never says anything is false by omission, which is honest and frustrating: it cannot answer "does Agent A bind only target P?" Under the closed world assumption, absence means falsity, which is convenient and false for any incomplete graph, which is every graph. Systems that mix the two without saying so produce answers whose meaning cannot be determined.

Graph embeddings and inference. Much current work derives numerical representations from graphs and predicts missing edges. Such predictions are proposals with scores, and the caution of Epistemology applies: a score is not a warrant. A predicted edge should carry provenance saying it was predicted, by what, from what, so that it is never mistaken for an asserted one.

Scope of this draft. The historical sketch is compressed and should be checked against the primary sources. The characterization of RDF 1.2 depends on a document whose status may have changed; the source checker should record the status at the time of check. The Hetionet summary is from the authors' own paper. The synthetic example makes no claim about any real agent, target, tissue, or expression pattern.

References

  1. Euler L. Solutio problematis ad geometriam situs pertinentis. Commentarii Academiae Scientiarum Imperialis Petropolitanae. 1741;8:128-140.
  2. Quillian MR. Semantic memory. In: Minsky M, ed. Semantic Information Processing. MIT Press; 1968:227-270.
  3. Sowa JF. Conceptual Structures: Information Processing in Mind and Machine. Addison-Wesley; 1984.
  4. Berners-Lee T, Hendler J, Lassila O. The Semantic Web. Scientific American. 2001;284(5):34-43.
  5. Singhal A. Introducing the Knowledge Graph: things, not strings. Google Official Blog. 16 May 2012.
  6. Ehrlinger L, Wöß W. Towards a definition of knowledge graphs. In: SEMANTiCS 2016 Posters and Demos. CEUR Workshop Proceedings. 2016;1695.
  7. Hogan A, Blomqvist E, Cochez M, et al. Knowledge graphs. ACM Computing Surveys. 2021;54(4):Article 71.
  8. Angles R, Gutierrez C. Survey of graph database models. ACM Computing Surveys. 2008;40(1):Article 1.
  9. Francis N, Green A, Guagliardo P, et al. Cypher: an evolving query language for property graphs. In: Proceedings of the 2018 International Conference on Management of Data (SIGMOD). ACM; 2018:1433-1445.
  10. ISO/IEC 39075:2024. Information technology. Database languages. GQL. International Organization for Standardization; 2024.
  11. W3C. RDF 1.1 Concepts and Abstract Syntax. W3C Recommendation, 25 February 2014. https://www.w3.org/TR/rdf11-concepts/ (access checked 10 October 2026).
  12. W3C. RDF 1.1 Semantics. W3C Recommendation, 25 February 2014. https://www.w3.org/TR/rdf11-mt/ (access checked 10 October 2026).
  13. W3C. RDF 1.1 Turtle: Terse RDF Triple Language. W3C Recommendation, 25 February 2014. https://www.w3.org/TR/turtle/ (access checked 10 October 2026).
  14. W3C. SPARQL 1.1 Query Language. W3C Recommendation, 21 March 2013. https://www.w3.org/TR/sparql11-query/ (access checked 10 October 2026).
  15. W3C. PROV-DM: The PROV Data Model. W3C Recommendation, 30 April 2013. https://www.w3.org/TR/prov-dm/ (access checked 10 October 2026).
  16. W3C. PROV-O: The PROV Ontology. W3C Recommendation, 30 April 2013. https://www.w3.org/TR/prov-o/ (access checked 10 October 2026).
  17. W3C. RDF 1.2 Concepts and Abstract Syntax. https://www.w3.org/TR/rdf12-concepts/ (access checked 10 October 2026; check the document status at the time of source check).
  18. Himmelstein DS, Lizee A, Khankhanian P, et al. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. eLife. 2017;6:e26726.
Review gate for this monograph

Graphs do not guarantee truth or completeness.

A full draft exists. It has not been source-checked or reviewed by a named domain expert.

Source list as recorded in the manuscript metadata (18)
  1. Euler L. Solutio problematis ad geometriam situs pertinentis. Commentarii Academiae Scientiarum Imperialis Petropolitanae. 1741;8:128-140.
  2. Quillian MR. Semantic memory. In: Minsky M, ed. Semantic Information Processing. MIT Press; 1968:227-270.
  3. Sowa JF. Conceptual Structures: Information Processing in Mind and Machine. Addison-Wesley; 1984.
  4. Berners-Lee T, Hendler J, Lassila O. The Semantic Web. Scientific American. 2001;284(5):34-43.
  5. Singhal A. Introducing the Knowledge Graph: things, not strings. Google Official Blog. 16 May 2012.
  6. Ehrlinger L, Wöß W. Towards a definition of knowledge graphs. In: SEMANTiCS 2016 Posters and Demos. CEUR Workshop Proceedings. 2016;1695.
  7. Hogan A, Blomqvist E, Cochez M, et al. Knowledge graphs. ACM Computing Surveys. 2021;54(4):Article 71.
  8. Angles R, Gutierrez C. Survey of graph database models. ACM Computing Surveys. 2008;40(1):Article 1.
  9. Francis N, Green A, Guagliardo P, et al. Cypher: an evolving query language for property graphs. In: Proceedings of the 2018 International Conference on Management of Data (SIGMOD). ACM; 2018:1433-1445.
  10. ISO/IEC 39075:2024. Information technology. Database languages. GQL. International Organization for Standardization; 2024.
  11. W3C. RDF 1.1 Concepts and Abstract Syntax. W3C Recommendation, 25 February 2014. https://www.w3.org/TR/rdf11-concepts/ (access checked 10 October 2026).
  12. W3C. RDF 1.1 Semantics. W3C Recommendation, 25 February 2014. https://www.w3.org/TR/rdf11-mt/ (access checked 10 October 2026).
  13. W3C. RDF 1.1 Turtle: Terse RDF Triple Language. W3C Recommendation, 25 February 2014. https://www.w3.org/TR/turtle/ (access checked 10 October 2026).
  14. W3C. SPARQL 1.1 Query Language. W3C Recommendation, 21 March 2013. https://www.w3.org/TR/sparql11-query/ (access checked 10 October 2026).
  15. W3C. PROV-DM: The PROV Data Model. W3C Recommendation, 30 April 2013. https://www.w3.org/TR/prov-dm/ (access checked 10 October 2026).
  16. W3C. PROV-O: The PROV Ontology. W3C Recommendation, 30 April 2013. https://www.w3.org/TR/prov-o/ (access checked 10 October 2026).
  17. W3C. RDF 1.2 Concepts and Abstract Syntax. https://www.w3.org/TR/rdf12-concepts/ (access checked 10 October 2026; check the document status at the time of source check).
  18. Himmelstein DS, Lizee A, Khankhanian P, et al. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. eLife. 2017;6:e26726.

These citations have not yet been verified by a named source checker. A citation existing is not the same as a citation supporting the precise claim.