Layered graphs, explained
A knowledge graph stores facts: Alice works at Acme. Agents need more than that. They need to say how sure am I, where did I read it, who told me, what do I conclude from it, and later I was wrong, it was Globex. Those are statements about statements. A layered graph is a graph where that is not a special feature, because every fact has an id and can be talked about like any other node.
The problem: a triple cannot point at itself
In a plain triple store, alice worksAt acme is a row with no name. To add a confidence you have three bad options:
- Reification (RDF 1.0): invent a node
_:r, then add_:r subject alice,_:r predicate worksAt,_:r object acme, and only then_:r confidence 0.8. Four extra rows, and the original fact is a separate row that nothing links to. - Quads: put the fact in a “context” graph and describe the graph. The confidence now belongs to a group of facts, not to this one.
- Property graphs: an edge can have properties, but an edge cannot be the target of another edge, so a belief cannot be “about” a relationship.
The idea: every statement has an id
In tiramemsu, a fact is a row (eid, s, p, o) plus its lifetime. The eid is its address. Since an id can be used anywhere a node can, another statement can have it as its subject or object:
e1. Layer 3 holds a belief whose object is e1. Each dashed arrow is one ordinary row.There is no second data structure. “Layer” only describes how statements point at each other, and the depth is not limited. A confidence in a source's reliability is a statement about e3, which is a statement about e1. A statement may not use its own id, but longer cycles are allowed.
Asking layered questions
In SPARQL, RDF 1.2 annotation syntax binds straight to the id. These queries are from the test suite:
# the fact's confidence and source, in one pattern
SELECT ?c ?s WHERE {
v:alice v:worksAt v:acme {| v:confidence ?c ; v:source ?s |}
}
# bind the id, then use it like any node
SELECT ?p ?c WHERE {
?p v:worksAt v:acme ~ ?r
OPTIONAL { ?r v:confidence ?c }
}
# a layer on a layer: the method behind the confidence
SELECT ?c ?m WHERE {
v:alice v:worksAt v:acme {| v:confidence ?c ~ ?r2 {| v:method ?m |} |}
}
In Cypher, the same id has two faces. It is a relationship, and it is also a node with the label :Statement, so a relationship can be the target of a pattern:
MATCH (a)-[r:worksAt]->(c), (b:Belief)-[:supportedBy]->(r)
RETURN a, c, r.confidence, r.txAdded, b
A literal-valued statement about r reads as a property (r.confidence), and a node- or statement-valued one reads as a relationship. Both languages see the same rows, and a differential test suite checks that equivalent queries return equal results.
Walking across layers
To find what a belief rests on, follow its statements down to the entities they mention. The virtual hops sys:subject and sys:object step from a statement to its parts, so a path can cross layers:
# SPARQL: everything behind belief9, through any number of layers
SELECT ?x WHERE { v:belief9 v:supportedBy/(sys:subject|sys:object)+ ?x }
# Cypher
MATCH p = (b:Belief)-[:SUPPORTED_BY|`sys:subject`*2]->(x) RETURN x
Layers have a life of their own
Because a layer statement is an ordinary row, it has its own transaction time and valid time. So the structure itself is versioned:
- Retracting a fact retracts its layers. The cascade follows subject and object positions, and every cascaded row gets the same retraction time, so
as_ofjust before it still shows the whole structure, including what a belief relied on. - Correcting a fact replays its layers.
supersederetracts the fact and everything hanging off it, then inserts them again on the new fact, in one transaction. The history shows one correction, and the confidence and source survive. - Changing a value is different. For a single-valued predicate like an age,
30becoming31is a new fact: the source of “30” does not apply to “31”, so its layers are retracted and not replayed. - Corroboration is a layer.
confirm(eid)adds(eid sys:confirmedBy tx), and the transaction carries the source and author. “Seen five times from three sources” is a query.
Named graphs are one more layer
Grouping by session, source document or agent is the same trick. A graph is a node, and membership is a statement (e sys:inGraph g) about the fact's id. A fact can be in many graphs or none, and it keeps a single id whatever the number. GRAPH, FROM and WITH in SPARQL work on it, and metadata about the graph is ordinary triples about the graph node.
What it costs
An id is not free. On a 750 000-triple test where 10 % of edges are annotated, the id-per-statement layout was about 50 % larger than a clustered (s,p,o) table with RDF 1.2 reifiers, and 2× slower on a set-semantics 2-hop join, because duplicates have to be removed. It was faster on annotation-heavy reads and writes (1.3× to 2.5×). The engine now skips duplicate removal for predicates that have never held two ids for the same triple. The full table is in the comparison article.
If you never annotate, never correct and never look at the past, a plain triple store is smaller and simpler. Layers pay off when provenance, belief and revision are the point.
Try it
The repository has a runnable example, cargo run -p tiramemsu --example quickstart, which asserts a fact with a confidence and a source, corrects it, and queries before and after in SPARQL and Cypher.