The Firmament

A knowledge graph of the Bible — where every connection carries its provenance

The Firmament is the Anselm Project's living map of Scripture: 215,802 nodes joined by 1,554,656 relationships. This page is its specification — what a node means, what a link means, which connections are deterministic, which are imported, which are AI-inferred, and exactly how much trust each class has earned.

IThe Shape

What a node and a link actually are.

Under the hood the whole graph is deliberately boring: two database tables and an alias index. A node is one canonical entity — a verse, a person, a place, an episode, a theme, an object, a Hebrew or Greek lemma — with a stable identifier like verse:genesis:1:1 or person:god, one of 24 node types currently in the public graph. The alias index maps the names authors actually write to those canonical entities, so "the LORD," "YHWH," and "Yahweh" resolve to one node instead of three.

A link (an edge) is where the interesting part lives. Every edge carries a relationship type (522 in public use — appears_in, quotes, thematic_parallel, and so on), a provenance label, a source reference — a citation handle saying exactly which process created it and from what — a stored evidence excerpt, and a confidence score. Edge identity is unique on the combination of source, target, type, and source reference: the same two entities connected by two different pieces of evidence are two edges. That is deliberate — evidence records are never merged into an unaccountable blur.

"Two entities connected by two different pieces of evidence are two links, on purpose."Evidence is the unit of account
IIProvenance

Five classes of connection, and what each is worth.

Every edge is labeled with one of five provenance values. In the map itself you can filter by them — the "Evidence source" control. If you want the taxonomy in one sentence: two classes are deterministic, one is gated AI, one is hand-authored, and one is private curation that never appears in the public graph.

1
Structural deterministic · 120,918 links
The corpus skeleton: book contains pericope, pericope contains verse, domain contains book. Pure code walking the canon — no model is involved anywhere. Confidence: certain, in the way a table of contents is certain.
2
Reference deterministic + imported · 771,179 links
Evidence-keyed indexing. Most of it is registries matching the translation's own text — the stored excerpt on such an edge is the exact verse text that triggered it. The imported scholarly layers live here too: per-word Hebrew and Greek morphology from the same sources as the translation, Strong's numbering, and the gazetteer behind the atlas. Deterministic given its source data.
3
Inference AI-proposed, gate-checked · 662,568 links
Connections proposed by a model and admitted only through the gates in the next section — quoted evidence, a confidence floor, a fixed relationship vocabulary. Shown in the map's filter as "Anselm," because these are the connections Anselm itself proposes rather than records.
4
Wikilink hand-authored
Explicit links written by a person inside a note or report, the way a wiki is written. A wikilink inherits the visibility of the document it was written in, so most live in private research layers.
5
Manual private curation
A curation layer maintained by hand. Manual edges are always private — they never appear in the public graph or in the counts on this page.
IIIThe Leash

What the AI may claim — and what it may not.

The inference class is the one that deserves suspicion, so it is the one wearing the most restraints. A model proposing a connection must satisfy every gate below; failing any one of them means the proposal is rejected before it ever touches the graph.

1
No invented targets
The model chooses from a supplied list of candidate entities. It cannot mint a new person, place, or theme by fiat.
2
A fixed vocabulary
The relationship must come from a twelve-item allowlist — discusses, develops, interprets, parallels, contrasts, foreshadows, fulfills, supports, challenges, defines, occurs_at, involves. Nothing else is accepted.
3
Evidence must be a quotation
The claimed evidence has to be a literal substring of the source passage — at least eight characters of it, verified by code, not by another model. A paraphrase is a rejection.
4
A confidence floor
Proposals below 0.72 confidence are discarded. The surviving score is stored on the edge, visible to everyone.
5
A recorded rationale
Every accepted connection carries the model's stated reason, kept with the edge rather than thrown away.
6
A named author
The model, pipeline version, and a digest of the exact input are recorded in the edge's metadata. An inferred edge can always answer the question "who said so, from what?"
IVHonest Wrinkles

Where the labels are imperfect.

A provenance taxonomy earns trust by admitting its edge cases, so here are mine. First: a large share of the cross-reference parallels — quotation, verbal, and thematic links across the canon — were originally generated by the translation committee's model (grok-4.3) during the v3 run, then deterministically validated against the canon inventory and imported. They carry the "reference" label because the import path is deterministic, but their origin is AI. The raw generation records for that stage are public in the audit trail on Hugging Face, so that origin is downloadable, not just admitted.

Second: the inference class currently mixes true model output with locally authored pattern ontologies that made no model call. The discriminator is the source reference prefix — ai: for model-proposed edges, codex: for authored patterns — which is served on every edge but was, until this page, explained nowhere.

Third, on human approval: I will not pretend there is per-edge human sign-off at this scale — this is a solo project. Human review exists as targeted layers instead: reviewed interpretive ledgers for ambiguous cases, and every entity prose page passes a two-model gate — one model writes, a different model verifies against the evidence, and every reference the page cites must be on a server-side allowlist built for that entity. Pages that fail are held back.

VNothing Hidden

Check any link yourself.

None of the above asks to be believed. Open the Firmament, select any connection, and the evidence inspector shows its provenance label, its source reference, its confidence, and the full stored evidence excerpt — to anonymous visitors, with no account. The counts on this page are the public graph only: private research layers are excluded, and a link is counted only when both of its endpoints are public.

The map is not asking to be believed.

Every connection in the Firmament carries its provenance, its evidence, and its confidence — and all of it is served to anyone who asks. Pick the link you find least plausible and open its receipts.