onlyDB: Graph-Vector Storage for Cognitive AI
Structure, provenance, and constraints — not just embeddings
onlyDB: Graph-Vector Storage for Cognitive AI
The Problem with Flat Vectors
Modern AI systems store knowledge as flat vector embeddings. You embed a document, store the vector, and search by cosine similarity. This works for retrieval — but it fails for cognition.
- No structure — A vector embedding collapses a document into a single point. The relationships between facts, the hierarchy of concepts, the provenance of claims — all lost.
- No provenance — When a vector database returns a similar document, you don't know where the information came from. Was it from a peer-reviewed paper? A Reddit comment? A hallucination?
- No constraints — Vectors don't enforce logical consistency. A vector database will happily return contradictory facts with similar embeddings.
onlyDB: The Cognitive Database Layer
onlyDB is built on HelixDB-style graph-vector storage. It is implemented in the Only Forum prototype — a Rust/Axum/Maud SSR forum with Postgres for canonical content and HelixDB for the graph/vector knowledge layer.
Dual Storage Architecture
- Postgres — canonical content storage. Posts, revisions, user accounts, and metadata. ACID guarantees for content integrity.
- HelixDB — graph/vector knowledge layer. Every post is embedded as a node in a knowledge graph with edges to related posts, authors, and topics. Vector embeddings enable semantic search.
Hybrid Search
- Lexical retrieval — traditional full-text search on Postgres (tsvector)
- Graph expansion — find related posts via graph traversal in HelixDB
- Vector similarity — find semantically similar posts via embedding cosine similarity
- Merge & rank — combine all three signals into a unified ranking
Searching for "Ricci curvature" finds not just posts containing those words, but also posts about "Ollivier curvature," "graph geometry," and "information flow bottlenecks" — because the graph knows these concepts are related.
Structure
Every fact is a node. Every relationship is a typed edge. The graph structure is explicit, queryable, and mathematically governed. When an agent retrieves information, it gets not just the fact but its full relational context.
Provenance
Every node carries provenance metadata: source, confidence, timestamp, and verification status. Facts from peer-reviewed papers have higher confidence than facts from unverified sources. Provenance is a first-class query dimension.
PIR Health Monitoring
The graph's structural health is monitored using PIR balance gap. When the balance gap spikes, it indicates the graph is becoming structurally unsound — too many unsupported claims, too many contradictory nodes, too many orphaned edges.
Implementation
The Only Forum prototype demonstrates onlyDB in production:
- Rust + Axum + Maud — server-rendered, type-safe, memory-safe
- Postgres — canonical content with ACID guarantees
- HelixDB — graph/vector knowledge layer
- Multilingual first-class support — posts in any language with automatic translation links
- Post revisions with edit history — every edit preserved with diff visualization
- Citation graph with author neighborhoods — discover related researchers
Why It Matters
Without structured memory, AI is trapped in the present moment. It can retrieve similar text, but it cannot reason about relationships, trace provenance, or maintain consistency. onlyDB gives AI the three things it needs for genuine cognition:
- Structure — relationships between facts, not just similarity
- Provenance — every fact traces to a source
- Constraints — the graph enforces logical consistency
This is the memory layer that makes durable, trustworthy AI possible.
Published by Only Institute