Every AI answer, with a receipt.
Know where every AI answer came from and who can see it, and withdraw any source everywhere at once. ReceiptDB is provenance-native storage for the chunks, embeddings, memories and answers your AI produces.
- PostgreSQL
- pgvector
- LangChain
- LlamaIndex
- MCP
- Okta
- Entra ID
- AWS, GCP and Azure KMS
// the problem
AI copies your data into places nobody tracks.
One document becomes chunks, vectors, extracted facts, memories, summaries and answers. The moment it does, today's stacks forget where that derived data came from.
Permissions leak through derived data
A summary of a legal-only contract is readable by anyone the vector index lets in. Access controls stop at the source.
Erasure doesn't reach the copies
Delete a person's record and their data lives on in chunks, embeddings, memories and cached answers.
Answers can't be audited
"Which sources did the agent use?" usually has no answer. Regulators increasingly expect one.
Stale and poisoned content keeps serving
An amended contract or an injected document stays in retrieval until someone remembers to rebuild the index.
// who it's for
One lineage graph. Three teams that need it.
One store with lineage, not four without.
Chunks, vectors, memories, caches and answers in one place, each linked to its sources, behind the LangChain and LlamaIndex interfaces you already use.
Access follows the data.
Every derived record inherits access from every input, so a summary or an answer is never readable by someone who couldn't read its sources.
Audit any answer. Erase anyone, everywhere.
A receipt for every answer, and erasure that reaches every chunk, vector, memory and answer built from a person's data, with a certificate.
// how it works
One retract. Every copy, every answer.
Records are identified by their content and linked to what they were derived from. Labels flow down that graph and fail closed: a derived record is never more trusted, or less restricted, than its least favourable input.
When the Globex MSA is amended, one call handles everything built from it, in a single transaction. All features →
A preview first: nothing has changed yet.
// security
Runs in your cloud. Your data never leaves it.
ReceiptDB is a layer on the PostgreSQL you already operate, not a hosted service. There's nothing to send out of your environment, and nothing new to put through a security review except the software itself.
- ✓Your PostgreSQL, your networkRecords, vectors and the lineage graph live in your own database, with no hosted dependency. Nothing leaves your network unless you switch on an integration that calls out, such as a connector.
- ✓Keys in your KMS EnterpriseText about each data subject is encrypted with its own key, under a master key in AWS KMS, Google Cloud KMS or Azure Key Vault. Destroying a person's key makes their stored text unreadable, even in backups.
- ✓Identity from your IdP EnterpriseSSO over OIDC and SCIM provisioning from Okta or Entra ID. Users always search as themselves.
- ✓A tamper-evident audit log EnterpriseEvery lifecycle change is recorded in a hash-chained log, with signed exports and certificates.
Record-keeping for AI systems
Each answer is stored with its inputs, model, asker and time, and period reports are built from those records.
Right to erasure
Erase a person from every chunk, vector, memory and answer derived from their data, with a certificate of what was removed.
Supervising AI output
An audit export per answer: what it said, who asked, and every source it relied on.
ReceiptDB produces the evidence; whether you meet a regulation depends on how you use it. It has not yet been independently audited.
// features
One lineage graph, many guarantees.
Permission-aware retrieval
Derived records inherit access from every input. Revoke a group's access to a source and it's revoked on everything built from it.
Erasure that reaches everything
One call erases a data subject across chunks, vectors, memories and answers, and lists the answers that had used their data.
Answer audit
One receipt per answer: question, asker, model and lineage to every source. Filter by period, user, or "every answer that used this document".
Versioned re-ingestion
Re-run ingestion as often as you like. Unchanged content is a no-op; a changed page supersedes its old version and everything derived from it.
Quarantine for suspect content
Scanners, including a prompt-injection classifier, hold flagged sources and their derivatives out of retrieval until someone reviews them.
Safe answer cache
A cached answer is served only while every source behind it is valid and readable by the caller. No invalidation code, no cross-tenant leaks.
Model-safe migration
Cross-model queries are refused. Backfill a new embedding model while the current one serves, then cut over atomically.
Agent memory with provenance
A Mem0-style memory API where updates supersede instead of contradicting, and untrusted sources can't overwrite trusted memories.
Validity windows
Content can carry start and end dates. A digest expires with its earliest-expiring input, and search skips what has expired.
// performance
Provenance without a speed tax.
Measured on a single node. Filtered searches never come back short: the index keeps scanning until it has enough results the caller may see.
median search time, finding 99% of the true top 10, on real embeddings
of the right results with trust and license filters on, at 1 million chunks, in about 5 ms
chunks ingested per second, each with its full lineage
added to adopt it in a popular LangChain RAG app, which also fixed a stale-answer bug
// editions
Start with the core. Add what your auditors ask for.
ReceiptDB
For teams building RAG and agents
- Receipts, label inheritance and cascading retraction
- Permission-aware, policy-filtered vector search
- Erasure, validity windows, quarantine and re-derivation
- Answer audit export and safe answer cache
- LangChain, LlamaIndex, MCP, HTTP, Python and TypeScript
ReceiptDB Enterprise
For regulated and multi-team deployments
- SSO (OIDC) and SCIM identity sync from Okta or Entra ID
- Connectors that sync content and permissions: Google Drive, SharePoint, Slack, Confluence (preview)
- Crypto-shredding with keys in AWS, GCP or Azure KMS
- Signed erasure certificates and a tamper-evident audit log
- EU AI Act and FINRA report templates, provenance explorer
- Multi-tenancy, scoped API keys, usage metering and quotas
// faq
Questions teams ask first.
Does it replace our vector database?
It can. ReceiptDB stores vectors in PostgreSQL with pgvector and serves search itself, with your policies enforced. It plugs into LangChain and LlamaIndex as a vector store, so adopting it is a small code change rather than a migration project.
Do we need to run a new database?
No. ReceiptDB runs on PostgreSQL with the pgvector extension, which your team may already operate. The lineage graph and audit log are ordinary tables, so your existing backups, monitoring and access controls apply.
What does "early preview" mean?
The core is built and tested, including at a million chunks, but the APIs may still change. We don't recommend production use yet unless you're working with us as a design partner.
Is it open source?
The licence will be announced before general availability.
What does it cost?
Pricing is set with design partners during the preview. Get in touch and tell us about your deployment.
// early access
Building AI on data you're accountable for?
We're looking for a small number of design partners. Tell us what you're building and where provenance hurts today.
- Early access to ReceiptDB and Enterprise
- A direct line to the people building it
- A real say in the roadmap