Counterfoil Labs Get early access
Early preview · runs on your own PostgreSQL

Every AI answer, with a receipt.

Know where every AI answer came from and who can see it, and withdraw any source everywhere at once. ReceiptDB is provenance-native storage for the chunks, embeddings, memories and answers your AI produces.

Q What's the termination notice period
in the Globex contract?
A 90 days' written notice, per §14.2
of the 2025 MSA.
answer
└─ chunk "§14.2 Either party may…"
└─ part msa-2025.pdf p.11 v3
└─ source contracts://globex/msa-2025
asked_bydana@acme.com
modelagent:legal-assist@2
trustverified
licenserestricted
readersgroup:legal
piifalse
status● active · 1 source
Works with
  • PostgreSQL
  • pgvector
  • LangChain
  • LlamaIndex
  • MCP
  • Okta
  • Entra ID
  • AWS, GCP and Azure KMS

// the problem

AI copies your data into places nobody tracks.

One document becomes chunks, vectors, extracted facts, memories, summaries and answers. The moment it does, today's stacks forget where that derived data came from.

01 · access

Permissions leak through derived data

A summary of a legal-only contract is readable by anyone the vector index lets in. Access controls stop at the source.

02 · erasure

Erasure doesn't reach the copies

Delete a person's record and their data lives on in chunks, embeddings, memories and cached answers.

03 · audit

Answers can't be audited

"Which sources did the agent use?" usually has no answer. Regulators increasingly expect one.

04 · freshness

Stale and poisoned content keeps serving

An amended contract or an injected document stays in retrieval until someone remembers to rebuild the index.

// who it's for

One lineage graph. Three teams that need it.

AI platform teams

One store with lineage, not four without.

Chunks, vectors, memories, caches and answers in one place, each linked to its sources, behind the LangChain and LlamaIndex interfaces you already use.

Security

Access follows the data.

Every derived record inherits access from every input, so a summary or an answer is never readable by someone who couldn't read its sources.

Compliance and legal

Audit any answer. Erase anyone, everywhere.

A receipt for every answer, and erasure that reaches every chunk, vector, memory and answer built from a person's data, with a certificate.

// how it works

One retract. Every copy, every answer.

Records are identified by their content and linked to what they were derived from. Labels flow down that graph and fail closed: a derived record is never more trusted, or less restricted, than its least favourable input.

When the Globex MSA is amended, one call handles everything built from it, in a single transaction. All features →

$ receipt impact contracts://globex/msa-2025
impact (nothing changed yet)
chunk 41 retract
embedding 41 remove from index
summary 1 queue rebuild
answer 17 flag · 6 askers

A preview first: nothing has changed yet.

// security

Runs in your cloud. Your data never leaves it.

ReceiptDB is a layer on the PostgreSQL you already operate, not a hosted service. There's nothing to send out of your environment, and nothing new to put through a security review except the software itself.

  • ✓Your PostgreSQL, your networkRecords, vectors and the lineage graph live in your own database, with no hosted dependency. Nothing leaves your network unless you switch on an integration that calls out, such as a connector.
  • ✓Keys in your KMS EnterpriseText about each data subject is encrypted with its own key, under a master key in AWS KMS, Google Cloud KMS or Azure Key Vault. Destroying a person's key makes their stored text unreadable, even in backups.
  • ✓Identity from your IdP EnterpriseSSO over OIDC and SCIM provisioning from Okta or Entra ID. Users always search as themselves.
  • ✓A tamper-evident audit log EnterpriseEvery lifecycle change is recorded in a hash-chained log, with signed exports and certificates.
EU AI Act · Art. 12

Record-keeping for AI systems

Each answer is stored with its inputs, model, asker and time, and period reports are built from those records.

GDPR · Art. 17

Right to erasure

Erase a person from every chunk, vector, memory and answer derived from their data, with a certificate of what was removed.

FINRA · Regulatory Notice 24-09

Supervising AI output

An audit export per answer: what it said, who asked, and every source it relied on.

ReceiptDB produces the evidence; whether you meet a regulation depends on how you use it. It has not yet been independently audited.

// features

One lineage graph, many guarantees.

Permission-aware retrieval

Derived records inherit access from every input. Revoke a group's access to a source and it's revoked on everything built from it.

Erasure that reaches everything

One call erases a data subject across chunks, vectors, memories and answers, and lists the answers that had used their data.

Answer audit

One receipt per answer: question, asker, model and lineage to every source. Filter by period, user, or "every answer that used this document".

Versioned re-ingestion

Re-run ingestion as often as you like. Unchanged content is a no-op; a changed page supersedes its old version and everything derived from it.

Quarantine for suspect content

Scanners, including a prompt-injection classifier, hold flagged sources and their derivatives out of retrieval until someone reviews them.

Safe answer cache

A cached answer is served only while every source behind it is valid and readable by the caller. No invalidation code, no cross-tenant leaks.

Model-safe migration

Cross-model queries are refused. Backfill a new embedding model while the current one serves, then cut over atomically.

Agent memory with provenance

A Mem0-style memory API where updates supersede instead of contradicting, and untrusted sources can't overwrite trusted memories.

Validity windows

Content can carry start and end dates. A digest expires with its earliest-expiring input, and search skips what has expired.

// performance

Provenance without a speed tax.

Measured on a single node. Filtered searches never come back short: the index keeps scanning until it has enough results the caller may see.

2.4ms

median search time, finding 99% of the true top 10, on real embeddings

99.5%

of the right results with trust and license filters on, at 1 million chunks, in about 5 ms

7,144/s

chunks ingested per second, each with its full lineage

12lines

added to adopt it in a popular LangChain RAG app, which also fixed a stale-answer bug

// editions

Start with the core. Add what your auditors ask for.

ReceiptDB

For teams building RAG and agents

  • Receipts, label inheritance and cascading retraction
  • Permission-aware, policy-filtered vector search
  • Erasure, validity windows, quarantine and re-derivation
  • Answer audit export and safe answer cache
  • LangChain, LlamaIndex, MCP, HTTP, Python and TypeScript

ReceiptDB Enterprise

For regulated and multi-team deployments

  • SSO (OIDC) and SCIM identity sync from Okta or Entra ID
  • Connectors that sync content and permissions: Google Drive, SharePoint, Slack, Confluence (preview)
  • Crypto-shredding with keys in AWS, GCP or Azure KMS
  • Signed erasure certificates and a tamper-evident audit log
  • EU AI Act and FINRA report templates, provenance explorer
  • Multi-tenancy, scoped API keys, usage metering and quotas

// faq

Questions teams ask first.

Does it replace our vector database?

It can. ReceiptDB stores vectors in PostgreSQL with pgvector and serves search itself, with your policies enforced. It plugs into LangChain and LlamaIndex as a vector store, so adopting it is a small code change rather than a migration project.

Do we need to run a new database?

No. ReceiptDB runs on PostgreSQL with the pgvector extension, which your team may already operate. The lineage graph and audit log are ordinary tables, so your existing backups, monitoring and access controls apply.

What does "early preview" mean?

The core is built and tested, including at a million chunks, but the APIs may still change. We don't recommend production use yet unless you're working with us as a design partner.

Is it open source?

The licence will be announced before general availability.

What does it cost?

Pricing is set with design partners during the preview. Get in touch and tell us about your deployment.

// early access

Building AI on data you're accountable for?

We're looking for a small number of design partners. Tell us what you're building and where provenance hurts today.

  • Early access to ReceiptDB and Enterprise
  • A direct line to the people building it
  • A real say in the roadmap
hello@counterfoillabs.com