How it works
Read locally, de-identified before storage, answered with its source.
Four capabilities that normally need four vendors, and the architectural decision underneath all of them.
What is actually in it
Four jobs, normally four vendors, normally four uploads.
Searching a confidential archive is not one problem. It is four, and each one is usually solved by a different service that wants a copy of your documents. Veldrun does all four on one machine.
OCR and transcription
Real archives are not tidy text. They are scanned agreements, photographed letters, spreadsheets, email, recordings and video. Veldrun reads all of those locally, so a photograph of a signed variation becomes searchable text without leaving the machine.
Handwritten prose gets a second, context-aware reading that is searchable but never quoted as the page’s words — measured at 95% of words on one transcribed page, against 59% without it. Handwritten forms and photographs taken in poor conditions remain under release validation.
Redaction, before anything is written down
This happens on the way into the index, not on the way out of it. Names, addresses, account numbers and dates of birth become placeholders before a passage is stored; the real values go into an encrypted map only you hold. A document that cannot be redacted confidently is set aside rather than indexed — the system fails closed. And the rule is enforced in software: text still carrying identifiers can only reach a model on your own machine, with no override.
Measured on a public, independently annotated benchmark and published in full — including the harder quasi-identifier figure, the metric definition and what it still misses. See the numbers and the method.
Retrieval and grounded answering
Ask in plain language and get an answer with the exact indexed passage it came from, so you can check it against the original. No folder discipline to maintain, no tagging to keep up.
Ranking by credibility, not just similarity
Ordinary search returns whatever most resembles your question. A signed order and an email arguing about that order look equally relevant to a similarity score. Veldrun is built to weigh which source deserves to be believed. See below for how far that has actually been proved.
There is a second reason, and it is structural rather than a score. An encoder cannot be instructed by the text it is reading. A language model asked to find the names in a document can be talked out of it by that document; an encoder has no instruction to follow, so a hostile file cannot persuade it to leave a name alone.
What makes it different
Not all documents deserve equal weight.
Ordinary search treats a forwarded email and a signed court order as equally credible. Veldrun doesn't.
Sources can be weighted by the question's subject, declared evidence type, recency and owner corrections. An owner can explicitly mark an older form as replaced after reviewing its differences; Veldrun does not guess that relationship from prose. Registered official records can outrank commentary, but deterministic statute and medical-reference resolution are still planned work.
Crucially, Veldrun shows you the ranking and the reasoning behind it, so you can overrule it. A judgement you can't inspect is a judgement you can't trust.
01 Recency, weighted by domain
An outdated statute is a serious problem. An outdated grammar reference usually isn't. The decay rate differs per domain rather than applying one rule everywhere.
02 Source-linked, not automatically proven
Veldrun records the exact passage a response cites and can refuse malformed citation links. Independent-origin corroboration and claim-level entailment validation are active benchmark work, not a shipping claim.
03 Opinion kept distinguishable
Commentary and opinion remain searchable and can receive different, bounded authority weights from official records. Those labels help review; they are not an infallible judgement of truth.
Architecture, not policy
Readable archive content stays on your machines.
Most services ask you to trust a privacy policy — a promise not to look. Veldrun is built so the question doesn’t arise. The current local product indexes and analyses documents on the Windows machine you control. The sync relay is designed to carry application-encrypted artifacts whose encryption key we do not hold; its production deployment and two-PC release walkthrough are still open. Account, device and operational metadata are separate and are listed below rather than hidden behind the phrase “nothing leaves your hardware.”
What our servers hold
- Waitlist/account email and sign-in records
- Licence, household membership and registered-device metadata
- Bounded operational/security logs and, when sync is enabled, encrypted relay objects
What they never hold
- Your documents, in readable form
- Questions or answers from the current local-only archive path
- Your decryption keys
The choice everyone else makes you make
Keep it private and use a weaker model, or use the best model and trust a promise.
That is the whole market. Veldrun removes the choice by changing where the protection sits — not between you and a server, but before anything is written down.
Why the choice exists
Anything that reads your documents can read your documents. Server-side search, summaries and AI assistants all work that way, so a provider genuinely unable to read your content cannot offer them — which is why encrypted cloud storage arrives with a list of what no longer works.
The alternative on offer is a contract: the provider promises not to keep your files and not to train on them. That is a promise about conduct, not a property of the system. It is only as good as the company making it.
Veldrun de-identifies before it stores
Documents are read and redacted on the way into the index, not on the way out of it. Names, addresses and account numbers become placeholders before a single passage is written down; the real values go into an encrypted map only you hold. A document that cannot be redacted confidently is set aside rather than indexed.
So the searchable archive — every chunk, every embedding, every export — contains no identifiers to leak. Not because we promise not to look. Because they are not in there.
Today every model Veldrun uses runs locally, and no hosted-model option is offered or enabled. The routing gate exists so that if one is ever added, the boundary is already structural rather than a setting somebody could get wrong. Your original files never move in either case.
Questions
Before you ask
How do I know it isn't making answers up?
Answers carry links to the exact indexed file and passage used, and the local Library can open the original file. That is source traceability, not proof that the generated sentence follows from the passage. Claim-support evaluators exist, but no verifier/model/holdout combination is approved for a semantic “grounded” claim. Verify consequential answers against the original and a qualified professional.
Is my data used to train models?
The current archive-analysis path runs locally and does not use your content for training. Encrypted sync artifacts are designed to reach the relay only as ciphertext. Any future optional hosted-model path would require separate, explicit disclosure and consent; it is disabled in the current product.