How it works

Read locally, de-identified before storage, answered with its source.

Four capabilities that normally need four vendors, and the architectural decision underneath all of them.

What is actually in it

Four jobs, normally four vendors, normally four uploads.

Searching a confidential archive is not one problem. It is four, and each one is usually solved by a different service that wants a copy of your documents. Veldrun does all four on one machine.

01 · Read it

OCR and transcription

Real archives are not tidy text. They are scanned agreements, photographed letters, spreadsheets, email, recordings and video. Veldrun reads all of those locally, so a photograph of a signed variation becomes searchable text without leaving the machine.

Handwritten prose gets a second, context-aware reading that is searchable but never quoted as the page’s words — measured at 95% of words on one transcribed page, against 59% without it. Handwritten forms and photographs taken in poor conditions remain under release validation.

02 · Take the identities out

Redaction, before anything is written down

This happens on the way into the index, not on the way out of it. Names, addresses, account numbers and dates of birth become placeholders before a passage is stored; the real values go into an encrypted map only you hold. A document that cannot be redacted confidently is set aside rather than indexed — the system fails closed. And the rule is enforced in software: text still carrying identifiers can only reach a model on your own machine, with no override.

Measured on a public, independently annotated benchmark and published in full — including the harder quasi-identifier figure, the metric definition and what it still misses. See the numbers and the method.

03 · Answer the question

Retrieval and grounded answering

Ask in plain language and get an answer with the exact indexed passage it came from, so you can check it against the original. No folder discipline to maintain, no tagging to keep up.

04 · Weigh the sources

Ranking by credibility, not just similarity

Ordinary search returns whatever most resembles your question. A signed order and an email arguing about that order look equally relevant to a similarity score. Veldrun is built to weigh which source deserves to be believed. See below for how far that has actually been proved.

On the redaction specifically, because it is the part most people assume needs an AI. It does not, and we measured that. The shipped configuration is deterministic rules plus a local encoder model, with no large language model in the redaction path at all. We tested the LLM approach on the same corpus: it scored 63.8–65.2% across six runs against the shipped configuration’s 99.3%. It was roughly thirty-five points worse, so it is not what ships.

There is a second reason, and it is structural rather than a score. An encoder cannot be instructed by the text it is reading. A language model asked to find the names in a document can be talked out of it by that document; an encoder has no instruction to follow, so a hostile file cannot persuade it to leave a name alone.

What makes it different

Not all documents deserve equal weight.

Ordinary search treats a forwarded email and a signed court order as equally credible. Veldrun doesn't.

Sources can be weighted by the question's subject, declared evidence type, recency and owner corrections. An owner can explicitly mark an older form as replaced after reviewing its differences; Veldrun does not guess that relationship from prose. Registered official records can outrank commentary, but deterministic statute and medical-reference resolution are still planned work.

Crucially, Veldrun shows you the ranking and the reasoning behind it, so you can overrule it. A judgement you can't inspect is a judgement you can't trust.

01 Recency, weighted by domain

An outdated statute is a serious problem. An outdated grammar reference usually isn't. The decay rate differs per domain rather than applying one rule everywhere.

02 Source-linked, not automatically proven

Veldrun records the exact passage a response cites and can refuse malformed citation links. Independent-origin corroboration and claim-level entailment validation are active benchmark work, not a shipping claim.

03 Opinion kept distinguishable

Commentary and opinion remain searchable and can receive different, bounded authority weights from official records. Those labels help review; they are not an infallible judgement of truth.

Architecture, not policy

Readable archive content stays on your machines.

Most services ask you to trust a privacy policy — a promise not to look. Veldrun is built so the question doesn’t arise. The current local product indexes and analyses documents on the Windows machine you control. The sync relay is designed to carry application-encrypted artifacts whose encryption key we do not hold; its production deployment and two-PC release walkthrough are still open. Account, device and operational metadata are separate and are listed below rather than hidden behind the phrase “nothing leaves your hardware.”

What our servers hold

  • Waitlist/account email and sign-in records
  • Licence, household membership and registered-device metadata
  • Bounded operational/security logs and, when sync is enabled, encrypted relay objects

What they never hold

  • Your documents, in readable form
  • Questions or answers from the current local-only archive path
  • Your decryption keys
The trade-offs, stated plainly. The current product target is Windows — Windows 10 or 11, 64-bit. Mac and Linux are planned, not abandoned, but neither exists today and neither has a date we would ask you to rely on; a narrow macOS OCR helper is development code, not a supported Mac release. If you are on one of those, join the list and say so — it is how the order gets decided. We hold no key to your content, and that is the point — but it also means key custody is a responsibility, not only a reassurance. Organisation-held key recovery is designed and not yet built: today, if every authorised device and recovery copy is lost, the archive is not recoverable by anyone, us included. Firms with legal-hold, supervision or records obligations should read that as a gap to be closed before deployment, not a feature. Because inference runs locally, Veldrun needs a reasonably capable always-on machine (roughly 16 GB of memory). Installed performance, signing and clean-PC checks remain release gates. Browser remote access through a tunnel terminates TLS at the edge; sync artifacts use a separate application-encryption design, but that relay still requires deployment and two-PC proof. We'd rather publish the limits than let you discover them later.

The choice everyone else makes you make

Keep it private and use a weaker model, or use the best model and trust a promise.

That is the whole market. Veldrun removes the choice by changing where the protection sits — not between you and a server, but before anything is written down.

Why the choice exists

Anything that reads your documents can read your documents. Server-side search, summaries and AI assistants all work that way, so a provider genuinely unable to read your content cannot offer them — which is why encrypted cloud storage arrives with a list of what no longer works.

The alternative on offer is a contract: the provider promises not to keep your files and not to train on them. That is a promise about conduct, not a property of the system. It is only as good as the company making it.

Veldrun de-identifies before it stores

Documents are read and redacted on the way into the index, not on the way out of it. Names, addresses and account numbers become placeholders before a single passage is written down; the real values go into an encrypted map only you hold. A document that cannot be redacted confidently is set aside rather than indexed.

So the searchable archive — every chunk, every embedding, every export — contains no identifiers to leak. Not because we promise not to look. Because they are not in there.

And that is what makes the model choice a performance decision instead of a privacy one. Veldrun routes work by whether the text has been de-identified yet, and the rule is enforced in the software rather than written in a policy: anything still carrying identifiers can only reach a model running on your own machine. There is no override. Once a document has been redacted, the work built on top of it — classification, insight, structured extraction — is running on text that no longer says whose it is.

Today every model Veldrun uses runs locally, and no hosted-model option is offered or enabled. The routing gate exists so that if one is ever added, the boundary is already structural rather than a setting somebody could get wrong. Your original files never move in either case.

Questions

Before you ask

How do I know it isn't making answers up?

Answers carry links to the exact indexed file and passage used, and the local Library can open the original file. That is source traceability, not proof that the generated sentence follows from the passage. Claim-support evaluators exist, but no verifier/model/holdout combination is approved for a semantic “grounded” claim. Verify consequential answers against the original and a qualified professional.

Is my data used to train models?

The current archive-analysis path runs locally and does not use your content for training. Encrypted sync artifacts are designed to reach the relay only as ciphertext. Any future optional hosted-model path would require separate, explicit disclosure and consent; it is disabled in the current product.