Curates a local archive into a publication-ready knowledge base.
You point it at a folder of documents and at a handful of examples of what you care about. It walks the archive, extracts the text, throws out the duplicates and the boilerplate, scores what is left against your examples, and hands you the judgement calls it cannot make for you.
It never guesses a threshold and it never edits your archive. Everything it decides, it records; everything it cannot decide, it measures and asks you about.
This file is how to operate the tool. For what each stage does when you run it, and how it resumes after a stop, see docs/running-stage-by-stage.md. For the state of the project — what is built, what is only designed — see AGENTS.md. For why it works the way it does, see docs/adr/.
Vespera has no idea what you consider relevant, and it will not try to work it out. You supply a seed folder: a handful of documents that are examples of what you want kept. That is the only domain knowledge the tool takes, and everything downstream is measured against it.
Put the path in profile.yaml in your working directory, under seedFolder, beside a line saying how you chose them. The working directory is .vespera under the directory you run the command from, unless you move it as "Where things live" describes. Create it if it is not there yet.
seedFolder:
value: 'D:\archive\exemplars'
provenance: "the twelve reports I would keep without reading them again"Every key in profile.yaml takes this shape: the answer under value, and how you arrived at it under provenance. A bare seedFolder: D:\archive\exemplars is refused, and no command starts until the key is written as above. Put a Windows path in single quotes, or write it with forward slashes: inside double quotes, YAML reads a backslash as the start of an escape.
Do this first. Nothing the tool prints can ask you for it, because by the time anything is printed the first invocation has already happened — and discovering at the second invocation that exemplars were needed all along is the worst way to meet this tool.
Getting from a folder of documents to a curated archive takes five invocations. Each one stops where a value is missing that only you can supply. Every stop is deliberate: the tool ends the invocation, keeps everything it learned, and tells you the one thing to do next.
| You set | You run | It does | |
|---|---|---|---|
| 1 | — | vespera run <root> |
walks the archive, extracts the text, takes the census. Stops before deduplication, because the boilerplate floor is read off what this invocation just measured |
| 2 | boilerplateDocumentFrequencyFloor, embeddingModel |
vespera run |
deduplicates, reads your seed folder, scores every survivor against it, groups them, and writes you sixty documents to judge. Removes nothing |
| 3 | sixty answers in relevance-labels.yaml |
vespera label |
records your answers. They belong to the documents, not to the run, so they survive everything afterwards |
| 4 | relevanceScoreFloor |
vespera run |
applies your threshold. This is the first invocation that removes anything for being irrelevant |
| 5 | arrangementApproved |
vespera run |
writes the connecting text over the arrangement you approved, and leaves the finished tree of Markdown in your working directory |
You will be told which value is missing, every time. You do not need this table in front of you and you do not need to know which stage you are at — read the last line of the output and it names what to set next.
Six is what you get if you set one value per invocation. All three of invocation 2's prerequisites are settable as soon as invocation 1 finishes, so setting them together is what makes the path five rather than six or seven.
Vespera writes reports beside the database. Each one measures something; none of them chooses for you.
| The value | Read | Which reports |
|---|---|---|
seedFolder |
your own knowledge of the archive | — |
boilerplateDocumentFrequencyFloor |
how many documents share the same passages | format-mix.html |
embeddingModel |
whichever model you can serve locally | — |
relevanceScoreFloor |
what each possible cut would cost you, in documents | relevance-labelling.html, cluster-sizes.html, seed-corpus-comparison.html |
degenerateOutputConfidenceFloor |
how well the text extraction went | confidence-distribution.html |
extractionAttempt |
whether the files the converter did not answer about are worth asking again | extraction-failures.html |
arrangementApproved |
whether the groups the tool formed are worth writing over | arrangement.html |
generationModel |
whichever model you can serve locally, if you want a different one | — |
generationContextWindow |
how much your own machine can read in one go | — |
logTimestampShareFloor |
how much of each text file begins with a date or a time | format-mix.html |
Five of these are optional and none of them is one of the five stops; all five are here so that you know they exist.
degenerateOutputConfidenceFloor, left unset, removes nothing for extracting badly. logTimestampShareFloor, left unset, removes no logs: set it, as a share between 0 and 1, to leave out as a log every text file of ten lines or more in which at least that share of the lines read begin with a date or a time. generationModel and generationContextWindow, left unset, write the connecting text with the model and the reading window Vespera ships with — they are the values that already have answers, and setting one only replaces the answer it already had.
extractionAttempt, left unset, is the first attempt at extracting text from your archive. Some files may not have been read because the converter was busy, failing or too slow at the time. extraction-failures.html lists those files beside the ones it could not convert. To have the converter asked about them again, write 2 into extractionAttempt, with why in provenance, and run again. Files it already answered about are not converted again. Everything after extraction is redone, because what was read may have changed: deduplication, scoring, grouping and the connecting text. The groups are new, so you approve the arrangement again before anything is written. Nothing from the first attempt is deleted. Remove the value, or write 1, and the next run is back on the first attempt, with nothing redone. If you had approved the second attempt's arrangement, write the first one's name back into arrangementApproved before the connecting text is written again; the last line of the output names it. Next time, write 3. Do not delete rows from vespera.db to make a stage run again.
The reading window is how much of a group goes into one request. Set it larger and more of each group is read in one go; leave it alone and Vespera uses a size any machine can serve. Groups too large to fit are still written about, from the documents nearest your exemplar, and the finished page says how many of them it was written from.
Every value you set carries a provenance field. Write down how you arrived at the number. Nothing checks that you read the report first — what stands between a guess and your archive is what you record there.
Invocation 2 leaves two files beside the database:
relevance-labelling.html— how the scores are spread, the five bands they fall into, and what cutting at each band boundary would cost you in documents kept and documents lost. Open it in a browser; every document links to the original.relevance-labels.yaml— sixty documents drawn evenly across the score range, each withrelevant: nullwaiting for atrueor afalse. A document you have already answered, against the same seed folder and from the same corpus root, shows your answer instead, so a later run's file starts from what you said last time. The file names the seed folder it was written for; if you changeseedFolderbefore runningvespera label, it is refused, so run again first.
Answer them, then run vespera label. At roughly two minutes a document that is about two hours, and the sixty is a ceiling rather than a target: the same run always asks about the same documents, so you can stop and come back.
Your answers outlive the run that asked. They are keyed to the documents, so changing the embedding model later re-scores the archive and re-reads the answers you already gave, without asking you anything twice.
Everything the tool writes goes in one working directory, never inside your archive:
profile.yaml the values you set, and how you arrived at them
vespera.db the ledger: every file seen, every verdict, every measurement
vespera.lock held by the command running now, so a second one is refused
format-mix.html what kinds of file the archive holds
extraction-failures.html the files that could not be read, and why
confidence-distribution.html how well the text came out
seed-corpus-comparison.html how far your exemplars resemble the archive
relevance-labelling.html the report you read to choose the threshold
relevance-labels.yaml the sixty questions you answer
cluster-sizes.html how the survivors grouped under each exemplar
arrangement.html the groups, named and in order, for you to approve
deliverable/<run>/index.md what was written, group by group, and what produced it
deliverable/<run>/documents.csv every surviving document, with its place in the order
Set it with --db-dir=<path>, which must be written with the =, or with vespera.working-dir in configuration. Both commands take --db-dir=<path>, so if you moved the working directory, name it on vespera label as well as on vespera run. A command given a --db-dir other than the directory it actually opened refuses and records nothing.
One command at a time uses a working directory. A second command started on it while the first is still running, from a second terminal or from an IDE, refuses at once with one line naming the command that holds it, its process id and when it started, and exits 1; the first carries on. Wait for the first to finish, or stop it, and run the second again. vespera.lock stays in the working directory after every command and is harmless: a command that crashed or was killed releases it as its process ends. Do not delete it while a command is running, because that lets a second one start. If a step fails saying the database file is held by another process, something other than Vespera, such as a database browser, has vespera.db open; close it and run the same command again.
While a command is running, two more files sit beside vespera.db: vespera.db-wal, SQLite's write-ahead log, which holds the most recent changes, and vespera.db-shm, its index. When the command ends they are folded back into vespera.db and deleted, unless something else, such as a database browser, still has the database open. To copy the working directory, copy it after the command has ended and vespera.db-wal is gone. If vespera.db-wal is still there, or you have to copy while a command is running, copy all three files together, or the copy is missing the latest changes. Keep the working directory on a disk attached to the machine that runs Vespera, not on a network share: the database relies on shared memory that only works on a local disk.
The database also keeps its temporary files in the working directory, not in the system's temporary folder. It writes them while a step has it sort a great many rows and removes them itself; on Windows they show there under names beginning etilqs_, and they are no part of a copy. They grow with the archive and can run to gigabytes on a large one, so the disk that holds the working directory needs that much free while a command runs.
vespera run <root> [--db-dir=<path>] walk a corpus and take it as far as the next missing value
vespera label [file] [--db-dir=<path>] record the answers you wrote into the label file
To have vespera run write an invocation account (a file with no document's name or words in it, which a hosted model may read), set vespera.account-dir to a folder outside every working directory, in application-local.yaml or with --vespera.account-dir=<path>; left unset, or set inside a working directory, no account is written.
vespera run takes the archive root as its argument, falling back to vespera.corpus-root in configuration. Given neither, it refuses rather than guessing — a census of the wrong tree reports success.
You need Java 26 and a Docker daemon. Run every command below from the root of this repository.
Build it. This builds the jar Vespera runs from:
./mvnw package
Start the sidecars, once, before the first invocation. Vespera needs two services running beside it:
- Ollama, which serves the models;
- docling-serve, the document converter.
Nothing in Vespera starts them or stops them. That holds for the jar, for a run from your IDE and for ./mvnw spring-boot:run alike. If your .env sets SPRING_DOCKER_COMPOSE_ENABLED, that line no longer does anything, and you can take it out. You start them yourself from compose.yaml:
docker compose -p vespera up -d --build
-p vespera names the Compose project vespera, whatever your checkout's directory is called. Give it on every docker compose command here, so that each one finds the same containers. Without it, Docker Compose names the project after the directory, and a second checkout starts a second set that fights the first for the same ports.
If your machine has an NVIDIA graphics card that Docker can use, let Ollama and the document converter run on it. Ollama runs its models several times faster there than on the processor, and the converter got through a sample of the archive two to two and a half times faster. Docker Desktop on Windows can use one through WSL 2; on Linux, Docker needs NVIDIA's Container Toolkit. Start the sidecars with a second file, compose.gpu.yaml, named after the first:
docker compose -p vespera -f compose.yaml -f compose.gpu.yaml up -d --build
Name both files every time you run up, this time and every time after. An up without compose.gpu.yaml replaces Ollama with one that runs on the processor. The models you gave it are kept, because they live in a volume of their own rather than in the container. It replaces the document converter with the processor build too. stop, exec and down need only -p vespera. Without such a card, leave compose.gpu.yaml out: with it, up stops with an error, and neither Ollama nor the document converter starts. Once a model has answered, docker compose -p vespera exec ollama ollama ps shows under PROCESSOR 100% GPU when the model is wholly on the card, a split such as 30%/70% CPU/GPU when only part of it fits, and 100% CPU when Ollama is not using the card.
With compose.gpu.yaml, the document converter is built from the same Containerfile on a base that can use the card, under a name of its own. Its output differs from the processor build's in a few places, and Vespera records beside every conversion the name of the image that made it. That name is the one you give Vespera, so you have to tell it which image you started. Set VESPERA_DOCLING_IMAGE in the shell you run vespera from, before you run it. In PowerShell:
$env:VESPERA_DOCLING_IMAGE = 'vespera/docling-serve-cu128-libreoffice:v1.32.0-docling-parse-7.17.0-r2'
In a POSIX shell:
export VESPERA_DOCLING_IMAGE='vespera/docling-serve-cu128-libreoffice:v1.32.0-docling-parse-7.17.0-r2'
If you run Vespera from the IDE instead, add this line to the file .env at the root of this repository, creating it if it is not there. The Local SpringApp and Local Vespera Label run configurations in .run/ read it, and java -jar does not:
VESPERA_DOCLING_IMAGE=vespera/docling-serve-cu128-libreoffice:v1.32.0-docling-parse-7.17.0-r2
The converter reports which image it runs, and Vespera checks that against the name you gave it before it converts anything. If you forget, or if you go back to compose.yaml alone and leave the variable set, the two do not match. The command then stops before converting a single file, and its last line names both images: the one the converter runs and the one Vespera was told. Set the variable to the image you meant to run, or start the sidecars again with the files that build the other one, and run the same command again. If you go back to compose.yaml alone, unset the variable and take the line out of .env again.
The first run after you switch to the card converts every file again, once, because conversions recorded under the processor build's name are not reused under the card's. On the card that is still quicker than finishing on the processor. Every step after the conversion runs again too. The answers you already wrote into the label file are kept, but the label file is written anew, and you are asked to approve the arrangement again. If you switch back, the processor build finds its own conversions where it left them.
To check that the converter is using the card, run this. It prints True when it is:
docker compose -p vespera exec docling-serve python -c "import torch;print(torch.cuda.is_available())"
To check that Vespera recorded the card's image, run this from the directory you run vespera from, once a run has converted something, with Python on your machine. It prints what the newest conversion run in the ledger was keyed on, and the image is the part after image=. After your first run on the card, that is vespera/docling-serve-cu128-libreoffice:
python -c "import sqlite3;print(sqlite3.connect('file:.vespera/vespera.db?mode=ro',uri=True).execute('select config_consumed from run where stage=? order by rowid desc limit 1',['extraction']).fetchone())"
If you moved the working directory, put its vespera.db in place of .vespera/vespera.db.
The document converter's image is not pulled. It is built on your machine from docker/docling-serve, because it adds LibreOffice to the published image, so that .doc and .ppt files convert. --build builds it the first time and rebuilds it if its Containerfile has changed. The first build is the slow one. After that, Docker reuses what it built.
The services listen on ports 11434 and 5001, which is where Vespera looks for them. If something else on your machine already holds one of those ports, the start fails and names the port.
Leave the sidecars up for all five invocations. They can be days apart.
If a sidecar stops on its own, Docker starts it again, and it does the same after your machine restarts, as long as Docker itself starts. When the document converter goes away while documents are being converted, Vespera waits up to three minutes for it to come back and then asks again about the document it was converting, so a restart of the converter does not stop the command. A document that makes the converter go away twice, or that the converter answers with an error, is marked, skipped and listed in extraction-failures.html, and the command carries on with the rest. If five documents in a row make it go away twice, or are turned down for the converter's own reasons, Vespera does not yet know whether the documents or the converter are to blame, so it checks the converter with a small document of its own. If the converter reads that one, the five were the documents' own doing: each is marked, skipped and listed in extraction-failures.html, and the command carries on. Only if the converter cannot read Vespera's own document either does the command stop, and it says so. One of your exemplars failing that way does stop the command, and the message names the file. If the converter is not back within three minutes the command stops and says so, and you can run the same command again once it is back. A sidecar you stop with docker compose -p vespera stop stays stopped, even across a restart of your machine.
Give Ollama its models. Ollama serves only the models it has been given, and Vespera does not fetch them for you:
- The embedding model you name in
embeddingModelhas to be there before invocation 2. - The model the connecting text is written with has to be there before invocation 5. That is
qwen3:8b, unless you setgenerationModel.
docker compose -p vespera exec ollama ollama pull <embeddingModel>
docker compose -p vespera exec ollama ollama pull qwen3:8b
Pulling the embedding model again can change it, because a name can be published again with other contents. Vespera then treats it as another model: the next vespera run embeds and scores every document again, the relevance threshold removes nothing until the sample is answered again from the new label file, and the arrangement has to be approved again. Pull it again only when you mean to, and not while vespera run is computing vectors: if the model changes under that step, the command stops there with a line saying so, and the next vespera run embeds again, storing the vectors that are missing under the model as it is now.
Run it. Each vespera in this file is this command:
java -jar target/vespera-0.0.1-SNAPSHOT.jar
So vespera run <root> is java -jar target/vespera-0.0.1-SNAPSHOT.jar run <root>, and vespera label is java -jar target/vespera-0.0.1-SNAPSHOT.jar label. Unless you set it as "Where things live" describes, the working directory is .vespera under the directory you run the command from.
Stop the sidecars when you are finished:
docker compose -p vespera stop
This keeps the models you gave Ollama. docker compose -p vespera down removes the containers and keeps the models too, because they live in a volume of their own. docker compose -p vespera down -v removes that volume as well, and the models with it. After that, the same ollama pull lines as above give them back.
When compose.yaml has changed since you started the sidecars, for example after you update your checkout, docker compose -p vespera up -d --build, or the up line above that names both files if you started them with both, replaces the container of each service whose part of compose.yaml or compose.gpu.yaml changed. Ollama's models stay. docker compose -p vespera exec ollama ollama list shows which it has. Nothing in your working directory is lost.
Approving the arrangement is the last thing you are asked for. The fifth invocation writes the connecting text over each group and leaves you a tree of Markdown in your working directory: an index and a listing of every surviving document at its root, and beneath it one directory per exemplar holding one file per group, each named and placed the way you approved them.
A group's file opens with the heading written for it, carries the prose with every citation resolved into a link to that document's numbered place in the list below, and then lists the group entire — including the documents the writing never mentioned. Each entry in that list links to the original where it already sits in your archive, so a sentence you doubt is two clicks from the document behind it. Where a group was written from part of its documents, the page says so and names both numbers. Where an answer was turned down, the group keeps its place and its file says plainly that nothing was written over it, rather than leaving a gap you have to notice.
Those files are where Vespera stops. It does not turn them into a wiki, a site or a page anywhere, and it does not upload or send them (ADR-101). Your archive is untouched throughout: the tree links to your documents where they already sit and copies none of them (ADR-104). What you do with any of it is yours.