Athena

Everything you own, searchable.
Nothing you own, touched.

Athena reads a folder — photographs, documents, audio, video, code — and turns it into a catalogue you can filter, graph, summarise and ask questions of. It opens every file read-only and re-hashes it afterwards, so indexing cannot change what it indexed.

The demo is a real, pre-built catalogue of 7,740 files — you are not uploading anything, and nothing here reads your machine. To index your own folder, run it locally.

Files catalogued
7,740
Tags derived
556
Teams sharing them
4
Bytes modified
0

Start here

Three things worth trying, each a live link into the demo. They take about a minute between them.

  1. Ask a question and watch the graph answer it

    Type “how do finance and legal overlap”. The agent turns it into a filter and picks which of three drawings answers it — then the catalogue does the counting. It is never asked how many there are.

    Open the graph →
  2. See one catalogue through two teams’ eyes

    Studio Ops sees all 6,120 files in the main library. Marketing holds a grant on a slice of the same one and sees fewer — and neither can see the other’s annotations. Sign in as Iris, who is in both.

    Switch between them →
  3. Tell it that a tag is wrong

    Open any file and remove a tag. The count, the facets, the summary and both graphs agree immediately — and the catalogue is not edited. Your correction is a lens your team looks through, not a change to someone else’s data.

    Open the library →

How it works

It walks, read-only

Every file is opened through one guarded reader and re-hashed afterwards. Nothing is renamed, moved, rewritten or re-timestamped. The index lives in a separate SQLite database.

It extracts

Text, document properties, EXIF, palettes, keyframes, OCR, duration, dimensions — about thirty detectors, each recorded so you can see which ones ran and which were skipped.

It decides — rules first

Deterministic rules settle most files in roughly three milliseconds each, with no GPU, no key and no network. A model is asked only about the residue the rules could not settle — about 11% of a real library.

Then you ask

Filter it, draw it as a graph, summarise a selection, or hand a question to the agent. Every figure you are shown is arithmetic over the index — the model writes prose, never numbers.

What it will not do

The limits are the product. They are listed here rather than in a policy page because they are the reason to trust any of the above.

It never writes to your files

Read-only handles, verified by re-hashing after every run. 0 integrity mutations across 7,740 catalogued files.

A correction is a lens, not an edit

Removing a wrong tag hides it for your workspace and leaves the catalogue alone. Another team sharing the same library is unaffected, and can disagree with you.

The model is given very little

A summary sends already-aggregated counts. The agent sends one file’s name, folder and type. Neither sends file contents, because this catalogue holds none — and both say so on screen.

No key reaches your browser

Model calls happen server-side only; an accidental import into client code fails the build rather than shipping a secret. The test suite asserts it on every run.

Three ways to look at it

Library

Facets on every axis at once — kind, topic, author, year, what a file contains. Counts update against the selection, so a refinement never leads somewhere empty.

Open →

Graph

The same selection as a picture, three ways: files pulled together by what they share, tags joined by co-occurrence, or a pyramid stacking broad tags above the narrow ones they contain.

Open →

Agent

Five verbs, and only three of them cost anything. It proposes labels for files that have none, explains what a file probably is, and finds what else is like it.

Open →

What the agent can do

Two of these need no model at all. That is the cost argument in one table: the expensive tier is asked only where arithmetic runs out.

VerbWhat it doesNeeds a model
askTurns a question into a filter and a drawing. Never reports a count.yes
labelProposes a kind and topic for a file that has neither. You accept or ignore.yes
explainDescribes a file from its name, folder and labels — and lists what it could not tell without opening it.yes
relatedWhat else is like this, ranked by how rare the shared tags are.no
repeatsOne name filed across many folders. Not duplicates — every hash here differs.no

With a speech key set, ask and explain — and any summary — can be read aloud, caveats and all. Listening never re-runs the model: it replays the answer already in front of you, or says it cannot.

Run it on your own files

The demo is a catalogue someone else built. This is the part that builds one from a folder of yours, and it never leaves your machine.

# Python 3.12+
git clone https://github.com/s4196552/athena
cd athena
python -m venv .venv && .venv\Scripts\activate
pip install -e .

athena serve

That is the whole install. The UI opens at 127.0.0.1:8731, bound to loopback and refusing any non-loopback origin. Point it at a folder and it starts indexing.

Optional: let a model read too

Athena is complete without this — the rules tier needs no key and the library is fully searchable without one. A model adds captions for photographs and settles the files the rules were unsure about.

# entirely local, nothing leaves the machine
set ATHENA_AI_PROVIDER=ollama

# or a shared gateway, so you need no key of your own
set ATHENA_AI_PROVIDER=gateway
set ATHENA_GATEWAY_URL=https://your-gateway.up.railway.app

Or drive the hosted demo from a terminal

The same catalogue, the same agent, over a versioned HTTP API. The client shares its types with the server, so it cannot drift from it.

cd cloud && npm install
npm run cli -- login
npm run cli -- ls --doctype invoice --date 2024
npm run cli -- ask "how do finance and legal overlap"
No demo gateway configured — Athena runs fully without one.