When No One Knows What's In There

Who's the likely user

There's a particular category of people who've inherited someone else's things. A collector died, or a researcher, an archivist, simply a person with a history — and their disks, hard drives, flash drives, external storage ended up with whoever took on the responsibility of preserving them. Sometimes it's a professional archivist. Sometimes a relative. Sometimes a trusted friend.

What they all have in common: they're holding someone else's legacy and don't know what's inside.

The professional archivist — a keeper of oral-history collections, sheet music archives, other collections. Drives from people who are no longer alive arrive regularly. They know they're obligated to go through it. They don't know when or how — because manually, this could take years.

The family of a researcher or scholar — a father spent his whole life collecting material on his subject, died, left behind terabytes. The children want to preserve it, but don't understand what's in there or who might need it.

A cultural institution — a library, museum, or foundation that received a donated collection. Formally accepted. In practice — a box of drives sits on a shelf waiting for a turn that never comes.

What the product does

The starting point is zero. No description, no table of contents, no person left who remembers what was put where. Just files.

First step — quick reconnaissance. In half a minute, the program scans the entire mass and produces a first map: what file types, what languages, what time periods, whether there's any internal structure at all or it's just one big pile.

Second step — careful deduplication. Inherited archives almost always contain copies: the same files in different folders, backups of backups. The program finds exact content-based duplicates and removes the excess carefully: it keeps the copy from the more reliable source, moves the rest to a separate folder — never deletes. Before any action — a preview: look first, act second.

Third step — layer-by-layer analysis. Documents and books are recognized and annotated by a local neural network. Material with no text layer is prepared for OCR. Audio and video are transcribed into text with timecodes on every phrase, with support for old and noisy recordings.

Fourth step — identification. For audio recordings: who's performing and what, by matching the transcribed text against a database of known works.

Fifth step — assessment and report. What's unique here and found nowhere else. What's typical. What's personal and not meant for other eyes. The output: a structured document about the archive, suitable for handoff to an institution or to heirs.

Everything runs locally. Files never leave.

What category it belongs to

This is a tool for initial archival reconnaissance — what used to be done only by a person who spent months on manual review. Existing archive-management systems assume you already know what you have. Here, the situation is different: you don't know anything.

Why it beats the alternatives

Path one: by hand. A person sits down and goes through it file by file. At a volume of several terabytes, this can take years. Most archives stay untouched — not out of indifference, but because there's no resource for it.

Path two: hire a specialist. Expensive, slow, and the specialist ends up doing the same manual work anyway.

The program offers a third path: automatic reconnaissance in a few days of machine time, after which a person sees the full map and makes decisions knowingly, not blindly.

The fundamental difference from cloud services: someone else's legacy is an especially delicate thing. It may contain personal correspondence, unpublished manuscripts, restricted-access documents. Sending it to a corporate cloud is unacceptable. Everything stays in place.

What tasks it can handle

Task Result
A first map of the archive In half a minute — types, languages, periods, overall structure
Find and remove duplicates Careful deduplication: excess goes to a separate folder, not the trash
Recognize documents OCR for books and scans with no text layer
Transcribe recordings Text with timecodes on every phrase
Identify who's performing Voice identification of performers
Identify what's being performed Matching against a database of known works
Flag what's personal and private What can't be shared or published without permission
Prepare an archive passport A document for handoff to an institution or heirs

How it helps

The archivist gets, in a few days, what would otherwise have taken years: a full picture of someone else's legacy. They see what's unique and needs careful preservation, what can be digitized and opened up, what's personal and should stay closed.

A scholar's family understands, for the first time, exactly what their father spent his life collecting — and can hand it to those who need it: a specialized archive, a university, colleagues. With understanding, not guesswork.

A cultural institution can finally answer the question that's been hanging since the donation arrived: what's here, what's unique, what to do with it.


The generation of people who spent their lives collecting, recording, preserving — is passing. With them goes the knowledge of exactly what they collected. Drives age. Formats become obsolete. The window in which this can still be saved and understood isn't infinite.

If the outcome of this reconnaissance is a decision to hand the archive to a library, museum, or foundation, the receiving institution faces the same task from the other side: Donated Collections on digercules.org.


Digeritage is a dedicated, restrained space for this specific situation — powered by the same Digercules technology already proven on real archives.