Persephone
I developed a Go VCS prototype around bounded parallel hashing, a deterministic binary index, and compressed content-addressed objects. The purr CLI currently supports staging and local commit history.
- Focus
- Bounded concurrency · binary formats · filesystem I/O
- Built with
- Go · content-addressed objects · CLI
- Status
- Prototype · local commands only
- Ownership
- Started with Chandranil Bakshi
Engineering focus
Persephone began with Chandranil Bakshi and continues in my repository. Its technical focus is the staging pipeline: parallel file I/O and hashing must produce a coherent index regardless of worker completion order. The prototype also implements the blob, tree, commit, and ref structures needed for local history.
- A fixed worker pool separates file discovery from hashing and compression.
- Modification-time and size checks avoid reading unchanged entries; final path sorting makes index output deterministic.
- A versioned binary index preserves nanosecond timestamps, while compressed objects and parent links form local history.
Architecture
-
Filesystem discovery feeds bounded workers
The staging path walks the selected files into a buffered job channel. It starts
runtime.NumCPU() * 5workers and sizes the channel to four times that count. Reads, hashing, and blob serialization happen in those workers rather than spawning unbounded work per file. -
Concurrent work converges on one index
Workers inspect cached file metadata and update a shared index map under an RWMutex. Expensive file I/O runs outside the map lock. After workers finish, the staging code sorts entries by path before serializing the index.
-
Content objects form a local history graph
The commit path builds a tree from index entries and stores zlib-compressed SHA-1 objects under
.purr/objects. A commit names its tree and parent, then advances HEAD. The log command traverses parent links.
Engineering challenges
Parallel staging still needs bounded resource use
A buffered jobs channel limits queued paths and a fixed worker count limits active file work. A separate error consumer collects worker failures while work proceeds. This bounds active hashing, but it is not a constant-memory guarantee: the index and collected results still grow with repository size.
Worker completion order must not change stored output
Concurrent workers finish in different orders, so the shared map is only an intermediate representation. Sorting paths before writing the index gives subsequent commit construction a stable input order. Locks protect map access without serializing the entire read-and-hash operation.
Timestamp precision affects the stat cache
The binary index has a fixed header, big-endian fields, variable-length paths, and alignment padding. Its version 3 format preserves modification times as seconds plus nanoseconds; version 2 discarded subsecond precision. The reader handles both versions and checks truncated headers, entries, and paths.
Content identity needs a defined byte representation
Blob hashing, tree construction, compression, and object paths form the storage format. Commits refer to tree and parent hashes rather than copying previous snapshots. Before recording a commit, the implementation compares its tree with the parent tree and rejects an unchanged snapshot.
Tradeoffs and current limits
The stat cache trusts modification time and size; matching metadata is an optimization, not proof that file contents are unchanged. The worker multiplier is an implementation choice, not evidence that it is optimal across disks and workloads. SHA-1 is part of this prototype's Git-style object format, not a security guarantee.
Persephone supports initialization, configuration, staging, removal, index listing, commits, and local history. Branches, merges, remotes, and diffs are unimplemented. The repository links to a separate benchmark project; I make no general speed claim here.
Code and documentation
Read the code and the documents behind this explanation.