Skip to main content
Portrait representing Abhishek Thakur
therealshek Abhishek Thakur (therealshek)
Abhishek Thakur · Software Engineer
Open to work · India / Remote
Version control CLI

Persephone

I developed a Go VCS prototype around bounded parallel hashing, a deterministic binary index, and compressed content-addressed objects. The purr CLI currently supports staging and local commit history.

Focus
Bounded concurrency · binary formats · filesystem I/O
Built with
Go · content-addressed objects · CLI
Status
Prototype · local commands only
Ownership
Started with Chandranil Bakshi

Engineering focus

Persephone began with Chandranil Bakshi and continues in my repository. Its technical focus is the staging pipeline: parallel file I/O and hashing must produce a coherent index regardless of worker completion order. The prototype also implements the blob, tree, commit, and ref structures needed for local history.

  • A fixed worker pool separates file discovery from hashing and compression.
  • Modification-time and size checks avoid reading unchanged entries; final path sorting makes index output deterministic.
  • A versioned binary index preserves nanosecond timestamps, while compressed objects and parent links form local history.

Architecture

  1. Filesystem discovery feeds bounded workers

    The staging path walks the selected files into a buffered job channel. It starts runtime.NumCPU() * 5 workers and sizes the channel to four times that count. Reads, hashing, and blob serialization happen in those workers rather than spawning unbounded work per file.

  2. Concurrent work converges on one index

    Workers inspect cached file metadata and update a shared index map under an RWMutex. Expensive file I/O runs outside the map lock. After workers finish, the staging code sorts entries by path before serializing the index.

  3. Content objects form a local history graph

    The commit path builds a tree from index entries and stores zlib-compressed SHA-1 objects under .purr/objects. A commit names its tree and parent, then advances HEAD. The log command traverses parent links.

Engineering challenges

Parallel staging still needs bounded resource use

A buffered jobs channel limits queued paths and a fixed worker count limits active file work. A separate error consumer collects worker failures while work proceeds. This bounds active hashing, but it is not a constant-memory guarantee: the index and collected results still grow with repository size.

Read the implementation: purrcommands/add.go

Worker completion order must not change stored output

Concurrent workers finish in different orders, so the shared map is only an intermediate representation. Sorting paths before writing the index gives subsequent commit construction a stable input order. Locks protect map access without serializing the entire read-and-hash operation.

Read the implementation: purrcommands/add.go

Timestamp precision affects the stat cache

The binary index has a fixed header, big-endian fields, variable-length paths, and alignment padding. Its version 3 format preserves modification times as seconds plus nanoseconds; version 2 discarded subsecond precision. The reader handles both versions and checks truncated headers, entries, and paths.

Read the implementation: index/index.go

Content identity needs a defined byte representation

Blob hashing, tree construction, compression, and object paths form the storage format. Commits refer to tree and parent hashes rather than copying previous snapshots. Before recording a commit, the implementation compares its tree with the parent tree and rejects an unchanged snapshot.

Read the implementation: purrcommands/commit.go

Tradeoffs and current limits

The stat cache trusts modification time and size; matching metadata is an optimization, not proof that file contents are unchanged. The worker multiplier is an implementation choice, not evidence that it is optimal across disks and workloads. SHA-1 is part of this prototype's Git-style object format, not a security guarantee.

Persephone supports initialization, configuration, staging, removal, index listing, commits, and local history. Branches, merges, remotes, and diffs are unimplemented. The repository links to a separate benchmark project; I make no general speed claim here.

Code and documentation