🗃️

Woodrat

GoKafkaS3PostgreSQL

Woodrat is the file cataloging and archival service for my self-hosted platform — one place to record what a file is, where its bytes live, and which downstream services should get a copy.

A file can be catalogued by metadata alone (name, MIME type, checksum, origin) before the bytes ever arrive. When I ask, Woodrat archives the contents to S3-compatible object storage and publishes a Kafka event so named targets — Owl’s paper, book, and stack libraries — can pull it in. It also reconciles: point it at a bucket that was filled by a raw sync and it backfills catalogue records for everything already there.

Woodrat is the file cataloging, archiving, and distribution service for the joel.holmes.haus platform. It provides a central location to store file metadata and blobs, and publishes lifecycle events when files are cataloged, archived, or pushed.

Features

  • File Cataloging — Record file metadata (name, MIME type, checksum, origin) without requiring the bytes upfront.
  • Archiving — Store file contents in S3-compatible object storage via a gocloud.dev/blob URL.
  • Distribution — Publish FilePushedEvent to Kafka for named downstream targets (owl-papers, owl-stacks, owl-books).
  • Folders — Organize files into named folders; bulk archive or push an entire folder.
  • Reconciliation — Backfill File records for objects that already exist in the bucket (e.g. a raw mc mirror/aws s3 sync) but were never cataloged through the API.
  • CLI — cmd/woodrat is a Cobra-based client (catalog, archive, push, folder, import, reconcile) that talks to the API over ConnectRPC; import chains catalog + archive + push in one step.

Documentation