Architecture

System Context

C4Context title System Context โ€” Pangolin within the joelholmes.dev platform Person(joel, "Joel", "Runs ad-hoc analytics queries; relies on public JSON exports on his static sites") Boundary(platform, "Self-Hosted Platform") { System(pangolin, "Pangolin", "Postgres-to-Parquet snapshot and analytics export service") SystemDb(weevildb, "weevil DB", "Books catalog") SystemDb(owldb, "owl DB", "Books and papers") SystemDb(magpiedb, "magpie DB", "Resources and labels") SystemDb(shrikedb, "shrike DB", "Search index records") SystemDb(greysealdb, "greyseal DB", "Conversations and messages") SystemDb(lynxdb, "lynx DB", "RSS feeds and websites") SystemDb(woodratdb, "woodrat DB", "Cataloged files") SystemDb(narwhaldb, "narwhal DB", "Products and design docs") SystemDb(rabbitdb, "rabbit DB", "Projects, sprints, tickets") System(weevil, "weevil API", "Book service โ€” own ConnectRPC API") System(lynx, "lynx API", "Feed/website service โ€” own ConnectRPC API") System(owl, "owl API", "Papers/books service โ€” own ConnectRPC API") } SystemDb(minio, "MinIO", "Private bucket โ€” date-partitioned Parquet snapshots (same host as the Postgres instances)") SystemDb(publics3, "Public AWS S3", "Public bucket โ€” per-entity JSON for static sites / browser widgets") Rel(pangolin, weevildb, "SELECT (pgx) hourly") Rel(pangolin, owldb, "SELECT (pgx) hourly") Rel(pangolin, magpiedb, "SELECT (pgx) hourly") Rel(pangolin, shrikedb, "SELECT (pgx) hourly") Rel(pangolin, greysealdb, "SELECT (pgx) hourly") Rel(pangolin, lynxdb, "SELECT (pgx) hourly") Rel(pangolin, woodratdb, "SELECT (pgx) hourly") Rel(pangolin, narwhaldb, "SELECT (pgx) hourly") Rel(pangolin, rabbitdb, "SELECT (pgx) hourly") Rel(pangolin, minio, "Writes Parquet snapshots; DuckDB httpfs reads back for queries") Rel(pangolin, weevil, "GET current books (ConnectRPC) every 6h") Rel(pangolin, lynx, "GET current websites (ConnectRPC) every 6h") Rel(pangolin, owl, "GET current papers (ConnectRPC) every 6h") Rel(pangolin, publics3, "Writes public per-entity JSON") Rel(joel, pangolin, "POST /api/v1/query (DuckDB SQL)") Rel(joel, publics3, "Static sites / widgets fetch JSON directly")

Container Diagram

C4Container title Pangolin โ€” Internal Containers Boundary(pangolin, "Pangolin") { Container(api, "cmd/api", "Go / net/http :9000", "health, info, manual snapshot/publish trigger, query, list snapshots") Container(worker, "cmd/worker", "Go / gocron", "Runs the actual snapshot and publish cron jobs unattended") Container(archiveOrch, "archive.Orchestrator", "Go", "Runs all 9 Snapshotters; aggregates errors, doesn't stop on one failure") Container(snapshotters, "Snapshotters", "Go / pgx", "One per service โ€” weevil, owl, magpie, shrike, greyseal, lynx, woodrat, narwhal, rabbit") Container(writer, "archive.WriteParquet / SnapshotKey", "Go / parquet-go", "Encodes rows to Parquet; builds service/table/year/month/day/unix.parquet keys") Container(publishOrch, "publish.Orchestrator", "Go", "Runs all Publishers; aggregates errors") Container(publishers, "Publishers", "Go / ConnectRPC", "weevil, lynx, owl โ€” read each service's own API, never its DB") Container(storageClient, "storage.Client", "Go / gocloud.dev blob + AWS SDK v2", "Wraps a MinIO or S3 bucket; used for both the private snapshot bucket and the public publish bucket") Container(engine, "query.Engine", "Go / DuckDB + httpfs", "In-process DuckDB query engine reading Parquet directly off MinIO") } SystemDb(minio, "MinIO", "Private bucket") SystemDb(publics3, "Public AWS S3", "Public bucket") SystemDb(postgres, "9 service Postgres DBs", "") System_Ext(svcapis, "weevil / lynx / owl APIs", "") Rel(api, archiveOrch, "POST /api/v1/snapshot") Rel(api, publishOrch, "POST /api/v1/publish") Rel(api, engine, "POST /api/v1/query") Rel(api, storageClient, "GET /api/v1/snapshots (list keys)") Rel(worker, archiveOrch, "cron: SNAPSHOT_CRON (default hourly)") Rel(worker, publishOrch, "cron: PUBLISH_CRON (default every 6h)") Rel(archiveOrch, snapshotters, "Run() each") Rel(snapshotters, postgres, "SELECT ... ORDER BY") Rel(snapshotters, writer, "WriteParquet(rows)") Rel(snapshotters, storageClient, "Upload(bucket, key, parquetBytes)") Rel(storageClient, minio, "s3blob (path-style, MinIO endpoint)") Rel(publishOrch, publishers, "Publish() each") Rel(publishers, svcapis, "ConnectRPC list/get calls") Rel(publishers, storageClient, "Upload(publicBucket, key, json)") Rel(storageClient, publics3, "s3blob (virtual-hosted, AWS default endpoint)") Rel(engine, minio, "httpfs SELECT ... FROM read_parquet('s3://...')")

Overview

Pangolin is a two-binary Go service sharing one internal codebase: cmd/api serves a small HTTP API, and cmd/worker runs the actual cron jobs. Both build the same set of nine archive.Snapshotters and three publish.Publishers from internal/config; the worker drives them on a schedule via gocron, while the API exposes manual trigger endpoints plus a read-only DuckDB query surface over what’s already been snapshotted.

There are two independent data paths, kept deliberately separate:

  1. Archive (internal/archive) โ€” reads each service’s Postgres database directly via pgx, one Snapshotter per service, and writes the rows out as Parquet to a private MinIO bucket. This is the “snapshot everything” path and runs hourly by default.
  2. Publish (internal/publish) โ€” reads a service’s own ConnectRPC API (never its database) and writes per-entity JSON to a public, internet-reachable AWS S3 bucket, for static sites and browser widgets to fetch directly. This only covers services that want a public-facing export (weevil, lynx, owl today) and runs every 6 hours by default.

Archive: Snapshot Pipeline

Each of the nine Snapshotters (internal/archive/*.go) follows the same shape:

  1. Opens its own *sql.DB against one service’s Postgres instance (database/sql + jackc/pgx/v5/stdlib), using a connection string from internal/config (e.g. WEEVIL_DATABASE_URL).
  2. Runs one or more SELECT ... ORDER BY queries against that service’s tables, scanning rows into small Go structs tagged for Parquet (e.g. WeevilBook, LynxFeed, RabbitTicket).
  3. Encodes the rows with archive.WriteParquet[T] (a generic wrapper around parquet-go).
  4. Uploads the resulting bytes to MinIO under a key built by archive.SnapshotKey(service, table, time.Now()), which partitions as service/table/year=YYYY/month=MM/day=DD/<unix>.parquet โ€” so every run adds a new file rather than overwriting the previous snapshot.

The archive.Orchestrator (internal/archive/archive.go) runs every registered Snapshotter in sequence, collects errors from all of them (one service’s failure doesn’t stop the others), and returns an aggregate error. It can run all of them (Run) or just one by name (RunOne, used by POST /api/v1/snapshot {"service": "..."}).

Snapshotted services and what they cover, per the code in internal/archive/:

ServiceTables snapshotted
weevilbooks
owlbooks, papers
magpieresources, labels
shrikeindexrecords
greysealconversations, messages
lynxfeeds, websites (joined with site metadata)
woodratfiles
narwhalproducts, design_docs
rabbitprojects, sprints, tickets

A comment on several snapshotters (e.g. weevil.go, rabbit.go) notes that the column lists were hand-fixed after an earlier version queried columns that don’t actually exist in the target service’s schema (see commit history: fix: archive snapshotters query columns that don't exist in their services' schemas) โ€” a reminder that these queries aren’t generated from the target schema and can drift.

Storage Layer

internal/storage/minio.go wraps gocloud.dev/blob over an AWS SDK v2 s3.Client, and is used for both buckets pangolin talks to:

  • The private snapshot bucket, addressed via a MinIO endpoint (MINIO_ENDPOINT, path-style addressing, UsePathStyle: true).
  • The public publish bucket, addressed with cfg.Endpoint left empty so the AWS SDK falls back to real AWS S3 endpoint resolution and virtual-hosted-style addressing (PUBLIC_S3_* env vars).

Both cmd/api and cmd/worker construct two separate storage.Clients at startup for this reason. Only the private/MinIO client gets EnsureBucket called on it (a HeadBucket-then-CreateBucket check); the public S3 client deliberately skips this, since its credentials are typically narrower (write-only to a prefix) and wouldn’t be allowed to CreateBucket anyway โ€” the bucket is assumed to already exist.

Query Layer

internal/query/engine.go opens an in-process DuckDB database (marcboeker/go-duckdb, sql.Open("duckdb", "")) and installs/loads the httpfs extension, then configures it to talk to the same MinIO endpoint/credentials as the snapshot bucket (s3_endpoint, s3_url_style='path', access/secret keys). This lets arbitrary SQL submitted to POST /api/v1/query run read_parquet('s3://bucket/weevil/books/...')-style queries directly against the Parquet files on MinIO โ€” there’s no separate warehouse or load step; DuckDB reads the object storage files in place.

Publish: Public Export Pipeline

internal/publish is architecturally distinct from internal/archive, per its package doc: a Publisher reads a service’s own API (ConnectRPC), never its database, and writes public, per-entity files meant for direct browser/static-site consumption โ€” not a dated archive. Today there are three: WeevilPublisher, LynxPublisher, OwlPublisher, each writing a <service>/list.json (the full list response) plus one <service>/<uuid>.json per entity, re-serialized verbatim with protojson so client code generated from the same service’s schemas can decode it unchanged. The publish.Orchestrator mirrors archive.Orchestrator’s run-all/run-one shape and is driven by PUBLISH_CRON in the worker, or manually via POST /api/v1/publish.

Scheduling

cmd/worker/main.go is the only process that actually schedules anything, using go-co-op/gocron. It registers two cron jobs from internal/config:

  • SNAPSHOT_CRON (default 0 * * * * โ€” hourly) runs archive.Orchestrator.Run.
  • PUBLISH_CRON (default 0 */6 * * * โ€” every 6 hours) runs publish.Orchestrator.Run.

cmd/api/main.go builds the same orchestrators but never schedules them โ€” it only exposes them for manual, on-demand triggering over HTTP (POST /api/v1/snapshot, POST /api/v1/publish), plus the DuckDB query and snapshot-listing endpoints. In other words: the worker is the cron daemon, the API is the manual/query surface, and they can be deployed and scaled independently (see docker-compose.yml โ€” pangolin-api and pangolin-worker are separate services/images).

Process Inventory

ProcessSourcePortNotes
API servercmd/api/main.go9000 (mapped to 9100 in docker-compose.yml)Manual snapshot/publish trigger, DuckDB query, snapshot listing, health/info
Workercmd/worker/main.goโ€”No HTTP surface; runs the gocron snapshot and publish jobs on a schedule

Known Limitations

The archive snapshots are not disaster-recovery backups. This is called out directly in the root README, and it’s worth repeating here because it shapes the architecture above:

  • MinIO’s data lives on the same disk/host as the Postgres instances it snapshots (see the minio and db services referenced in configuration) โ€” a host or disk failure takes out the primary data and the “backup” together.
  • There is no restore path implemented anywhere in this codebase (or any other service’s) โ€” the Parquet output only feeds the read-only DuckDB query engine. There is no coded way to rehydrate a service’s database from these files.

These snapshots are an audit/analytics export, not a backup, until a real off-host pg_dump-based path with a documented restore command exists.

External Dependencies (key)

PackageRole
jackc/pgx/v5Postgres driver used by every Snapshotter
parquet-go/parquet-goEncodes snapshot rows to Parquet
marcboeker/go-duckdbEmbedded DuckDB engine + httpfs for querying Parquet on MinIO
gocloud.dev/blob (s3blob)Bucket abstraction shared by the private MinIO client and the public S3 client
aws/aws-sdk-go-v2Underlying S3 client, credentials, and config resolution
connectrpc.com/connectConnectRPC clients used by the publish package to read weevil/lynx/owl’s own APIs
go-co-op/gocron/v2Cron scheduling in cmd/worker
rs/corsCORS handling for cmd/api’s HTTP mux