Architecture
System Context
Container Diagram
Overview
Pangolin is a two-binary Go service sharing one internal codebase: cmd/api serves a small HTTP API, and cmd/worker runs the actual cron jobs. Both build the same set of nine archive.Snapshotters and three publish.Publishers from internal/config; the worker drives them on a schedule via gocron, while the API exposes manual trigger endpoints plus a read-only DuckDB query surface over what’s already been snapshotted.
There are two independent data paths, kept deliberately separate:
- Archive (
internal/archive) โ reads each service’s Postgres database directly viapgx, oneSnapshotterper service, and writes the rows out as Parquet to a private MinIO bucket. This is the “snapshot everything” path and runs hourly by default. - Publish (
internal/publish) โ reads a service’s own ConnectRPC API (never its database) and writes per-entity JSON to a public, internet-reachable AWS S3 bucket, for static sites and browser widgets to fetch directly. This only covers services that want a public-facing export (weevil, lynx, owl today) and runs every 6 hours by default.
Archive: Snapshot Pipeline
Each of the nine Snapshotters (internal/archive/*.go) follows the same shape:
- Opens its own
*sql.DBagainst one service’s Postgres instance (database/sql+jackc/pgx/v5/stdlib), using a connection string frominternal/config(e.g.WEEVIL_DATABASE_URL). - Runs one or more
SELECT ... ORDER BYqueries against that service’s tables, scanning rows into small Go structs tagged for Parquet (e.g.WeevilBook,LynxFeed,RabbitTicket). - Encodes the rows with
archive.WriteParquet[T](a generic wrapper aroundparquet-go). - Uploads the resulting bytes to MinIO under a key built by
archive.SnapshotKey(service, table, time.Now()), which partitions asservice/table/year=YYYY/month=MM/day=DD/<unix>.parquetโ so every run adds a new file rather than overwriting the previous snapshot.
The archive.Orchestrator (internal/archive/archive.go) runs every registered Snapshotter in sequence, collects errors from all of them (one service’s failure doesn’t stop the others), and returns an aggregate error. It can run all of them (Run) or just one by name (RunOne, used by POST /api/v1/snapshot {"service": "..."}).
Snapshotted services and what they cover, per the code in internal/archive/:
| Service | Tables snapshotted |
|---|---|
| weevil | books |
| owl | books, papers |
| magpie | resources, labels |
| shrike | indexrecords |
| greyseal | conversations, messages |
| lynx | feeds, websites (joined with site metadata) |
| woodrat | files |
| narwhal | products, design_docs |
| rabbit | projects, sprints, tickets |
A comment on several snapshotters (e.g. weevil.go, rabbit.go) notes that the column lists were hand-fixed after an earlier version queried columns that don’t actually exist in the target service’s schema (see commit history: fix: archive snapshotters query columns that don't exist in their services' schemas) โ a reminder that these queries aren’t generated from the target schema and can drift.
Storage Layer
internal/storage/minio.go wraps gocloud.dev/blob over an AWS SDK v2 s3.Client, and is used for both buckets pangolin talks to:
- The private snapshot bucket, addressed via a MinIO endpoint (
MINIO_ENDPOINT, path-style addressing,UsePathStyle: true). - The public publish bucket, addressed with
cfg.Endpointleft empty so the AWS SDK falls back to real AWS S3 endpoint resolution and virtual-hosted-style addressing (PUBLIC_S3_*env vars).
Both cmd/api and cmd/worker construct two separate storage.Clients at startup for this reason. Only the private/MinIO client gets EnsureBucket called on it (a HeadBucket-then-CreateBucket check); the public S3 client deliberately skips this, since its credentials are typically narrower (write-only to a prefix) and wouldn’t be allowed to CreateBucket anyway โ the bucket is assumed to already exist.
Query Layer
internal/query/engine.go opens an in-process DuckDB database (marcboeker/go-duckdb, sql.Open("duckdb", "")) and installs/loads the httpfs extension, then configures it to talk to the same MinIO endpoint/credentials as the snapshot bucket (s3_endpoint, s3_url_style='path', access/secret keys). This lets arbitrary SQL submitted to POST /api/v1/query run read_parquet('s3://bucket/weevil/books/...')-style queries directly against the Parquet files on MinIO โ there’s no separate warehouse or load step; DuckDB reads the object storage files in place.
Publish: Public Export Pipeline
internal/publish is architecturally distinct from internal/archive, per its package doc: a Publisher reads a service’s own API (ConnectRPC), never its database, and writes public, per-entity files meant for direct browser/static-site consumption โ not a dated archive. Today there are three: WeevilPublisher, LynxPublisher, OwlPublisher, each writing a <service>/list.json (the full list response) plus one <service>/<uuid>.json per entity, re-serialized verbatim with protojson so client code generated from the same service’s schemas can decode it unchanged. The publish.Orchestrator mirrors archive.Orchestrator’s run-all/run-one shape and is driven by PUBLISH_CRON in the worker, or manually via POST /api/v1/publish.
Scheduling
cmd/worker/main.go is the only process that actually schedules anything, using go-co-op/gocron. It registers two cron jobs from internal/config:
SNAPSHOT_CRON(default0 * * * *โ hourly) runsarchive.Orchestrator.Run.PUBLISH_CRON(default0 */6 * * *โ every 6 hours) runspublish.Orchestrator.Run.
cmd/api/main.go builds the same orchestrators but never schedules them โ it only exposes them for manual, on-demand triggering over HTTP (POST /api/v1/snapshot, POST /api/v1/publish), plus the DuckDB query and snapshot-listing endpoints. In other words: the worker is the cron daemon, the API is the manual/query surface, and they can be deployed and scaled independently (see docker-compose.yml โ pangolin-api and pangolin-worker are separate services/images).
Process Inventory
| Process | Source | Port | Notes |
|---|---|---|---|
| API server | cmd/api/main.go | 9000 (mapped to 9100 in docker-compose.yml) | Manual snapshot/publish trigger, DuckDB query, snapshot listing, health/info |
| Worker | cmd/worker/main.go | โ | No HTTP surface; runs the gocron snapshot and publish jobs on a schedule |
Known Limitations
The archive snapshots are not disaster-recovery backups. This is called out directly in the root README, and it’s worth repeating here because it shapes the architecture above:
- MinIO’s data lives on the same disk/host as the Postgres instances it snapshots (see the
minioanddbservices referenced in configuration) โ a host or disk failure takes out the primary data and the “backup” together. - There is no restore path implemented anywhere in this codebase (or any other service’s) โ the Parquet output only feeds the read-only DuckDB query engine. There is no coded way to rehydrate a service’s database from these files.
These snapshots are an audit/analytics export, not a backup, until a real off-host pg_dump-based path with a documented restore command exists.
External Dependencies (key)
| Package | Role |
|---|---|
jackc/pgx/v5 | Postgres driver used by every Snapshotter |
parquet-go/parquet-go | Encodes snapshot rows to Parquet |
marcboeker/go-duckdb | Embedded DuckDB engine + httpfs for querying Parquet on MinIO |
gocloud.dev/blob (s3blob) | Bucket abstraction shared by the private MinIO client and the public S3 client |
aws/aws-sdk-go-v2 | Underlying S3 client, credentials, and config resolution |
connectrpc.com/connect | ConnectRPC clients used by the publish package to read weevil/lynx/owl’s own APIs |
go-co-op/gocron/v2 | Cron scheduling in cmd/worker |
rs/cors | CORS handling for cmd/api’s HTTP mux |