Project · Self-directed · Infrastructure · Automation · Data
Pi Ops: a self-hosted platform on one Raspberry Pi
14 services, 38 active workflows and one 8 GB computer. Resource-aware scheduling so nothing heavy starts just because a clock fired.
- containers
- 14
- active workflows
- 38
- RAM, total
- 8 GB
Summary
Everything here runs on a Raspberry Pi 5: n8n, Arkham, a Discord bot, voice cloning, three render services, a podcast renderer, websites and a media server. I moved it over from n8n Cloud and Google Cloud Run. Pi Ops is the layer that keeps them from fighting each other: a heavy-resource lock, production and publication schedulers, a health monitor, a lock sweeper and a resource diary.
The problem
The Pi fully rebooted twice. The audit found the same short regenerated 17 times, a story rendered twice (about 5 hours of Pi time lost), and failed renders recorded as successful runs.
Why I built it
Cloud render costs and scattered services were the wrong shape for one person. A single box is cheap and easy to reason about, but only if it's managed as a constrained resource.
The solution
One principle: the Pi is the constrained resource, n8n is the orchestrator, production is asynchronous and publication is scheduled. Heavy jobs take a lock, check the time window and check they'll fit. A systemd sampler records CPU, memory, swap and temperature so the rules are based on data.
What it does
- Audited every workflow against execution history and wrote up the findings in a resource-diary workbook.
- A heavy-resource lock with a sweeper for stale locks, so voice rendering, video rendering and the podcast never overlap.
- Found the cause of a crash: TTS inference froze the box and the hardware watchdog reset it. Capped voice-render at 3 CPUs and gave n8n priority CPU shares.
- Moved public access from Tailscale Funnel to a Cloudflare Tunnel after a DNS outage. Admin panels are locked to home and tailnet addresses.
- Credential audit across 25 migrated workflows: credentials with the right name but the wrong ID were failing silently.
Technologies, and what they're for
- Raspberry Pi 5 (8 GB)The whole platform
- Docker Compose14 services with memory and CPU limits and health checks
- Cloudflare TunnelHTTPS public hostnames without opening ports
- TailscalePrivate admin access
- n8n data tablesLocks, queue ledger, config, resource history
- systemd timersResource sampler
AI involvement
- Claude Code as a pair engineer for audits, migrations and incident investigations
Automation
- Health monitor
- Error alerts
- Queue sync
- Lock sweeper
- Production scheduler (dry run first)
Infrastructure
- Everything on this page
Hard parts
- The Pi's firmware disables cgroup memory accounting, so per-container memory reads zero. Sampling is done at host level instead.
Outcome
Stable multi-tenant hosting for all projects, with alerting. The production scheduler is running in dry-run mode and publication switches over once it's approved.