Case study · Self-directed · Software · AI · Simulation

Arkham: A Living Survey

A persistent 1926 New England town that simulates itself in real time. You can watch it, but you can't interfere.

LiveSep 2026 – presentVisit ArkhamCase archiveDiscord
investigators
128
mapped places
300
scripted incidents
160
lines of JavaScript
~17k
The live Arkham map on 28 September 1926 (Arkham time): lamplit investigator markers, the Dossier notebook with active investigations and this morning's Gazette headline
The live Arkham map on 28 September 1926 (Arkham time): lamplit investigator markers, the Dossier notebook with active investigations and this morning's Gazette headline

Summary

One Arkham, kept by a server and watched by anyone who opens the page. 128 investigators with their own diaries, preoccupations and relationships move around 300 places on a real clock set exactly a century behind ours. Disturbances build slowly at locations until a case opens. A Gazette reports on what happened, a Discord server gets a live channel for every investigation, and every third day a podcast is written and narrated from the town's own records.

The problem

Most 'living worlds' are really waiting rooms that only move when a player shows up. I wanted a town that carries on without an audience: people with memories, rumours that spread, trouble that builds slowly. It also had to stay coherent for months without me writing everything by hand.

The idea

Make the visitor an observer, not a player. The server owns a single world, and the page is only a window onto it. The simulation should be deterministic enough to trust but rich enough to surprise, and AI supplies new raw material around it without ever running the world.

The architecture

  1. 01

    Engine

    25 engine modules: geography, people, minds, diaries, the brewing system, cases and content packs. The same engine code runs in the browser's standalone build.

  2. 02

    World host

    Steps the simulation every few seconds, stores snapshots and permanent history in SQLite, and catches up on missed time after a restart.

  3. 03

    Public API

    World state, a live SSE stream, diaries, newspapers, scene images and the crawlable archive pages.

  4. 04

    Integration API

    Token-protected endpoints that n8n uses to validate and publish content, pull podcast scripts and relay case events.

  5. 05

    n8n

    Monthly content, Discord relay, podcast render and feedback collection. The heavy jobs share the Pi Ops resource lock.

  6. 06

    Edges

    The Discord bot runs in its own container, and the podcast renderer runs in its own container with voice-render.

The build

Node 22 with no framework, SQLite through the built-in node:sqlite, a build script that inlines the engine and UI into a single page, Docker Compose profiles for the bot and the podcast, and test scripts (84 Discord checks, 170 podcast checks, headless multi-year story simulations).

  • Node.js 22Simulation engine and HTTP server, with no framework
  • SQLite (node:sqlite)World snapshots, permanent history, diaries, Gazette, analytics
  • Server-Sent EventsLive world stream to every open browser
  • n8nMonthly content, Discord feed, podcast render, feedback
  • GeminiStorylines, investigators, portraits, Polaroids, voice lines
  • Discord gateway + REST (hand-written)A bot with no library dependency: live investigation channels, slash commands
  • Chatterbox TTS + ffmpegPodcast narration and vertical clips
  • Docker + Cloudflare TunnelHosting on the Pi, public at arkham-survey.com

The AI

AI supplies material but the world decides what happens. Gemini drafts storylines, investigators, voice lines, portraits and Polaroids, and every pack goes through the server's own validator before it can reach the world. A failed image is skipped, not fatal. The podcast writer deliberately uses no LLM because the town's records already hold the story.

  • Gemini writes monthly storylines and investigator packs. The server validates them and sends each one back to Gemini once for a fix if it's refused.
  • Gemini image models draw investigator portraits and black-and-white investigation Polaroids.
  • Around 6,750 voice and coping-style lines were generated in batches through n8n, then filtered for repetition.
  • Claude (through the Agent SDK, with a hard spending cap) can take over batch writing jobs from Gemini, such as voice and coping lines.

The automation

New content arrives on the 1st of each month, Discord channels open by themselves, the podcast goes from script to Drive every third day, and backups run nightly. I get an email when something goes live or fails.

  • The monthly update generates content, validates it, publishes it, posts a release note and emails me, and skips itself if that month is already done.
  • The Discord feed and a nightly reconcile keep the channels in step with the world.
  • The podcast render is written, voiced, clipped and uploaded to Drive every third day.
  • Daily backups cover the database, content packs and scene images.

The data

Every diary entry, newspaper and case is permanent history. A 'Day in Arkham' chronicle records inner changes, conversations and decisions, and exports as CSV. Visitor analytics are cookieless daily hashes kept for 180 days.

The result

A public, persistent world that has been running continuously since late September 2026, with an archive that grows by itself, a Discord server, a podcast and a monthly content cycle, all from one Raspberry Pi.

  • Built an event-driven relay that spots each new investigation, has the Discord bot open a channel for it, and streams the investigation's events into that channel live.
  • The world runs on the real clock, 5,218 weeks behind. After a restart the server fast-forwards through the missed time, so the town never pauses.
  • Locations carry hidden state (dormant, uneasy, active, escalating, manifesting…). Talk spreads, the Gazette picks it up, and when a case finally opens its file already holds everything that led up to it.
  • On the 1st of every month an n8n workflow has Gemini write a storyline and an investigator, draw a portrait and ten Polaroids, and submit the lot. The server rejects anything that fails validation.
  • Diary anti-repetition cut self-repeated phrasing from 45% to 13% across about 5,550 generated voice lines.
  • A podcast writer that uses no LLM turns three days of town records into a script, which is narrated in my own cloned voice, rendered on the Pi and filed in Drive.

Hard parts

  • A corruption pass escaped every template literal in the source. After the fix, a single missing '>' in the HTML head still made browsers treat the entire JavaScript bundle as plain text. I found it by working out why a literal ${…} was showing up in the markup.
  • Rendering podcast clips as one ffmpeg graph queued frames without limit (2.9 GB for one clip) and froze the Pi. Splitting it into one ffmpeg process per shot plus a concat pass kept memory flat at about 120 MB per shot.

What I learned

Headless multi-year runs showed the real bottleneck: the authored storylines run out after 12 to 21 months, and diary phrasing wears out fastest. That told me where AI-generated content was worth adding (voices and phrasing) and where it wasn't (the core rules). I also learned to protect the one always-on service from the heavy batch jobs sharing its machine.