Give your AI agent a private messaging home: Signal and self-hosted Mattermost¶
Every AI agent eventually needs a place to live where you can talk to it from your phone, where it can talk to your family, and where nothing you type crosses someone else's server. Mine lives in two places: a self-hosted Mattermost server for the project channels and interactive bits, and Signal for the private stuff. This is the whole setup — the bot account, the admin permissions, the two custom Mattermost plugins, the Google Voice number that keeps my personal phone out of it, and the model-routing decisions that keep the privacy story honest.
Why not just the web UI¶
The web chat that ships with most agent harnesses is fine for debugging and useless for daily life. My agent needs to:
- receive messages from my phone while I'm out of the house
- be a member of project channels where it gets mentioned like a coworker
- deliver scheduled jobs (morning briefs, watchdog alerts) to a channel I actually read
- let other family members talk to it without handing them a server login
That's a messaging platform, not a chat window. And once the agent lives in a messenger, the platform choice becomes a privacy decision — which is most of this post.
Self-hosted Mattermost¶
Mattermost is the piece that does the heavy lifting: teams, channels, threads, bots, and — unlike Signal — message editing and interactive buttons. I run it in Docker on the personal VM using the official mattermost/docker compose layout — a mattermost service plus its own Postgres, everything driven by an .env:
# ~/docker/mattermost/docker-compose.yml (abridged — the real one is upstream's)
services:
postgres:
image: postgres:18-alpine
volumes:
- ${POSTGRES_DATA_PATH}:/var/lib/postgresql
mattermost:
image: mattermost/mattermost-team-edition:11.11.1
ports:
- "${APP_PORT}:8065"
volumes:
- ${MATTERMOST_CONFIG_PATH}:/mattermost/config:rw
- ${MATTERMOST_DATA_PATH}:/mattermost/data:rw
- ${MATTERMOST_PLUGINS_PATH}:/mattermost/plugins:rw
Both services run read_only with no-new-privileges and a memory cap — the upstream compose ships those hardening flags, I just left them on. The config directory is bind-mounted on the host, so config.json is editable without touching the container. TLS comes from swag: chat.thomasjwilde.com proxies to port 8065, same pattern as every other service behind the gateway.
The bot account and its permissions¶
The agent connects as a dedicated bot user, not a shared human account. In the Mattermost UI: Integrations → Bot Accounts → Add Bot, then generate a personal access token for it. That token goes into the agent's .env and nowhere else.
The part people get wrong is permissions. A default bot can post in channels it's been added to — and that's it. It can't create channels, can't create teams, can't invite people. My agent creates project channels on demand (every homelab project gets one), so the bot gets system admin rights: System Console → Users → find the bot → change system role to admin.
An admin bot is a real blast radius
A system-admin bot token that leaks is a full server compromise. Two things contain it: the token lives in a chmod-600 .env on the agent box and never appears in shell command text, and the agent's own allowlist (allowed_channels in Hermes config) restricts which channels it will even respond in. If you don't need channel creation, don't grant admin — a plain account with channel-creation permission in specific teams is enough for most setups.
Mention gating¶
One server, many channels, one bot — the bot must not reply to everything. Hermes gates by platform config:
mattermost:
require_mention: true
free_response_channels: <homelab-channel-id> # the one channel where it can chime in
allowed_channels: <comma-separated channel ids>
channel_prompts:
<channel-id>: "This channel is dedicated to the LLM Slot Proxy project..."
require_mention: true means the bot only wakes when @-mentioned everywhere except the free-response channel. channel_prompts injects a per-channel system note so the agent knows what room it's standing in. This is what makes one bot plausible across a dozen project channels without becoming noise.
Two custom Mattermost plugins¶
Mattermost's defaults fight an agent in two specific ways, and I patched both. These are Hermes-side plugins and gateway patches, not Mattermost server modifications — the server stays stock.
The ! prefix for slash commands in DMs¶
The Mattermost mobile app refuses to send any message starting with a bare / unless it matches a registered slash command. So /model typed in a DM to the bot just... doesn't send. !model does.
The fix is a small Hermes plugin that hooks pre_gateway_dispatch and rewrites !model into /model before dispatch. The safety details matter:
- it only fires when the message starts with
! - it validates the word against the gateway's actual command list —
!hellostays prose, never a mangled command - it only touches DMs; channels already work with native
/cmdand@bot /cmd - every error path falls through to normal dispatch — a broken plugin can never eat a message
It's generic — any DM platform that blocks bare / benefits — and it lives in ~/.hermes/plugins/, which survives harness updates.
In-screen button menus (the model picker)¶
Typing /model qwen3.8-27b with an exact model name is a bad UX on a phone. So /model, /reasoning, and /fast render as real Mattermost interactive buttons and dropdowns instead of text menus. You tap, the setting changes, no typing.
The architecture detail that trips people up: Mattermost buttons are server-proxied. Your phone never calls the agent's callback URL — the Mattermost server does the outbound POST. So the wiring is three pieces:
- a webhook receiver running inside the agent's gateway process (port 8644)
- a swag block fronting it with TLS:
mm-hook-ted.thomasjwilde.com - the adapter posting blocks with an action registry pointing at that URL
The payoff: a button tap calls a stored Python closure — the actual model-switch function — directly. No LLM in the click path, no re-parsing a slash command, deterministic. Same pattern drives button menus that trigger whole agent runs from a channel.
Where this code lives
These are patches to the Hermes gateway plus one user plugin. The canonical versioned copy is a private repo (hermes-mm-buttons: four patches, the plugin, an apply script, and a handoff doc). If you want the upstream version of the idea, the interactive-block plumbing is the interesting part — the button registry and the per-picker token validation are where the bugs live.
Signal, with a Google Voice number¶
Signal is the right home for the private conversations: end-to-end encrypted, minimal metadata, and no server for me to maintain — the bridge is just signal-cli running as a linked device.
The first decision is the phone number. Do not link the bot to your personal number. A linked-device setup means the bot shares your number's contact graph — anyone who messages the bot is, from Signal's perspective, messaging you. I use a dedicated Google Voice number for the agent:
- Get a Google Voice number (or any VoIP number you control — Signal accepts VoIP numbers for registration).
- Install Signal on a spare Android device (or emulator) and register it with the GV number — verification arrives as an SMS to Voice.
- On the agent box, install
signal-cli(Java 17+), then link it as a secondary device of that spare install:
signal-cli link -n "HermesAgent" # prints a QR code
# spare phone: Signal → Settings → Linked Devices → scan it
signal-cli --account +1XXXXXXXXXX daemon --http 127.0.0.1:8080
- Point Hermes at the daemon:
# ~/.hermes/.env
SIGNAL_HTTP_URL=http://127.0.0.1:8080
SIGNAL_ACCOUNT=+1XXXXXXXXXX
SIGNAL_ALLOWED_USERS=+1YYYYYYYYYY # my number; nobody else gets in
SIGNAL_GROUP_ALLOWED_USERS=<group-id> # the family "Home" group
SIGNAL_HOME_CHANNEL=<group-id> # cron deliveries land here
Two Google Voice caveats worth knowing before you build on it. Google reclaims numbers that sit idle, so keep the number alive with occasional activity — a bot that only receives messages can look dead to Google. And the number is tied to a Google account, which is a metadata thread back to Google; keep that account separate from your main one. If either bothers you, a prepaid SIM in the spare phone is the fully-decoupled version.
The allowlist is the security boundary here — Signal has no channel concept, so SIGNAL_ALLOWED_USERS is the whole gate. Unknown senders get a pairing code that I approve manually.
The model routing that keeps this honest¶
A private messenger in front of an agent that phones a cloud API isn't private — it's just encrypted until the agent's provider call. Three routing decisions make the story true.
Main brain: local, always¶
The primary model is a 27B running on my own GPU box, exposed as an OpenAI-compatible endpoint the harness treats like any provider. Conversations over Signal and Mattermost go to my hardware, full stop.
Aux models: small jobs, thinking off¶
The harness quietly uses auxiliary models for mechanical side-jobs: describing images, extracting text from web pages, compressing old conversation context. These are not reasoning tasks — they're transforms — and a thinking model on them wastes seconds and tokens per call, sometimes producing worse structured output because it "reasons" its way out of the requested format. I disable thinking per aux task:
auxiliary:
web_extract:
extra_body:
chat_template_kwargs:
enable_thinking: false
compression:
extra_body:
chat_template_kwargs:
enable_thinking: false
On a Qwen-family endpoint that flag is the documented switch; on other families it's thinking: {type: "disabled"} or a -nothink model variant. The principle is what matters: the big brain thinks, the little hands don't.
OpenRouter fallback: only with zero data retention¶
When the GPU box is down — and it has been down — the agent needs some brain or it stops being an agent. My fallback is OpenRouter, and that's a deliberate privacy trade: conversations leave the house. The floor I won't go below is zero data retention. OpenRouter's provider pool includes ZDR endpoints (the provider doesn't log or train on the prompt), and you can require them at the routing layer:
data_collection: deny means the router will refuse any endpoint that declares data collection, even when it's cheaper. On OpenRouter's side you can also flip the org-level "require zero data retention" setting so the guarantee holds for anything else using the same key. It doesn't make the fallback private — it makes it not retained, which is the best you can do off-prem. The real fix for the fallback problem is a second GPU, which is a whole other post.
What this buys¶
The day-to-day result: the family talks to the agent in a Signal group, project work happens in Mattermost channels where it's @-mentioned like a teammate, scheduled briefs land in the same places humans read them, and the only thing that leaves my network is a fallback path that's explicitly ZDR and rare. The custom plugins are the difference between "a bot that works" and "a bot my family actually uses" — buttons on a phone beat typing model names, and the ! trick means the mobile app stops eating commands.
The video version of this walkthrough is coming to the channel; the handoff doc behind the Mattermost patches is linked from the repo above.