Skip to content

The 401 spam that wasn't HA

Gotify kept pinging with 401s. The natural read is "an app's token is broken" or "gotify is down." Neither was true. The gotify server is healthy, and the 401 was content inside a forwarded log line. The thing posting the alert was loggifly, the log-monitor container, matching a keyword in a container's access log and forwarding it through Apprise.

This is the class of work: gotify ping, who actually posted it, what container logged it, who the real client is, what the true cause is.

Read the sender first

The gotify database is a readable SQLite file, and you can query it with the stdlib sqlite3 module without an app token or admin login. The application_id on each message tells you which gotify app posted it. If it's the Loggifly app, the alert is a log match, not an app failing. The title often literally names what matched:

'Regex: \b401 Unauthorized\b' and 'unauthorized' found in mealie

That tells you both a per-container regex keyword and the global unauthorized keyword fired. To silence one app you have to neutralize both layers.

The real client

The message body is the raw log line, and the [ip:port] in it is the real client. Cadence is the strongest fingerprint. A web SPA or browser iframe polls /sw.js every couple of minutes plus /api/app/about plus one authenticated call. That's a signed-in display, a Mealie iframe in HA, a smart-screen. Not an API client.

And that's the cause. A Mealie kiosk 401 is an expired web session, not a bad API token. GET /api/app/about returns tokenTime: 48, the signed-in web-session JWT TTL in hours. A signed-in display holds a browser session, and after 48 hours the authenticated query 401s while the cached SPA shell still returns 200. A 30-day tally looked like about 2,600 200s for static and SPA assets against a handful of 401s for the authed query. The fix at source is to re-sign in on that display, or give the kiosk a long-lived API token so it stops relying on the 48-hour session.

The matcher was the problem

The gotify spam itself was the log matcher. global_keywords: unauthorized is a sledgehammer, it pings on the word "unauthorized" from every watched container, including benign session expiries and the agent's own test calls. The durable fix is to tighten the matcher, drop the global keyword, keep only the per-container nginx 401/403 regex for real external threats. Not to "fix" each app.

One gotcha that cost me a real mistake: loggifly is per-host and multiple instances exist. An instance runs on the gateway, the personal VM, and the media VM, each with its own config. The instance that generates the ping is the one watching the container that logged it, on the same host. Mealie runs on the personal VM, so the personal-VM loggifly fired. I edited the gateway config, fired a test 401, and it still paged a second later. Enumerate instances on each VM before editing.

The same rule, different story

The same triage discipline caught a different one. Every Google calendar under one integration went unavailable at the same time while the integration itself still read loaded. When every calendar under one account drops together and the integration loads fine, it's almost always the OAuth token for that account going stale, the background refresh stopped producing a valid access token. I ruled out the network first, googleapis.com answered in about 50ms with a 401, which means reachable but no token. DNS, TLS, and routing were all healthy. That ruled out firewall and DNS before touching the token.

Don't jump to "the app's token is broken." Read who posted the alert, read the log line, identify the client by cadence, and rule out the network before you blame a token.

The wiring is on the notifications page and the Mealie page.

Comments