Skip to content

AI

Do AI coding agents phone home? I grepped five of them

The opencode privacy panic made the rounds this year: "local-first" coding agent, allegedly proxying everything through their cloud, PostHog and Honeycomb baked in. Then the author of the original source audit publicly walked it back. Both things were true, which is exactly why I didn't want to take anyone's word for it — including the walk-back. So I cloned every agent harness running in my house, grepped the actual source for telemetry endpoints, and cross-checked the claims against 24 hours of whole-LAN DNS logs from AdGuard Home.

Four open-source coding agents: serve modes, Android clients, and chat integrations

Last post was about what these harnesses phone home. This one is about what they're for — the features I care most about in a self-hosted setup: can the agent serve a web app I can reach from anywhere, can it keep working on a goal across turns until the work is actually done, and can it talk to my chat server. Plus the third-party Android ecosystem that's quietly grown up around driving coding agents from the couch.

Method is the same: I cloned all four (opencode, pi, DeepSeek Harness, qwen-code) and read the source and shipped docs, and checked each Android client's repo directly. Versions tested: opencode dev branch, pi main, dsh main, qwen-code 0.25.0 — all October 2026.

A container runs what it was last created with

This one applies well beyond the box it happened on, so I'm writing it down separately.

I had a swap script that started a container with docker start. I'd edited the compose file to point at a new image and a new model path, ran the swap, and it came up running the old pair. Nothing was wrong with the compose file. Nothing was wrong with the script. The container was simply doing what containers do.

The bug that wasn't in upstream

After a slot restore, the first request came back at 0% cached. The second one was warm at 99%. Nothing looked broken — the KV rehydrated fine, the logs were clean, and yet every restore paid a full re-prefill on the turn right after it.

A front layer for local inference

The llama slot proxy is the piece I'd least expect someone to build and the piece I'd least like to work without. It sits on port 8081 and fronts whatever engine currently owns the GPU, and it's the only thing the clients talk to.