Four open-source coding agents: serve modes, Android clients, and chat integrations¶
Last post was about what these harnesses phone home. This one is about what they're for — the features I care most about in a self-hosted setup: can the agent serve a web app I can reach from anywhere, can it keep working on a goal across turns until the work is actually done, and can it talk to my chat server. Plus the third-party Android ecosystem that's quietly grown up around driving coding agents from the couch.
Method is the same: I cloned all four (opencode, pi, DeepSeek Harness, qwen-code) and read the source and shipped docs, and checked each Android client's repo directly. Versions tested: opencode dev branch, pi main, dsh main, qwen-code 0.25.0 — all October 2026.
Serve mode: the feature that separates the pack¶
| Harness | Version tested | Serve mode | Transport | Auth story | Source |
|---|---|---|---|---|---|
| opencode | dev, Oct 2026 | opencode serve — "the v2 API and web server"; hostname/port/CORS flags; web UI assets baked into the binary |
HTTP REST + SSE, typed SDK | none built in — you front it with your own proxy | repo |
| pi | main, Oct 2026 | experimental pi-server package — routes clients to durable sessions; --mode rpc for stdio JSON-RPC |
unix socket / stdio | n/a (local plumbing) | repo |
| DeepSeek Harness | main, Oct 2026 | dsh web — node:http server with optional TLS, serves the SPA "Web shell", /api bridge, plugins register named routes and WebSocket upgrades |
HTTP(S) + WS | bind address is a single concrete literal; wildcard binds are rejected at load | repo |
| qwen-code | 0.25.0 | qwen serve daemon + built-in Web Shell; --local-control prints a QR pairing code for phones |
HTTP + SSE, plus authenticated webhook ingress (x-qwen-webhook-secret) |
bearer token, revocable pairing token, TLS cert flags | docs |
Three of the four serve a real web app. pi is the honest outlier — its server package is plumbing for durable sessions over a unix socket, explicitly experimental, and the project's philosophy is a great terminal and nothing else.
dsh's web server deserves more credit than I initially gave it. The command is dsh web (not "serve", which is why my first grep missed it), and the architecture is arguably the cleanest of the bunch: TLS is core config, the bind address must name one concrete interface — a wildcard address is rejected at load instead of silently opening every NIC — and every feature including the SPA is a plugin registering routes on a plain node:http server. It also ships a profile matrix for programmatic clients: dsh web, headless, sdk, sdk-minimal, acp.
qwen-code is the one that clearly wants to be remote. qwen serve --local-control picks a private LAN interface, mints a revocable pairing token, and prints a QR code; your phone scans it and lands in the same Web Shell session — conversation, file diffs, tool calls, and permission prompts. The pairing credential rides in the URL fragment, not a query param, so it stays out of request logs. For HTTPS (needed for voice input in mobile browsers), you pass --tls-cert/--tls-key with the LAN IP in the SAN.
Multi-turn goal prompts: who keeps working until done¶
A "goal prompt" feature means you hand the agent one objective and the session keeps taking its own turns until the objective is verifiably done — not just a todo list inside one turn. Only two of the four have it, and one of those two has the most rigorous implementation I've seen anywhere in this space.
| Harness | Goal system | Auto-continuation | Completion check | Source |
|---|---|---|---|---|
| opencode | none — todo tool tracks tasks within a session, no goal object | none — works until the model decides it's done | model self-assessment only | repo |
| pi | none — pi-durable is crash-resilience (resume where you stopped), not goal-driven continuation |
none | n/a | repo |
| DeepSeek Harness | one durable goal per session, survives restarts/resumes/forks; model tools get_goal/create_goal/update_goal; human /goal command that costs no model turn |
separate goal-round-driver package turns an active goal into sequential Rounds — opt-in, not default |
goal records completion state; the driver just keeps the rounds coming | repo |
| qwen-code | /goal <objective> with pause/resume/edit/clear, plus /goal-draft to turn a fuzzy intention into a verifiable objective |
yes — the session keeps going on its own until the verifier accepts or a budget stops it | independent verifier reads the transcript and rejects unsupported completion claims | docs |
qwen-code's verifier design is worth stealing from even if you never run the tool. When the model proposes "done," a separate model call judges that claim using only what's already in the transcript — visible output, tool results, and your own messages; not the model's hidden reasoning. The rules are brutally good:
- "Tests pass" needs the actual test output in the transcript. Printed text proves only that text was printed.
- "You approved X" needs a real message from you; proposals that assume a confirmation that never happened get rejected.
- Missing evidence means the verdict is "not yet," not "done" — the loop keeps running.
- Budgets are enforced by the runtime, not the model's goodwill: a token window shown live (
1.2k/30.0m), optionalmodel.goalMaxTurnsandmodel.goalMaxActiveMinutescaps, and a single wind-down hand-off turn when any window runs out.
Their docs even give the objective template: Outcome: one sentence, Done when: numbered binary checks naming a command and its expected output, Must not: files and irreversible actions off-limits, On block: what to report and which decision needs a human. That's the anti-pattern for agents grading their own homework, written down.
dsh's design choice is the opposite knob: the goal records completion state but never schedules work by itself — automatic continuation is a separate package you enable deliberately. An anti-runaway default in a different shape.
opencode and pi just don't have the concept. opencode's todo tool is task tracking inside a single run; pi's durable package survives a crash but won't take a turn on its own.
Messaging integrations: who talks to your chat server¶
This is where the four diverge hardest.
| Harness | Built-in chat channels | Build-your-own path | Source |
|---|---|---|---|
| opencode | none | plugin system (event hooks) + REST/SDK — a bot against opencode serve is DIY |
repo |
| pi | none | DIY against --mode rpc / pi-server |
repo |
| DeepSeek Harness | none | ACP server — JSON-RPC stdio for programmatic clients: create/resume/close sessions, attach MCP, answer permission prompts without a human | repo |
| qwen-code | Telegram, DingTalk, Feishu, WeCom, WeChat, QQ, email, GitHub, GitLab — daemon-managed channel workers (qwen serve --channel telegram), hot-swappable without restarting the daemon |
@qwen-code/channel-base: subclass, implement connect / sendMessage / disconnect |
repo |
qwen-code is the only one with a first-class messaging architecture — channel workers are separate processes the daemon owns, inbound webhooks are authenticated by design, and the channel SDK docs say a custom adapter's only dependency is channel-base. Alibaba built this to live inside workplace chat (DingTalk/WeCom/Feishu are the Chinese equivalents of the ask), and the shape transfers: a Mattermost channel adapter is roughly three methods plus their existing webhook ingress.
Nobody in this lineup speaks Mattermost out of the box. The harness that does is the one I run at home — Hermes ships Mattermost, Telegram, and Signal gateways natively — and it can also act as a client of the others' serve APIs. So the architecture question isn't really "which coding agent can join chat"; it's whether you want coding-agent sessions reachable from chat, and the shortest paths there are qwen-code's channel SDK or a small bridge bot in front of opencode serve.
The Android ecosystem¶
opencode has the most mature third-party mobile scene, and it dragged the others along faster than expected.
| App | Harness | Stack | Status | Source |
|---|---|---|---|---|
| OpenCode Mobile | opencode | Android, talks to your own opencode serve over its open API |
Google Play, F-Droid, or APK; stream sessions, review diffs, approve tool calls | repo |
| opencode-android | opencode | thin Android client — chat, terminal, file browse/edit, all server-driven | GitHub | repo |
| deepseek-harness-app | dsh | Flutter, MIT — connects to a stock dsh web host, no patches; live run watching, tool-call approvals, workspace/model/goal/subagent management, on-device voice input (sherpa-onnx) |
v0.1.3 released Oct 10 2026, 699 commits, CI + fastlane metadata (F-Droid-ready) | repo |
| dsh-android-app | dsh | native Android — chat, approvals, workspace, files, models; Chinese-first | GitHub | repo |
| DSHBox | dsh | runs the whole harness on the phone — Debian + Node + PRoot + WebView in one APK, no server | site | |
| (qwen-code) | qwen-code | no native app — official path is the --local-control QR-paired Web Shell in the phone browser; an official Android companion is an open proposal |
issue #11704 | |
| (pi) | pi | none — no web server to talk to | — |
The dsh mobile scene is the surprise: the ChanceFlow client is exactly the opencode-mobile architecture — point an APK at an unmodified dsh web on your own box — and it's under heavy active development. DSHBox is the opposite philosophy: no server at all, the harness runs on-device. Worth knowing both exist.
The three-feature pattern¶
Every remote-control client in this space converges on the same three features: stream the session, review the diff, approve the tool call. That's the minimum to keep a human in the loop while away from the desk, and it's a good lens for evaluating any of these — if a client does more than that, check whether the extras are worth the attack surface.
The security note nobody should skip¶
A serve API is code execution as the daemon user. qwen-code's own docs say the quiet part plainly: anyone holding the pairing credential has the authority granted to that Web Shell, Local Control is for a trusted LAN, and it is not a shortcut for exposing the daemon to the internet. The harnesses handle this with different levels of care — dsh rejects wildcard binds at load, qwen-code gates every route behind a bearer token and puts pairing secrets in the URL fragment, opencode leaves auth entirely to you and hands you CORS flags instead.
My pattern for any of these: serve on a homelab box, front it with swag/pangolin behind the existing auth wall, phone connects over that or Tailscale. Same rule as everything else on this site — the agent gets an HTTPS-usable surface with auth in front, never a raw port.
What I'd actually run where¶
For the family's existing opencode habit, OpenCode Mobile against Roy's box is the finished product. For trying the newest ecosystem with the least fuss, the dsh Flutter APK against a dsh web container pointed at the llama slot-proxy is the closest thing to opencode-mobile outside opencode. For a chat-native experiment, qwen-code's channel SDK is the shortest path to "the agent lives in Mattermost" — though Hermes already lives there, so honestly the interesting version of that experiment is Hermes driving one of these serve APIs, not replacing it.
The telemetry side of all this — what each harness phones home by default, and what I blocked — is in the companion post.