AI AgentsHermesTelegramDiscordSlackWhatsAppMessaging

Hermes Agent Messaging: Choosing Your Platform — Telegram vs Discord vs Slack vs WhatsApp vs the Rest

One agent, 21+ messaging platforms — but "supports everything" is not a decision framework. The full tradeoff breakdown: setup friction, security posture, the capability matrix, and reliability across the five most common surfaces, plus the long tail and a decision guide that matches each surface to the working style it suits.

September 20, 2026Michel Laclé12 min read
🎧 Audio Version — Listen to This Report

1. The Architecture: One Gateway, Every Surface

Before comparing platforms, it's worth understanding what's actually being compared. Hermes' messaging layer is a single background gateway process that holds a live connection to every configured platform simultaneously. One agent core, one session store, one cron scheduler — and any number of fronts [Docs].

That architecture shape drives most of the decision. You're not choosing between five different agent products. You're choosing which door the same agent answers at — and each door has different plumbing, different lock quality, and different neighbors.

THE GATEWAY, IN ONE DIAGRAM ──────────────────────────────────────── Telegram ──┐ Discord ───┤ Slack ─────┤ ┌─────────────────────────┐ WhatsApp ──┼──────► │ messaging gateway │ Signal ────┤ one │ (background process) │ SMS ───────┤ proc │ • per-chat sessions │ Email ─────┘ │ • cron scheduler │ │ • delivery ledger │ └───────────┬─────────────┘ │ one AIAgent core (tools, memory, skills)

Three properties of the gateway matter for the platform choice:

  • Setup is wizard-driven. hermes gateway setup walks every platform with arrow-key selection, writes the config, and offers to restart. No public URL is required for any of the big five — Telegram, Discord, and WhatsApp run outbound connections from your machine, and Slack runs over Socket Mode WebSockets. The agent can live on a laptop, a private server, or behind a firewall.
  • Sessions are per-chat and persistent. Each DM, channel, or thread gets its own conversation context that survives restarts. In shared channels, Hermes isolates sessions per user by default (group_sessions_per_user: true) — two people in the same Slack channel or Discord thread do not share a transcript unless you explicitly opt into a shared room.
  • Delivery is at-least-once, with honest labeling. A durable delivery ledger records each response around platform send. If the gateway dies mid-send, the next boot redelivers — and if the platform may already have received the original, the redelivery carries a visible "♻️ Recovered reply — may be a duplicate" prefix. Ambiguity is labeled, never silently resent.

💡 The operational baseline

Every platform gets the same in-chat command set — /new, /model, /stop, /approve, /sessions, /background, and any installed skill invoked as a slash command. A model switch via /model provider:model persists across gateway restarts for that session. So the "what can I do once I'm connected" question is identical on every platform. The differences are entirely in how you connect, what the platform's API lets you do, and what can break.

2. The Capability Matrix — What Actually Differs

Hermes' own docs maintain a capability matrix across all supported platforms. It's the right starting point because it separates what the platform API offers from what Hermes implements. Seven dimensions matter in practice: voice (TTS replies and/or voice-message transcription), images, files, threads, reactions, typing indicators, and streaming (progressive message updates via editing) [Matrix].

PlatformVoiceImagesFilesThreadsReactionsTypingStreaming
Telegram ✅✅✅✅—✅✅
Discord ✅✅✅✅✅✅✅
Slack ✅✅✅✅✅✅✅
WhatsApp (Baileys bridge) —✅✅——✅✅
WhatsApp Cloud API ✅✅✅——✅✅
Signal —✅✅——✅—
SMS (Twilio) ———————
Email (IMAP/SMTP) —✅✅✅———

Three observations from the matrix, which do most of the heavy lifting for the decision:

  • The top tier is nearly a tie. Telegram, Discord, and Slack all support every capability dimension. Telegram's only gap is reactions; nothing else separates them in raw capability. If your workflow needs media, voice, threads, and streaming, all three qualify.
  • "Official" doesn't always mean "more capable." The WhatsApp Business Cloud API is Meta's sanctioned path — no ban risk, voice included — but the unofficial Baileys bridge, despite being the riskier path, ships things the Cloud API can't do: native polls (including rendering the agent's multiple-choice clarify questions as tappable poll options) and location pins. The two adapters can even run in parallel on different phone numbers.
  • The floor is brutal. SMS supports nothing — no media, no voice, no typing indicator, no streaming. Email supports images and threads but no real-time affordances at all. These platforms aren't for conversation; they're for reach.

Below the capability matrix, the five platforms differ along four axes that matter more in daily use: setup friction, security posture, reliability characteristics, and access-control model. Each section covers its platform on all four.

3. Telegram: Lowest-Friction Personal Surface

Telegram is the natural default for a personal agent, and the setup reflects it — it's the shortest path from zero to working on any platform [Setup]:

  • Setup: message @BotFather, run /newbot, paste the token, add your numeric user ID to TELEGRAM_ALLOWED_USERS. Done in under five minutes. No OAuth scopes, no event subscriptions, no app manifests, no WebSocket server.
  • Security posture: fully official Bot API. Access control is a numeric user-ID allowlist (discoverable via @userinfobot), with separate gates for group senders (group_allow_from) and whole groups (group_allowed_chats). The token is the only secret, and @BotFather can revoke it instantly if it leaks.
  • Capability notes: voice memos auto-transcribe; agent replies can come back as spoken audio; streaming edits the message progressively; files up to the platform limit ship as native attachments. A status_indicator option even writes "Online"/"Offline" into the bot's profile description on connect/disconnect — the closest the Bot API gets to a presence dot.
  • Reliability: long-lived bot connections are the most boring failure mode in the set. There are no privileged intents to toggle, no protocol versions to chase.

⚠️ The one trap: group privacy mode

Telegram bots default to privacy mode ON, which means the bot only sees /-commands, direct replies to itself, and service messages — not ordinary group messages. It's the single most common source of "the bot ignores my group" confusion. The fix is one of: disable privacy mode in @BotFather and remove/re-add the bot to each group (Telegram caches the privacy state at join time), or promote the bot to group admin. Neither is hard, but neither is discoverable until it bites.

Telegram in one line

Best for: personal use, voice-first operation, mobile-first access. Trade: no reactions, and group behavior requires the privacy-mode detour. The least setup of any full-capability platform.

4. Discord: Richest Team Surface, Most Setup Fragility

Discord is the most expressive platform in the matrix — it's the only one of the big four with reactions and voice channels (the agent can join a Discord voice channel and hold conversations there, not just DMs) [Setup]. It's also the one whose setup has the most sharp edges.

  • Setup: Developer Portal → new application → bot page → enable both privileged intents (Message Content and Server Members) → copy token → generate OAuth2 invite URL. The docs are blunt about it: the Message Content Intent being disabled is "the #1 reason Discord bots don't work" — the bot connects, receives message events, and the text is empty. The bot literally cannot see what you typed.
  • Security posture: official bot API with allowlists by user ID or by role ID (DISCORD_ALLOWED_ROLES — useful when a moderation team churns and you don't want to edit the allowlist for every hire). Notably, if you set neither, the gateway denies all users by default. There's also a built-in guard against the LLM accidentally emitting @everyone — by default the bot can't ping everyone or roles, only individual users and the person it's replying to.
  • Behavior model: DMs get per-DM sessions and answer everything. Server channels answer only on @mention by default, with free-response channels and per-user session isolation as knobs. Threads get isolated session namespaces. The bot defaults to staying silent on messages that mention other users but not it — so it doesn't jump into conversations directed elsewhere.
  • Reliability: this is where Discord is genuinely different. REST and the Gateway WebSocket are separate transports, and a successful REST call proves nothing about whether the bot can receive events. Hermes therefore runs a liveness check that combines ready state, socket closure, heartbeat ACK age, and heartbeat latency — configurable via websocket_liveness_interval_seconds and friends — and emits a retryable fatal event after consecutive unhealthy samples so the gateway's reconnect watcher rebuilds the adapter. You're trading setup simplicity for a liveness subsystem that exists purely because Discord's transport model is two independent failure domains.

💡 Where Discord earns its keep

The combination of voice channels + threads + reactions + streaming makes Discord the right surface for anything watchable — a live build session people can lurk in, an agent demonstrating a workflow, or a team channel where the agent's intermediate work is part of the value. The cleanup_progress option (auto-delete tool-progress bubbles after the final response lands) exists specifically so the channel stays clean for that use.

Discord in one line

Best for: team spaces, live sessions, voice-channel conversations, anything with multiple humans watching. Trade: the two privileged-intent toggles, the invite dance, and a liveness subsystem you have to know about — the most setup fragility in the group.

5. Slack: The Enterprise-Grade Path with the Longest Setup

Slack is the platform you pick when the agent has to live inside an organization's existing communication fabric — where the boss, the on-call rotation, and the vendor channel all already exist [Setup]. It's the most setup work of the five, for a reason: Slack's app model is the most granular.

  • Setup: Hermes generates a paste-in app manifest (hermes slack manifest --agent-view --write) that declares every built-in slash command, every OAuth scope, and every event subscription at once — which is the intended path and collapses most of the work. Manual setup, for reference, is: ~10 bot token scopes, an app-level token for Socket Mode, five event subscriptions, and the Messages Tab enabled in App Home. The docs are explicit that any scope or event change requires re-installing the app — the Install page shows a banner prompting you, but if you miss it, nothing changes and there's no error.
  • Transport: Socket Mode — outbound WebSocket, no public URL, no webhook endpoint exposed. For an agent with shell access, that's a real security property: one less attack surface, and it works from behind corporate firewalls and on a laptop with no inbound ports.
  • Security posture: official API, allowlist by Slack member ID (SLACK_ALLOWED_USERS). Deny-all default if unset, same as Discord. The two most commonly missed scopes have specific symptoms: without channels:history/groups:history the bot works in DMs but is deaf in channels; without files:read it chats but can't read uploaded attachments (including voice notes).
  • Capability notes: every Hermes command is a native Slack slash command with autocomplete. With the optional assistant:write scope, the agent controls its own "thinking…" status line instead of Slack's generic rotating placeholders. Threads, reactions, voice, and streaming all work.

⚠️ The Slack tax

The reinstall-everything-on-every-change model is the defining operational cost. It's also where "it works in DMs but not channels" comes from — the classic symptom of a missing message.channels event subscription plus an uninstalled app. The troubleshooting checklist is eight items long, and every one of them has failed a real setup. Budget time for it, or use the manifest path and never think about it again.

Slack in one line

Best for: production teams, orgs that live in Slack, anything where the agent must be reachable by colleagues without asking them to install a new app. Trade: the longest setup, the reinstall-every-change tax, and a scope/event matrix where one missed entry silently disables a whole surface.

6. WhatsApp: Meeting People Where They Live, with a String Attached

WhatsApp is the odd one out in the set, because its default integration is not an official API. Hermes' built-in bridge is based on Baileys and works by emulating a WhatsApp Web session — no Meta developer account, no Business verification, no public URL [Setup]. That's simultaneously its biggest advantage and its biggest caveat.

  • Setup: hermes whatsapp → pick mode (dedicated bot number or self-chat) → scan a QR code from your phone. The session persists to disk, survives restarts, and includes the encryption keys — treat ~/.hermes/platforms/whatsapp/session like a password (it grants full account access). For a bot number, the cheap options are a Google Voice number, a $5–15 prepaid SIM that sits in a drawer (make a call every 90 days to keep it active), or a VoIP number — some VoIP providers get blocked, so keep a spare.
  • Security posture: this is where it diverges from the other four. WhatsApp does not officially support third-party bots outside the Business API, and using the bridge carries a small but real risk of account restrictions. The mitigations are the standard ones: dedicated number (never your personal one), conversational use only, no outbound automation to people who haven't messaged first. Note the access control too: without WHATSAPP_ALLOWED_USERS the gateway denies everyone, and unauthorized DMs get a pairing-code reply unless you set unauthorized_dm_behavior: ignore — which is the right choice for a private number.
  • Capability notes: no TTS voice replies on the Baileys bridge (the Cloud API variant has them), but incoming voice notes are transcribed (local faster-whisper, Groq, or OpenAI Whisper). Streaming works via message editing, long responses auto-split at 4,096 characters, and standard Markdown gets converted to WhatsApp's native formatting. The bridge uniquely adds native polls (including the agent's clarify questions rendered as tappable options) and location pins.
  • Reliability: the honest story is "it's a moving target." WhatsApp periodically changes its Web protocol, which can break third-party bridges; when it does, you pull the latest Hermes and re-pair. Devices also get unlinked after long inactivity. Temporary network blips are handled automatically; protocol drift is not.

💡 The two-adapter escape hatch

If the ban-risk profile is unacceptable but you want WhatsApp reach, the WhatsApp Business Cloud API is the official Meta-supported path: no account-ban risk, voice included — but it requires a Meta Business account and a public webhook URL, and it drops the native polls/location extras. Hermes runs both adapters, and they can run in parallel on different phone numbers if a use case genuinely needs both postures.

WhatsApp in one line

Best for: reaching people who will never install Telegram or join a Discord; consumer-adjacent and non-technical users; field and mobile-first workflows. Trade: unofficial-protocol ban risk, no TTS on the bridge, and protocol-drift maintenance you didn't have on any other platform.

7. The Long Tail: Signal, SMS, Email, iMessage, ntfy, and Teams

The big five get the deep treatment, but the long tail exists precisely for the gaps the big five leave. The same gateway process adds any of these without changing the agent:

  • Signal — runs through a signal-cli daemon. No voice, no streaming, but the strongest privacy posture of any option in this post, which is the whole point of choosing it [Setup].
  • SMS via Twilio — the last-resort reach channel. No media, no voice, nothing — but it works on a device that's out of battery and out of data and in a pocket, and it requires the recipient to install nothing. The right tool for "the agent must be able to reach me always" [Setup].
  • Email via IMAP/SMTP — the async surface. Threads work (a thread is a conversation), images and files work, but there are no real-time affordances. Good as a notification sink and for batch-style interactions; wrong for interactive work [Setup].
  • iMessage via BlueBubbles or Photon — the macOS-native option. BlueBubbles (server-based) covers images/files but no voice; Photon (native) adds voice and reactions. Only makes sense if your whole life is on Apple devices and you refuse to carry a second messaging app [BlueBubbles] [Photon].
  • ntfy — the self-hosted push-notification broker. The pattern: the agent's cron jobs, monitors, and background runs push one-way notifications to your phone through a server you own. No conversation, no dependency on a platform company — just reliable signal delivery with your data staying on your infrastructure [Setup].
  • Microsoft Teams — the Slack alternative for organizations standardized on M365. Images and threads, no voice in the base adapter; Teams Meetings adds a Graph-webhook-driven meeting-summary pipeline [Setup].
  • The rest — Matrix (full capability, self-hostable, E2E-encrypted — the "Slack that's yours" option), Mattermost (self-hosted Slack-style, full capability), plus the Asian-market stack (LINE, DingTalk, Feishu/Lark, WeCom, QQ, Weixin), Google Chat, SimpleX, IRC, Buzz, and Open WebUI via the API server. Same gateway, same session model, same command set [Overview].

The long tail's real role: channel diversification

The underrated pattern isn't "pick one." It's primary + fallback + sink: a full-capability primary surface (Telegram or Discord) for conversation, a reach channel (SMS or ntfy) for the moments when the primary isn't, and an async sink (Email) for scheduled reports and batch results. The cron scheduler in the gateway already knows how to deliver to any configured platform, so a scheduled research digest can land in your email at 7am while live operation happens on your phone — one agent, zero extra infrastructure.

8. Decision Guide: Matching the Surface to the Use

Collapsed to a decision procedure. The honest version, based on the tradeoffs above:

Use casePickWhy — and what you're paying
Personal agent, mobile-first, voice memos, minimal setup Telegram Five-minute setup, official API, full capability, voice in both directions. You pay: privacy-mode detour for groups, no reactions.
Team channel, live/watchable sessions, voice conversations Discord Only platform with voice channels + threads + reactions + streaming. You pay: two privileged-intent toggles, the invite dance, liveness config, per-user session isolation to untangle.
Org that lives in Slack; agent must reach colleagues Slack Native slash commands, Socket Mode (no public endpoint), role-free member-ID auth. You pay: the longest setup and the reinstall-on-every-change tax.
Reaching people who are only on WhatsApp WhatsApp (Baileys), or Cloud API if ban risk is disqualifying Meets them where they live; QR-pair setup; polls and location pins. You pay: unofficial-protocol ban risk and drift maintenance — or, on the Cloud API, a Meta Business account plus public webhook and loss of native extras.
Critical reach: agent must get through even when apps won't SMS / ntfy as a fallback layer Nothing to install, nothing to update, nothing to trust at protocol level. You pay: zero capability (SMS) — it's a pipe, not a conversation.
Privacy-first personal channel Signal (or self-hosted Matrix for team use) Strongest privacy posture in the set; Matrix adds E2E team rooms you control. You pay: no streaming on Signal, and the daemon or server you now run.
Apple-only life, zero new apps iMessage via BlueBubbles/Photon Native surface, no new app on the phone. You pay: a Mac-based server (BlueBubbles) and a capability ceiling below the big four.

Three cross-cutting rules that hold regardless of which surface you pick:

  • Start with one home channel, then add. Run /sethome on the chat that should receive cron deliveries and scheduled results. Adding a second platform later is additive — the gateway connects them in parallel and sessions stay separate. Starting with five is how people end up debugging five stacks when one would have sufficed.
  • Allowlist before you expose. Every platform defaults to deny-all with no allowlist — the right default, and the reason "it just works" can't be the goal. Set *_ALLOWED_USERS (or roles, for Discord) before inviting anyone, and remember authorized users have full agent capabilities, including tool use and system access. Treat the allowlist as your auth system, not a detail.
  • Match the surface to the recipient, not your preference. The platform you enjoy is irrelevant if the person who needs the agent's output lives somewhere else. An agent that files a report into a Discord server nobody on the team has joined is a toy; the same agent pushing to the Slack channel where the decision actually happens is infrastructure. The capability matrix tells you what's possible; the recipient tells you what's necessary.

The one-paragraph version: the gateway makes platform choice a UX question, not an architecture question — one agent, one session store, N doors. Pick Telegram for the fastest personal path with full capability, Discord where the work is live and watchable, Slack where the organization already lives, WhatsApp where your users are and you accept the unofficial-protocol string, and a reach channel (SMS/ntfy) plus an async sink (Email) so the agent can still get through when the door is closed. Every choice trades setup friction, security posture, and maintenance burden against capability — and the only wrong answer is picking by brand familiarity instead of by who's on the other end.

References

Research by Michel Laclé · ThinkSmart.Life · September 2026