How the SDK integrates with Hermes Agent
The SDK treats Hermes Agent as an external dependency and needs no changes to Hermes. Everything goes through surfaces Hermes already publishes for third-party platforms. This page lists those surfaces, explains how each one is used, and states plainly where the SDK relies on behavior that is not a formal API. It ends with the few small upstream changes that would remove those workarounds.
The integration point: a platform plugin
Hermes's gateway runs "platform adapters": Telegram, Slack and about twenty others. A directory plugin with kind: platform can register a new adapter through ctx.register_platform(...) without touching core code; see gateway/platforms/ADDING_A_PLATFORM.md in Hermes. The SDK's plugin/ directory is such a plugin:
| Plugin file | Role |
|---|---|
plugin.yaml |
Manifest: kind: platform, optional env vars surfaced in hermes config |
__init__.py |
register(ctx): platform entry, agent tools, hermes gadget CLI |
adapter.py |
GadgetAdapter(BasePlatformAdapter): the Hermes-facing side |
hub.py |
WebSocket server, auth, audio, images, actions (no Hermes imports) |
tools.py |
gadget_devices, gadget_display, gadget_action |
store.py, protocol.py, audio.py, textfmt.py, imaging.py |
Support modules (no Hermes imports) |
Each device is a direct-message chat whose chat_id and user_id are the device id. Everything Hermes does for a DM on any platform then applies unchanged: sessions and memory, /new, /stop, /voice, model routing, approvals and delivery.
Configuration
# ~/.hermes/config.yaml
plugins:
enabled: [gadget] # or: hermes plugins enable gadget
platforms:
gadget:
enabled: true
extra:
host: 0.0.0.0 # interface devices connect to
port: 8765
path: /gadget
speak_replies: true # spoken replies for devices with a speaker
auto_home: true # first approved device becomes the gadget home channel
unauthorized_dm_behavior: pair
# heartbeat_s: 20
# max_utterance_s: 60
# tls_cert: /path/cert.pem # serve wss://
# tls_key: /path/key.pem
Secrets go in ~/.hermes/.env, following Hermes's rule that .env is only for secrets:
GADGET_ACCESS_TOKEN: optional shared token every device must present.GADGET_ALLOWED_USERS: comma-separated device ids admitted without pairing.GADGET_ALLOW_ALL_USERS=true: admit any device (development only).
If you set
GADGET_ALLOWED_USERS(orGATEWAY_ALLOWED_USERS), Hermes switches unknown senders from "pair" to "ignore". Keepunauthorized_dm_behavior: pairin the platform'sextraso new devices still receive pairing codes.
Hermes surfaces the SDK uses
| Need | Hermes surface | Notes |
|---|---|---|
| Register a transport | ctx.register_platform(...) → PlatformEntry |
allowed_users_env, allow_all_env, platform_hint, max_message_length, parse_target_ref_fn |
| Setup wizard | PlatformEntry.setup_fn |
hermes gateway setup lists Hermes Gadget and runs its step: enable the platform, pick the port (written with hermes config set's writer), print the device URL and the installer link. The wizard then offers the gateway restart |
| Inbound messages | BasePlatformAdapter.handle_message(MessageEvent) |
text → MessageType.TEXT; voice → MessageType.VOICE with a WAV in media_urls |
| Store voice audio | cache_audio_from_bytes_async |
The gateway's own audio cache |
| Speech-to-text | Gateway STT for VOICE messages (stt.* config) |
Hermes transcribes; the device never runs ASR |
| Transcript echo | Gateway STT echo (stt_echo_transcripts) |
Forwarded to the device as transcript |
| Per-device context | MessageEvent.channel_prompt |
Stable text: screen size, speaker, action names. The ephemeral prompt is pinned per session, so prompt caching holds |
| Turn boundaries | on_processing_start / on_processing_complete |
Become turn.start / turn.end |
| Live status | supports_status_text = True + set_status_text() |
Becomes status frames |
| Streaming text | supports_draft_streaming() + send_draft(); edit_message() |
Used when streaming is enabled; becomes reply.delta |
| Final replies | send() |
Markdown stripped and folded to ASCII for small screens |
| Whole-file TTS | play_tts() / send_voice() |
WAV decoded in Python; other formats through ffmpeg; resampled to the device rate |
| Streaming TTS | supports_streaming_tts / begin_ / write_ / finish_ / abort_streaming_tts |
PCM chunks forwarded as they are generated, for TTS providers Hermes can stream |
| Images from the agent | send_image_file() / send_image() |
Converted to RGB565 to fit the screen (needs Pillow) |
Confirmations (/new, /undo, model switches) |
send_slash_confirm() + tools.slash_confirm.resolve() |
Shown as a prompt; TALK resolves once, CANCEL cancel. The handler's reply goes back with send() |
| Dangerous-command approvals | _send_exec_approval_prompt(ExecApprovalPrompt) + tools.approval.resolve_gateway_approval() |
Shown as a prompt; TALK resolves once, CANCEL deny. On timeout Hermes calls edit_message() on the prompt, which withdraws it |
| Pairing | Hermes DM pairing (gateway/pairing.py, hermes pairing approve; hermes gadget pair approves through the same PairingStore) |
See below |
| Authorization state | BasePlatformAdapter._is_sender_authorized() |
The runner-installed check; used to mirror approval and revocation to the device |
| Agent tools | ctx.register_tool(toolset="gadget") |
Part of the implicit hermes-gadget toolset; also usable from other chats |
| Tool session context | gateway.session_context.get_session_env |
HERMES_SESSION_PLATFORM / HERMES_SESSION_CHAT_ID pick the default device |
| Durable state | plugins.plugin_storage.plugin_data_dir("gadget") |
Enrolled device keys, pending pairing codes, and firmware staged for updates (updates/) |
| Host CLI | ctx.register_cli_command("gadget", ...) |
hermes gadget devices / forget / info / pair / update. pair waits for a device to show a code and approves it after asking. update stages an image in the plugin's data directory (a file, or with --latest the newest release's image for the device's board, checked against the release manifest); the gateway's adapter installs it once the device is online and writes progress back for the command |
| Send and cron targets | parse_target_ref_fn |
gadget:hg-0123456789abcdef works with send_message and cron delivery |
Voice round trip
- Device: push-to-talk streams PCM16 at 16 kHz to the hub, which assembles a WAV.
- Adapter: caches the WAV with
cache_audio_from_bytes_async, then callshandle_message(MessageEvent(VOICE, media_urls=[wav])). - Gateway:
_enrich_message_with_transcriptionruns the configured STT provider. The transcript is echoed (device shows "> what you said"), then the agent turn runs. - Reply audio:
- When the TTS provider streams,
StreamingTTSConsumer→write_streaming_tts→ PCM to the device while text is still being generated. - Otherwise whole-file auto-TTS →
play_tts(file)→ decode → resample → paced PCM.
- When the TTS provider streams,
- Reply text:
send(text)→replyframe. Thenturn.end.
Because a gadget is voice-first, GadgetAdapter answers True from _should_auto_tts_for_chat for devices that declare a speaker. The gateway then reads every reply aloud for that chat, both typed and spoken input. /voice off in the device's chat, or speak_replies: false, turns it off.
Pairing
The SDK reuses Hermes's DM pairing instead of inventing its own:
- When an unknown device connects, the adapter dispatches a harmless
/statusmessage from it. - Hermes's unauthorized-sender path answers with a pairing code. The adapter recognizes the
hermes pairing approve <platform> <CODE>command in that reply and sends the device apairingframe. - The device shows the code. The owner runs
hermes gadget pairon the Hermes host, which finds the waiting device and approves its code after asking, or runs thehermes pairing approvecommand the device shows. - The adapter polls the runner's authorization check every 2 s, so approval reaches the device as
pairedwithout a reconnect. Revocation (hermes pairing revoke gadget <id>) reaches it asunpaired.
Home channel. Hermes opens each new session on a platform without a home channel with a "type /sethome" notice, which a device without a keyboard can't act on. So when a device is approved and the gadget platform has no home channel (platforms.gadget.home_channel or GADGET_HOME_CHANNEL), the adapter makes that device the home channel, as /sethome would, and saves it to config.yaml. Cron results and messages sent to gadget then go to that device. Run /sethome from another device to move it, or set auto_home: false to turn this off.
The pairing code authorizes the chat in Hermes. The device key, enrolled on first contact and proven by HMAC afterwards, keeps another device from impersonating an approved one.
Confirmations
Hermes asks before destructive commands (/new, /undo), costly model switches and dangerous shell commands. Chat platforms render those as buttons; the gadget renders them as a yes/no prompt (see protocol.md):
- Holding CANCEL sends
session.new, which the adapter turns into/new. Hermes then asks "Confirm /new". The hold already was a deliberate confirmation, so when the confirmation fornewarrives within 15 s of the device's own request, the adapter resolves it withonceand forwards the reset reply. A/newtyped or spoken any other way still asks. - Everything else goes to the device as a question. The adapter keeps only the header and detail of Hermes's message, dropping the choice list and typed-reply instructions meant for chat apps.
- One at a time: questions queue per device in arrival order, which matches Hermes's first-in, first-out approval queue. A question asked while the device was offline is shown when it reconnects.
- Elsewhere: answering with
/approveor/denyfrom another chat still works; the device's copy then reports "That question has expired" if pressed.
Agent tools and Hermes tool search
Hermes keeps the core tool schema narrow. Plugin tools are therefore deferred behind its tool-search bridge (tool_search / tool_describe / tool_call), and their names and descriptions appear in the bridge's manifest. A model invokes gadget_action through tool_call, and the end-to-end test exercises exactly that. The per-device context names the available actions so the model knows they exist.
Making replies fast
A voice turn is speech-to-text, the model turn, then text-to-speech. Most of the wait is usually the model, and the settings that matter live in Hermes:
| Setting | Effect |
|---|---|
/reasoning low (or none), typed on the device |
Session-only reasoning override for the gadget's chat. A deep-reasoning default (agent.reasoning_effort: high) can spend 10–20 s thinking about a one-line answer. Other chats keep their setting |
streaming.enabled: true |
Streams reply text to messaging platforms as it is generated, so words appear on the device within a second or two. Hermes's per-platform display.platforms.<name>.streaming can only narrow this global switch, so on a shared gateway set the global on and switch the other platforms off. The Desktop app streams on its own and is not affected |
| A TTS provider Hermes can stream | Speech starts after the first sentence instead of after the whole reply. While whole-file TTS synthesizes, the reply text waits for it, so a working streaming provider also shows the text sooner. Check the gateway log for streaming TTS clause failed; it usually means a bad API key |
stt.provider |
Cloud STT takes about 2–3 s for a short clip. Local faster-whisper can be quicker on a fast CPU |
Where the SDK leans on behavior that is not a formal API
These work on current Hermes and are covered by tests/test_adapter_hermes.py and tests/test_gateway_e2e.py, but they are conventions rather than documented contracts:
- Spoken replies by default. The adapter overrides
BasePlatformAdapter._should_auto_tts_for_chat, a private method the runner consults. Hermes's own docs endorse overriding private hooks such as_keep_typingfor platform UX, but this one is not listed. - Pairing trigger and code capture. Starting pairing takes a synthetic message, and the code is extracted from the human-readable reply by matching the approve command.
- Approval detection polls
_is_sender_authorizedevery 2 s, because there is no "pairing approved" event. - Transcript echo detection matches the echo format: microphone emoji plus quoted text. If the format changes, the transcript is shown as an ordinary reply; nothing breaks.
- Self-confirming
/newreadstools.slash_confirm.get_pending()to check that the pending confirmation is for thenewcommand. It relies on the runner registering a confirmation before it callssend_slash_confirm, which Hermes does on purpose so that fast button presses can't race it. If that changes, the device simply asks. - Question text is cut down from Hermes's chat-formatted confirmation by dropping the bullet-list and italic paragraphs. A different layout only means a longer question.
Suggested upstream changes (optional; nothing here blocks the SDK)
Each would replace a workaround above with a small, generic hook that every platform plugin could use. None of them is specific to gadgets. The first four are open pull requests on Hermes Agent; until they merge, the SDK keeps the workarounds.
Pairing API for adapters: NousResearch/hermes-agent#132448, open.
- What:
BasePlatformAdapter.request_pairing(source) -> PairingOffer | Nonereturns the code the DM path would send (code,command,expires_in), behind the same gates and limits, without sending anything to the chat.on_pairing_changed(user_id, approved)is called once for each approval and revocation, from one watcher in the runner.
- Why: devices, kiosks and any screen-first platform could show codes natively, without synthetic messages, text parsing or polling.
- SDK change afterwards:
on_readycallsrequest_pairing, and_watch_pairingis deleted.
- What:
Voice-first platforms: NousResearch/hermes-agent#132441, open.
- What: two adapter class attributes.
speaks_replies_by_default = Truespeaks replies withoutvoice.auto_tts, and/voice offstill silences a chat.supports_voice_replies = Falsekeeps replies text, replacing the name checks for A2A. - Why: voice-native platforms (gadgets, phone bridges) want spoken replies by default without overriding a private method.
- SDK change afterwards: the
_should_auto_tts_for_chatoverride is removed, andspeak_repliesmaps onto the attribute.
- What: two adapter class attributes.
Transcript echoes for adapters: NousResearch/hermes-agent#132435, open.
- What:
BasePlatformAdapter.send_transcript_echo(chat_id, transcript, metadata=None). By default it makes the samesend()call as today; an adapter that overrides it gets the raw transcript. - Why: adapters could render transcripts distinctly without matching localized text.
- SDK change afterwards: the adapter overrides it, and the echo-text matching is deleted.
- What:
Direct toolsets for a platform's own sessions: NousResearch/hermes-agent#132449, open.
- What:
PlatformEntry.direct_toolsets. A toolset listed there is direct only in sessions on that platform. - Why: a platform-specific tool (here, controlling the device the user is holding) is part of that surface, and a search round trip costs latency on voice turns. Sessions on other platforms keep deferring it, so the core schema stays narrow.
- SDK change afterwards:
register_platform(..., direct_toolsets=("gadget",)).
- What:
Per-platform reasoning default.
- What: a
platforms.<name>.reasoning_effort, resolved after the session override and before the per-model and global settings. - Why: voice devices want fast, shallow turns while desktop chats keep deep reasoning. Today that needs a
/reasoningper session. - Upstream status (October 2026): an open pull request, NousResearch/hermes-agent#65576, adds
reasoning_efforttochannel_overrides, which is per chat (per device). A platform-wide default would be a small follow-up to it rather than a separate change. - SDK change afterwards: document a
platforms.gadget.reasoning_effort: lowdefault.
- What: a
Running against a profile or a multiplexed gateway
The adapter binds its own port, so give each profile that serves gadgets a different platforms.gadget.extra.port.
Device keys and pending pairing codes live in the plugin data directory. It is resolved with plugin_data_dir("gadget") when the adapter connects, so it follows whichever Hermes home the adapter is started under.
The agent tools look up the adapter for the session's profile (HERMES_SESSION_PROFILE). Multi-profile setups have only been exercised in unit tests so far.