Choosing an LLM Backend for a Self-Hosted Hermes Agent: Cost, Security, and Policy
Who this is forDevelopers and automation builders who run a self-hosted AI agent and need a cost-effective LLM backend that stays within provider policy.
If you run a self-hosted AI agent around the clock, the LLM backend you choose determines three things at once: what you pay, where your prompts end up, and whether the provider’s terms allow the automation at all. This article walks through a real decision. A Mac mini runs Hermes (NousResearch’s hermes-agent), which powers a Discord bot named Jarvis2. Its main backend was the openai-codex provider, which uses a ChatGPT subscription to reach gpt-5.5. When the authentication for that route expired, the setup needed a replacement. The result is a three-axis evaluation covering cost, data security, and policy. You will get the decision rules, the supporting tables, the pitfalls specific to Hermes, and a reusable way to judge any new low-cost model.
Summary
The security of a self-hosted agent’s LLM backend is decided by the jurisdiction of the servers that data passes through, not by the model’s country of origin. A Chinese-origin open-weight model served from US inference hosts can keep its low cost while removing the jurisdictional risk. By contrast, using a Claude subscription as a backend sits in a policy gray area, so it cannot serve as the main backend for an automated deployment.
Key diagram

Key data
Background
The starting point was a replacement problem. The Mac mini Hermes instance had relied on openai-codex, which reaches gpt-5.5 through a ChatGPT subscription, and its authentication had expired. The selection rested on three principles, applied in order. The first was cost-effectiveness. The second was security. The third, which I treat as a fallback rule (2-1), was that if a security problem arises, security grades should be separated by channel, so that a sensitive channel can run on a stricter backend while a lower-risk channel keeps a cheaper one.
The research method was deliberately adversarial. It used a multi-agent workflow with 15 agents working across five axes in parallel. The core claims, ten in total, were checked against official primary sources, and that verification produced two rebuttals. Those rebuttals changed some of the conclusions, which is why the sections below include corrections as well as findings.
Claude -p and Claude subscription policy verdict
Policy was the first axis to resolve, because it determines which backends are even eligible for an always-on agent. The table below sorts each usage pattern by its verdict.
| Setup | Verdict |
|---|---|
claude -p programmatic invocation itself |
Allowed — officially documented for the Agent SDK and headless mode |
| Another tool directly using a subscription OAuth token | Prohibited — codified in February 2026; enforcement blocked 4 of 4 third-party harnesses |
| Hermes built-in anthropic provider (reuses Claude Code credentials) | Gray area — after the 4 of 4 blocks, a June 16 rollback kept it “for now”; however, a case of actual extra-usage charges (hermes issue #47260) is ongoing |
| Single user automating their own subscription | No explicit prohibition — whether it counts as “ordinary use” is unresolved |
The conclusion from this table is that the only clearly compliant path is pay-as-you-go billing through ANTHROPIC_API_KEY. A community claim says that starting June 15, 2026, Agent SDK and claude -p usage was split off from subscription limits into separate credits. I could not verify it, and I rate my confidence in it as medium, so it should not drive a decision.
The gray-area row deserves attention because it looks usable. It worked for a while, then the enforcement picture changed, and actual charges started appearing for some users. A backend that can flip from free to billed, or from allowed to blocked, is a poor foundation for an agent that runs unattended.
Structural security risks of Chinese services (company-independent, common legal basis)
The second axis concerns legal structure rather than any single company. The table lists three Chinese laws and what each implies for data sent to a service that falls under them.
| Law | Implication |
|---|---|
| National Intelligence Law (2017), Article 7 | Obligates all Chinese organizations and citizens to cooperate with intelligence work; contractual privacy commitments cannot offset this |
| Data Security Law (2021) and Cybersecurity Law (2017) | Government access rights; no external means to audit or verify how data is actually handled |
| Anti-Espionage Law (2023 revision) | The scope of what may be collected cannot be predicted |
The key point is that these risks come from the legal environment, so they apply regardless of how a particular vendor describes its practices. A privacy agreement is a contract between two parties, and it cannot override a statute that binds one of them.
Several company-level details follow from this. Z.ai’s Singapore entity (JINGSHENG HENGXING) is often presented as a barrier against obligations to its Chinese parent company, but no independent verification supports that claim. The API data processing agreement says prompts are not retained, but that statement is self-attestation, meaning the company asserts it without outside audit. Zhipu AI was added to the US Commerce Department’s Entity List on January 16, 2025, as a standard listing. An earlier claim that it was designated under Footnote 4 was a misreport. That designation applied to Sophgo-affiliated entities added the same day, and the verification process corrected the error.
The most useful distinction in this section separates open-weight models from hosted services. Models such as GLM, MiniMax, Kimi, and DeepSeek publish their weights, which means US inference providers such as DeepInfra and Fireworks can serve them on their own servers. On that path, traffic never crosses Chinese jurisdiction. OpenRouter works differently. It is a router, not a server that runs the model itself. It forwards requests to whichever provider it selects, so the security grade only holds if the provider allowlist is pinned to US-based providers.
Measured prices (July 2026, per 1M tokens, input / output)
Cost is the first selection principle, and the price table shows how quickly the field moves. Prices are per million tokens, input first, then output.
| Service | Price | Notes |
|---|---|---|
| Gemini 2.5 Flash-Lite | $0.10 / $0.40 | Lowest current price |
| Gemini 2.5 Flash | $0.30 / $2.50 | Shutdown on October 16, 2026 |
| Gemini 3.5 Flash | $1.50 / $9.00 | Successor at 5x the price; the “Flash means cheap” pattern breaks down |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 | Successor in effect to 2.5 Flash’s price range |
| xAI Grok 4.1 Fast | $0.20 / $0.50 | Lowest-priced non-Chinese direct option; 2M context |
| MiniMax M2.7 (OpenRouter) | $0.24 / $0.96 | Can be routed through US hosting |
| Z.ai GLM-4.7 API | $0.60 / $2.20 | GLM Coding Plan Lite $18/month |
| Claude Haiku 4.5 | $1 / $5 | Sonnet 5 is $2 / $10 through August 31, then $3 / $15 |
| ChatGPT Plus | $20/month flat | Includes codex |
Two details matter more than the raw numbers. First, the Gemini free tier uses submitted prompts for model training and human review, and Google itself says sensitive information should not be submitted to it. A “sensitive data OK” classification therefore holds only for paid usage with billing enabled. Second, the verification rebuttal found that the free-tier limit for Gemini 2.5 Pro had already been reduced to 2 RPM and 50 RPD, not the 5 RPM and 100 RPD that older figures suggested. Since December 2025, the official limits table has been removed from the documentation altogether, so any limit figure older than that should be treated as stale.
hermes-agent connectivity and community (checked against source code)
Connectivity was the last axis. The hermes-agent source defines 33 native providers, including gemini, zai, anthropic, xai, openrouter, nous, and deepinfra. Every candidate in this evaluation is therefore reachable without a proxy. The source has no backend that works as a local CLI subprocess, with one exception, copilot.
Connecting GLM directly to Hermes exposes two Hermes-specific traps. The first is that the z.ai coding endpoint rejects requests whose system prompt contains the string “Hermes Agent,” returning what looks like a 429 rate-limit error that is actually a fake rejection (#60118). The second is an endpoint auto-detection bug that bills the wallet instead of drawing from the subscription quota (#42536). Both problems occur only in combination with Hermes, so testing a provider outside Hermes will not reveal them.
Codex token expiry is a recurring pain point in the community (#65346, #63413). The most important detail is that the failure happens even when the subscription is still active, because the token refresh path breaks. The subscription staying alive does not guarantee that the agent can keep running.
Finally, the community has no dominant backend. Codex OAuth, low-cost Chinese plans, OpenRouter, and local Qwen all coexist, and no testimonies report Gemini as a primary backend.
Insights
-
The unit of security evaluation is the data path, not the model. A model is not rejected because it is Chinese. It is rejected when prompts travel to servers under Chinese jurisdiction. The most reusable idea from this research is that the same open-weight model can change jurisdictions simply by changing where it is served. Any future low-cost model can be judged with two questions: are the weights public, and where is the serving server located?
-
A router does not inherit the security grade of the providers behind it. The question is not whether OpenRouter is a US company, but which provider a request is routed to. Without an allowlist pinned to specific providers, the security grade cannot be determined at all.
-
Any path that reuses a flat-rate subscription as if it were an API carries policy volatility. Anthropic’s rules changed within half a year, moving from blocking to rollback to a shift toward billing. OpenAI’s unofficial backend-api route also carries similar risk. For an always-on automation, the main backend should sit on an explicitly permitted interface, such as a pay-as-you-go API key, because that choice keeps maintenance costs low.
-
Generational turnover in low-cost model lineups can bring price increases. Gemini 2.5 Flash at $0.30 per million input tokens became Gemini 3.5 Flash at $1.50, which is five times the price. An automation pipeline’s default model should therefore be tracked together with its shutdown schedule and the pricing of its successor.
Bottom line
For an always-on Hermes agent, the evidence supports a clear rule. Judge each backend by the jurisdiction of the servers that handle its data, not by the nationality of the model. Open-weight models served by US inference hosts can keep costs low without the jurisdictional exposure of a Chinese-operated service, provided any router in the path is locked to those US providers. Reusing a Claude subscription should not be the main backend, because its policy status is unresolved and has already shifted once. A pay-as-you-go API key is the compliant choice for unattended use. Track shutdown dates and successor prices alongside the cost of each model, since the cheapest tier can change within months.
Sources
- hermes-agent source (providers, credential reuse, fallback): https://github.com/NousResearch/hermes-agent
- Anthropic Claude Code Legal & Compliance: https://code.claude.com/docs/en/legal-and-compliance
- Anthropic Consumer Terms (Automated Access clause): https://www.anthropic.com/legal/consumer-terms
- Claude Agent SDK and headless documentation: https://code.claude.com/docs/en/agent-sdk/overview · https://code.claude.com/docs/en/headless
- Anthropic consumer data policy change (August 2025): https://www.anthropic.com/news/updates-to-our-consumer-terms
- Gemini API pricing: https://ai.google.dev/gemini-api/docs/pricing
- Gemini API terms (free and paid data handling): https://ai.google.dev/gemini-api/terms
- Verification of Gemini free-tier limit reduction (Wayback Machine, December 1, 2025): http://web.archive.org/web/20251201213053/https://ai.google.dev/gemini-api/docs/rate-limits
- Z.ai pricing: https://docs.z.ai/guides/overview/pricing · Coding Plan: https://docs.z.ai/devpack/overview
- Z.ai privacy policy and API data processing agreement: https://docs.z.ai/legal-agreement/privacy-policy
- Zhipu Entity List addition (90 FR 4617): https://www.federalregister.gov/documents/2025/01/16/2025-00704/addition-of-entities-to-and-revision-of-entry-on-the-entity-list
- Claude API pricing: https://platform.claude.com/docs/en/about-claude/pricing · subscriptions: https://claude.com/pricing
- xAI Grok pricing (secondary aggregator): https://www.aipricing.guru/xai-pricing/
- MiniMax M2.7 pricing: https://pricepertoken.com/pricing-page/model/minimax-minimax-m2.7
- hermes community threads: https://www.reddit.com/r/LocalLLaMA/comments/1ro9lph/anybody_who_tried_hermesagent/ · https://news.ycombinator.com/item?id=48419000
- hermes issues: GLM branded-string rejection #60118 · wallet billing #42536 · codex token expiry #65346 and #63413 · Claude extra-usage billing #47260
- Full report (HTML): report-hermes-llm-backend.html
Frequently asked questions
- Does the security of an LLM backend depend on the model's country of origin?
- No. The security question is the jurisdiction of the servers that process the data. An open-weight Chinese model served by US inference hosts keeps its traffic out of Chinese jurisdiction, while sending the same model's traffic through a Chinese-operated server does not.
- Can a self-hosted Hermes agent use a Claude subscription as its main backend?
- The note treats reusing Claude subscription credentials as a policy gray area, so it should not be the main backend for an always-on agent. The compliant route it identifies is a pay-as-you-go API key set through ANTHROPIC_API_KEY.
BuildnWrite helps teams build AI agents that keep running. About BuildnWrite ›