Don’t want to host it yourself? DeepInfra can run Hermes for you — see Hosted Agents: Hermes-Agent.
provider: deepinfra, export your key as DEEPINFRA_API_KEY, and pick any LLM from our catalog. Hermes reads the model list, context lengths and pricing live from the DeepInfra catalog, so there is no base URL or context length to configure.
Configure ~/.hermes/config.yaml
~/.hermes/.env:
- The DeepInfra model id passes through verbatim in
default— no reformatting needed. - Reasoning is controlled through DeepInfra’s
reasoning_effortfield, soagent.reasoning_effort,/reasoning <level>and--reasoningwork in both directions: an effort turns thinking on for models that default off (DeepSeek-V4.x), and/reasoning noneturns it off for models that default on. - Hermes needs roughly 64k of context for agent functionality. Check a model’s window and pricing via
/v1/openai/models?filter=with_meta&sort_by=hermes. - On a Hermes release older than the built-in provider, use
provider: customwithbase_url: https://api.deepinfra.com/v1/openai,api_key: ${DEEPINFRA_API_KEY}and an explicitcontext_length.
Run it
Learn more
AI Providers
Hermes’ provider overview, including the DeepInfra section.
Configuration
The full
config.yaml reference.Chat Completions
DeepInfra’s OpenAI-compatible API.
NVIDIA OpenShell
Run Hermes in a sandbox whose only allowed destination is DeepInfra.
Tested with Hermes Agent v0.21.5 (2026.9.24).