Files
Tony0410 77250ea8c9 incident: SFTPGo outage 2026-09-02 — CT299 tailscaled death + stale lxc-attach recovery
- New incident page documenting CT299 SFTPGo outage (02:18 AWST)
- Root cause: swap exhaustion (43Gi, 100%) killed tailscaled; 5 stale lxc-attach PIDs blocked pct start
- Resolution: killed PIDs 3037951/3334928/3352203/3402511/4004148, pct stop/start 299
- Updated current-state.md with Recent Changes (2026-09-02), bumped updated to 2026-09-04
- Updated log.md with incident entry
- Updated index.md last-updated date

No secrets written. Verified end-to-end: tailscale direct, SFTPGo WebAdmin HTTP 401, SFTP banner SSH-2.0-SFTPGo_2.7.0.
2026-09-04 00:05:32 +08:00

88 lines
5.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: Model Providers
type: system
status: active
created: 2026-07-22
updated: 2026-08-20
verified_on: 2026-08-20
confidence: high
tags: [model, provider, inference, routing]
sources: [live-probe, config, web-search]
---
# Model Providers
## Purpose
Inference providers available to Hermes, their routing paths, and notable model updates.
## Current Provider Stack
| Provider | URL | Role |
|---|---|---|
| `custom:litellm` | `http://litellm:4000/v1` | Main inference (slower path, 626s/call) |
| `custom:omni` | `http://omni:20128/v1` | Auxiliary roles (vision, skills, approval, mcp, web_extract, compression) — factory-reset 2026-08-14 |
| `opencode-zen` | (internal) | Free-tier model access for cron jobs |
| `charm-hyper` | `hyper.charm.land/v1` | Coding-focused inference (Charmbracelet) — passthrough models, free tier 100 credits/mo |
| Guanaco | `ollama:11434` | LLM proxy/router on CT205 (Ollama stack) |
## Default Model
- `resilient-free` — resolves through OmniRoute (changed from `auto/best-free` on 2026-08-14 after factory reset)
- Previous default `auto/best-free` was misconfigured — had no combo definition and routed to paid `gpt-5.6-sol`
- `freellmapi-auto` (LiteLLM path) still available but slower
## Notable Model Updates
### HyperCharm (Charmbracelet) — Provider Research (2026-08-20)
- **What:** Hyper (hyper.charm.land) is Charmbracelet's LLM inference service, purpose-built for agentic coding models
- **OmniRoute integration:** Already merged — PR #5961 (2026-07-03) added as `charm-hyper` provider with passthrough models from `hyper.charm.land/v1/models`
- **Free plan:** 100 Hypercredits/month (~$5 value), no rollover — use it or lose it monthly
- **Credit mechanics:** 1 Hypercredit = $0.05; balance queryable via `GET /v1/credits` (returns `{"balance": 100}`)
- **Catalog pricing (cheap models):** deepseek-v4-flash $0.20/$0.40 per 1M in/out, qwen3.7-flash $0.20/$0.80, gemma-4-26b $0.12/$0.42
- **No config changes made** — research only; `charm-hyper` is available but not wired as a Hermes default
- **Status:** Candidate for free-tier OmniRoute combo; passthrough model import means zero maintenance
### GPT-5.6 Luna — Model Research (2026-08-14)
- **Community consensus:** Luna is cleaner/more polished; DeepSeek V4 Flash is more persistent/cheaper
- **Luna strengths:** visual/UI work, final code review, bounded high-stakes tasks, cleaner structured output
- **DeepSeek strengths:** automation, persistent executor work, high-volume subtasks, long-context economics
- **Key caveat:** Provider route matters more than model name — `openai-codex/gpt-5.6-luna`, `openrouter/openai/gpt-5.6-luna`, `deepseek/deepseek-v4-flash`, and `openrouter/deepseek/deepseek-v4-flash` are four different tests
- **Pricing:** Direct OpenAI $0.20/M input, $1.20/M output; OpenRouter $0.10/M input, $0.60/M output
- **Context:** 1.05M tokens, 128K output, function calling, structured outputs, image input, configurable reasoning effort
- **Recommendation:** Luna plans and reviews; DeepSeek executes — but test on a clean direct route before judging
### Guanaco v0.8.14 (2026-08-14)
- Updated from v0.8.11 on CT205 — three incremental bug-fix releases focused on DSML/tool-call output hygiene
- DSML tag leakage fixes: malformed tool-call output no longer leaks raw syntax junk into model context
- Dashboard performance fix: cached content size, fixed wrong column names in history query
- No new features, no breaking changes — low-risk update
- See [[systems/guanaco]]
### DeepSeek V4 Flash 0731 (2026-07-31)
- **Same model ID:** `deepseek-v4-flash` — no ID change, just a post-training bump
- **Same architecture:** 284B total params, 13B active per token, 1M context window
- **Same pricing:** $0.14/M input, $0.28/M output
- **Benchmark jump:** Terminal Bench 2.1: 82.7 (up from 61.8 preview — 34% relative gain)
- **Now surpasses V4-Pro preview** (72.1) on agent/coding benchmarks
- **Stricter tool name validation:** Rejects tool names not matching `^[a-zA-Z0-9_-]+$` — may cause issues with OmniRoute-transformed tool names
- **Impact:** Your existing cron jobs using `deepseek-v4-flash` get the upgrade automatically. The empty-stream fallback issue (`opencode-go/deepseek-v4-flash` returning empty 200s) is a provider plumbing problem, not a model quality problem.
### OmniRoute Free-DeepSeek Silent-Stop Mitigation (2026-08-04)
- Free-DeepSeek upstreams returning empty-after-tool responses since Aug 1 → silent turn stops on `deepseek-flash-free` and `smart-route`/`good-free` across both instances
- Reworked both combos: strategy `priority``fill-first`, dropped proven culprits (`238d5978/deepseek-v4-flash` + `-0731`, `nvidia` deepseek), healthy free models promoted, deepseek demoted to last resort
- See [[incidents/2026-08-04-omniroute-free-deepseek-silent-stops]]
## Context Enforcement
- Three layers: model native → combo `data.config.context_length` → Hermes client cap
- Effective context is the minimum of the three
## Combo Behavior
- `priority` retries same provider up to `maxRetries`, then errors; it does NOT fall through
- `fill-first` walks the list on failure
- `auto`/`lkgp` recalculates per request using health scoring
- Same-provider entries in a combo share rate-limit buckets
- Session affinity caches provider in `session_model_history`
## Related
- [[concepts/model-routing]]
- [[systems/omniroute]]
- [[systems/opencode-go]]