Recommendation for General assistant

a General Assistant

Our top recommendation for a General Assistant, based on the public evidence we track, is OpenAI: GPT-5.6 Sol.[1][2][3] Choose for extended daily chat sessions where the model's tendency to catch and correct its own mistakes improves reliability over earlier OpenAI versions. Watch out: Verify your thinking level is actually active, as the UI may default to 5.5 instant with effort indicator hidden, requiring manual selection of 'think harder' to access Sol's capabilities. Google: Gemini 3.6 Flash is the next-ranked alternative. Select when benchmark-validated instruction following is prioritized, as it scores 75.37% on LiveBench Instruction Following (#8 of 58).

About this recommendation

Updated
Sep 25, 2026
Evidence through
Sep 25, 2026
Sources
24
Revision
v80

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

22

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

2 of 6

OpenAI

Established source breadth. 13 citation families and 4 practitioner families support the top result; 3 cautionary threads is retained. The largest citation family contributes 25%.

Sources evaluated

The task sets these weights before any model is scored.

winner: GPT-5.6 Sol
Evaluation feedWeightWinner resultField measured
LMArena Instruction Following
30%
#1221/22
LiveBench Reasoning
25%
#422/22
LiveBench Instruction Following
20%
#1722/22
LMArena Text
15%
#1321/22
OpenRouter usage
10%
97/10022/22

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

OpenAI33%
  • OpenAI2 models
  • Anthropic1 model
  • Google1 model
  • Meta1 model
  • Qwen1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01GPT-5.6 SolOpenAI
90
100%7 threads · 4 families · 3 cautions#4 LiveBench Reasoning · #12 LMArena Instruction Following
02Gemini 3.6 FlashGoogle
81
100%5 threads · 3 families · 3 cautions#8 LiveBench Instruction Following · #16 LMArena Text
03Claude Opus 4.6Anthropic
81
100%2 threads · 1 families · 2 cautions#1 LMArena Instruction Following · #2 LMArena Text
04Muse Spark 1.2Meta
78
100%3 threads · 3 families · 2 cautions#4 LMArena Text · #9 LiveBench Reasoning
05GPT-5.6 LunaOpenAI
77
100%6 threads · 5 families · 3 cautions#25 LiveBench Reasoning · #38 LMArena Instruction Following
06Qwen3.8 27BQwen
75
100%2 threads · 2 families · 0 cautions#14 LiveBench Instruction Following · #37 LMArena Instruction Following

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. GPT-5.6 Sol provides reliable daily conversational use with strong self-correction behavior, though benchmark standings trail top competitors in instruction following.

    Best when: Choose for extended daily chat sessions where the model's tendency to catch and correct its own mistakes improves reliability over earlier OpenAI versions.

    Tips

    • Choose for extended daily chat sessions where the model's tendency to catch and correct its own mistakes improves reliability over earlier OpenAI versions.
      Source 1
      “5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.”
    • Use at medium thinking settings when balancing response speed against quality for general questions.
      Source 1
      “5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.”
    • Leverage its calibrated tone adaptation, as users report it adjusts explanation depth based on inferred domain knowledge (simpler for law questions, more technical for politics).
      Source 1
      “5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.”

    Watch out for

    • Verify your thinking level is actually active, as the UI may default to 5.5 instant with effort indicator hidden, requiring manual selection of 'think harder' to access Sol's capabilities.
      Source 2
      “For me on a paid plan, the effort indicator was hidden and the model was 5.5 instant. I had to press the + button to select “think harder” before the dial that allowed Sol medium or high to be selected to show. It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?”
    • Expect intermittent transport errors in long sessions, with observed failure rates in extended conversations (up to 6 consecutive fetch failures in a 311-message session).
      Source 4
      “## 背景 Pi Web Desktop 中 openai-codex(ChatGPT 订阅 OAuth)会话频繁出现「输出一半停止 / fetch failed / terminated」,其他模型(aliyun/deepseek)偶发但次数少。 ## 实锤数据(2026-08-01,scripts/session-stops.mjs 修复后全量扫描 185 会话) - transport-error 共 37 条:openai-codex/GPT 19(gpt-5.6-sol×11 + gpt-5.6-terra×8)、aliyun 9、pi-router 8、deepseek 1。 - 最大异常会话:311 消息 / ~2.2M tokens,会话开头连续 6 次 fetch failed/terminated。 - compaction 仅 1 次且在非中断点,**不是主因**;主因是 transport error。 ## 工具修复(commit bfe020d) session-stops.mjs 原判定只在独立 type=error 事件时识别 transport-er…”
    • Confirm connection stability before relying on it, as test connections and normal chat requests may fail with certain OpenAI-compatible provider configurations.
      Source 5
      “### Environment - ReadAware: v0.2.10 - OS: macOS - Provider: Custom (OpenAI-compatible) ### Configuration - API format: OpenAI-compatible API format provided by the Sub2API upstream - Smart model: `gpt-5.6-sol` - Fast model: `gpt-5.6-luna` - API key: valid and redacted ### Steps to reproduce 1. Configure a Custom (OpenAI-compatible) provider with the settings above. 2. Save the configuration. 3. Click **Test Connection**. 4. Try a normal chat request. ### Observed behavior Both the connection t…”
  2. Gemini 3.6 Flash achieves strong LiveBench instruction-following scores but exhibits integration fragility and streaming behavior that may disrupt conversational flow.

    Best when: Select when benchmark-validated instruction following is prioritized, as it scores 75.37% on LiveBench Instruction Following (#8 of 58).

    Tips

    • Select when benchmark-validated instruction following is prioritized, as it scores 75.37% on LiveBench Instruction Following (#8 of 58).
      Source 3
      “Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗

    Watch out for

    • Avoid for new user onboarding flows, as API key configuration during setup triggers permanent 'Applying your changes' failures due to provider ID mismatches in admin/catalog routing.
      Source 6
      “## Summary A brand new customer who connects **Google Gemini with an API key** during onboarding is permanently stuck at "Applying your changes" and can never reach chat. The customer plane pushes the provider id `gemini`; admin only accepts `google`, and rejects the push with a 502. This is onboarding-blocking for the default path: `gemini-3.6-flash` is the `is_default` model in admin's catalog, and a fresh wizard run configures exactly one model, which routes through the broken code path. ##…”
    • Watch for client deserialization failures in streaming mode, as heavy-reasoning responses may emit usage chunks without completion_tokens, breaking strict OpenAI-compatible clients.
      Source 7
      “## Problem When proxying Antigravity/Gemini responses to the OpenAI chat-completions format, the `usage` object in a streamed chunk (and potentially non-stream responses) can be emitted **without** `completion_tokens`. Strict OpenAI clients that deserialize `usage` with required fields (e.g. Rust serde in Grok CLI) fail the whole turn: Observed with `gemini-3.6-flash-high` via Antigravity OAuth. Heavy-reasoning responses appear to trigger it: upstream sends a `usageMetadata` chunk containing `p…”
    • Expect 'No response: RAW' errors in book chat features on certain e-reader platforms, where streaming responses complete but fail to render.
      Source 8
      “**Describe the bug** Returns “No response: RAW” whenever using ‘book chat’ feature. First it does show a streaming response to my prompt, but when it finishes I get the error. **To Reproduce** Using Openrouter and Gemini 3.6 Flash, when I use the ‘Book Chat’. **Expected behavior** It to show me the response to the chat I submitted. **e-reader (please complete the following information):** - Device: Musnap Ocean - OS: Android 14 - KOReader **Additional context** Haven’t had any issues with any o…”
      rockettemortonOpen original ↗
    • Anticipate latency in search-augmented conversations, as the current architecture waits 15-45 seconds for full search results before displaying any response.
      Source 9
      “## 背景 現在は単一Gemini呼び出し(検索1回+抽出1回)の結果が全て揃ってから画面に表示される。検索フェーズだけで15〜45秒かかるため、ユーザーが待たされる体感が長い。 ## やりたいこと ユーザーを飽きさせないよう、パースできた情報から部分的にストリーミング出力したい。一気に全部出るより、分かった情報から順次出す体験にしたい。 ## 論点(要設計検討) - 現在の設計は「検索1回→抽出1回」の単一呼び出しに意図的に統合した経緯がある(多段パイプラインからの移行、PR #98)。ストリーミング化はこの設計判断と衝突しうる - Gemini APIのstreamGenerateContentを使うか、部分JSONの逐次パースが必要になる - 既存のSuspense分割(改札欄・出口欄が別Promiseで先出しされる仕組み、RouteGateStat/RouteExitStat)を活用できる可能性がある - レイテンシ改善(gemini-3.6-flash移行、PR #101)である程度は緩和されている点も考慮 ## スコープ 実装前に設計方針を固める必要あり。まずは調査・設計…”
  3. Claude Opus 4.6 leads the instruction-following category in human preference rankings while scoring strongly on reasoning benchmarks, positioning it as a capable general conversationalist.

    Best when: Use for tasks requiring careful adherence to complex instructions, as it ranks #1 on LMArena's instruction-following leaderboard with an Elo of 1523.

    Tips

    • Use for tasks requiring careful adherence to complex instructions, as it ranks #1 on LMArena's instruction-following leaderboard with an Elo of 1523.
      Source 10
      “Ranks #1 of 146 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.”
      LMArena instruction-following categoryOpen original ↗
    • Deploy for reasoning-heavy conversations where objective ground-truth performance matters, given its 88.67% score on LiveBench Reasoning.
      Source 11
      “Scores 88.67% on LiveBench Reasoning (#16 of 58), an objective, ground-truth-scored evaluation refreshed with new questions.”
      LiveBench ReasoningOpen original ↗

    Watch out for

    • Watch for fabricated self-justifications when challenging its outputs, as the model may generate false explanations about its own reasoning process rather than admitting errors.
      Source 12
      “### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…”
      NStestUser1954Open original ↗
      Source 13
      “### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…”
      NStestUser1954Open original ↗
  4. Muse Spark 1.2 delivers strong reasoning performance on LiveBench while maintaining competitive instruction-following capabilities in human evaluations.

    Best when: Select for reasoning-intensive queries where benchmark-validated performance is key, as it scores 90% on LiveBench Reasoning (#9 of 58).

    Tips

    • Select for reasoning-intensive queries where benchmark-validated performance is key, as it scores 90% on LiveBench Reasoning (#9 of 58).
      Source 14
      “Scores 90% on LiveBench Reasoning (#9 of 58), an objective, ground-truth-scored evaluation refreshed with new questions.”
      LiveBench ReasoningOpen original ↗
    • Use as a fallback when other models face availability issues, as it was confirmed functional in free-tier onboarding scenarios alongside other backup models.
      Source 15
      “## 背景 真实新用户纯净态验收:清空全部 kito + opencode 本地数据(备份在案)→ dev `kito-preview` 冷启动 → 多 agent 扮演新用户实测。sidecar `127.0.0.1:64830`,确认 `first launch onboarding pending: true`。 ## 实测结果 ### ✅ 通过项 - **未登录免费模型发消息可用**(原始投诉场景):7/8 免费模型正常返回回复(space-bunny / mimo-v2.6-flash / muse-spark-1.3 / ling-3.0-flash / nemotron-3.5-lightning / muse-spark-1.2 / big-pickle) - Onboarding 自动完成(建 `~/Documents/Default Project` + 草稿会话),全简体中文 UI,品牌 Kito - 数据隔离:kito 目录自重建(data 0700),opencode 目录零再生,注册文件一致,无越界写 - Telegram 登录对话框正常(二维码+校验码+…”
      chenjiangjiang85-jpgOpen original ↗

    Watch out for

    • Expect potential integration friction with certain proxy configurations, as the model triggered invalid_request_error on OpenCodex proxy due to parameter forwarding issues.
      Source 16
      “Proxy: opencodex 2.51.0, Codex 0.147.0, default provider openai. Model: opencode-go/muse-spark-1.3-contributor (also affects 1.2, both enabled in catalog). Error on every call, old and new threads, all thinking levels (including explicit max): Upstream request failed: [invalid_request_error] unknown parameter Notes: - DeepSeek Flash via the same opencode-go provider works, so auth and proxy routing look fine. - Catalog maps Spark under an opus Claude alias family; suspect the proxy forwards a f…”
    • Verify availability on your specific provider, as it returned 181001 errors or silent failures on some backend configurations in multi-model testing.
      Source 17
      “### 分支选择 dev-stable ### 模块选择 LLM wiki ### Checklist - [x] 我已经搜索过相关问题,但没有得到预期的帮助。 - [x] 最新版本中该错误尚未修复。 - [x] 请注意,如果您提交的Bug描述缺少相应的环境信息和最小可复现的demo,我们将很难复现和解决该问题,从而降低收到反馈的可能性,甚至该问题将被关闭。 ### 🐞 问题详细描述 ●内置模型来源:后端 Zen 提供的 9 个 free 模型(前端"内置免费模型"列表) 问题汇总 # 前端显示名 后端标识符(已知/待确认) 现象 错误码 错误类型 1 DeepSeek V4 Flash deepseek-v4-flash-free 调用报错 181001 / HTTP 400 BadRequestError: Model is unavailable 2 OX Alpha Free 待确认(推测带-free后缀) 对话无反应 / 无输出 无(无报错) 静默失败(疑似超时或未捕获错误) 3 Muse Spark 1.2 待确认(推测带-free后缀) 调用报错 181001 / H…”
      openjiuwen-collaboration-bot[bot]Open original ↗
  5. GPT-5.6 Luna offers solid reasoning performance with broad deployment flexibility, though operational observations reveal cache efficiency concerns in high-volume settings.

    Best when: Deploy for reasoning tasks where benchmark scores guide selection, as it achieves 85.64% on LiveBench Reasoning.

    Tips

    • Deploy for reasoning tasks where benchmark scores guide selection, as it achieves 85.64% on LiveBench Reasoning.
      Source 18
      “Scores 85.64% on LiveBench Reasoning (#28 of 58), an objective, ground-truth-scored evaluation refreshed with new questions.”
      LiveBench ReasoningOpen original ↗

    Watch out for

    • Monitor token costs in high-throughput applications, as production analysis showed only 0.84% cache hit rate versus 20.5% for predecessor models when using legacy prompt-cache contracts.
      Source 19
      “## Problem High-volume conversation-processing routes were moved to `gpt-5.6-luna`, but their callers still use the pre-GPT-5.6 prompt-cache contract: a global `prompt_cache_key` plus `prompt_cache_retention="24h"`, without GPT-5.6 request-wide cache options or explicit cache breakpoints. **Operational observation from a privacy-safe closed production analysis (2026-07-23 17:00–23:00 UTC):** - the Luna population read **0.84%** of input tokens from cache, versus **20.5%** for the preceding matc…”
      Git-on-my-levelOpen original ↗
    • Check your client library's compatibility, as default agent initialization may fail with 500 errors when reasoning configuration is omitted for this model.
      Source 20
      “## Summary Chatting with the default model `openai:gpt-5.6-luna` fails with a 500. The underlying OpenAI error is: Reproduced against a running instance as `admin@example.com`: `openai:gpt-5.6-luna` is currently the system default returned by `GET /api/llm/models`, so this is the out-of-the-box chat path. ## Root cause Every agent model is built by `init_graph()` in `backend/src/agents/__init__.py`: No reasoning configuration is ever passed, and `ChatOpenAI` defaults to the Chat Completions end…”
    • Watch for capacity-related failures during peak usage, as Pro tier users reported 'Selected model is at capacity' errors even with quota remaining.
      Source 21
      “### What version of the Codex App are you using (From “About Codex” dialog)? 1 ### What subscription do you have? 1 ### What platform is your computer? _No response_ ### What issue are you seeing? 我是 ChatGPT Pro 20X 用户,在 Windows 桌面端 Work 持续收到错误: “Selected model is at capacity. Please try a different model.” 原项目中 GPT‑6 Astra 和 GPT‑5.6 Sol 均报错;新建对话后,仅发送“在吗”,GPT‑5.6 Sol 和 GPT‑5.6 Luna 仍然报同样的错误。用量页面显示共享周额度剩余 88%。 我已完成新对话、极短消息和切换模型测试。请结合这些失败请求的后台记录,核查是否存在其他用量限制、模型访问问题或服务异常,并提供具体处理办法。 ### What steps…”
      shr723shr-cloudOpen original ↗
  6. Qwen3.8 27B offers open-weight deployment flexibility with competitive instruction-following performance, particularly suited for local hosting scenarios with memory constraints.

    Best when: Deploy locally when data residency or cost control is critical, as it runs efficiently on consumer AMD GPUs (RX 7900 GRE) with quantized formats fitting in VRAM without OOM risk.

    Tips

    • Deploy locally when data residency or cost control is critical, as it runs efficiently on consumer AMD GPUs (RX 7900 GRE) with quantized formats fitting in VRAM without OOM risk.
      Source 22
      “## Destination A single lean llama.cpp `llama-server` (HIP, gfx1100) serving Qwen3.8-27B (unsloth `UD-Q3_K_XL`) on the RX 7900 GRE — KV-cache pre-sized to fit in VRAM so it physically cannot OOM, exposed OpenAI-compatible, fronted by the Hollama browser UI as the opencode replacement, and launched under a systemd unit with a hard memory cap. "Done" = a usable model I can chat/code with: no swap, no 27GB RAM spike, correct context actually provisioned. ## Notes - Domain: local LLM serving, llama…”
      daryleriversOpen original ↗
    • Use for interactive chat workloads requiring low time-to-first-token, as it maintains resident model state with 262K context on appropriate hardware.
      Source 23
      “## Why Two workload shapes with opposite requirements are currently served by one always-on GPU: - **Sprinkled interactive traffic** (Discord, chat, vision, classifier, pi turns). Needs a resident model and low TTFT. Served well today by the 4090 running Qwen3.8-27B at 262K context. - **Clumped batch work** (`model-bench` sweeps, bulk classification, embeddings backfills, autonomous queue jobs). Latency-insensitive, preemption-tolerant, and wants a much larger model than the 4090 can hold. The…”
    • Select for instruction-following tasks where open-weight options are required, scoring 72.66% on LiveBench Instruction Following.
      Source 24
      “Scores 72.66% on LiveBench Instruction Following (#15 of 58), including paraphrasing, simplifying, story generation, and summarization.”
      LiveBench Instruction FollowingOpen original ↗

Frequently asked

What is the top-ranked model for a General Assistant?
OpenAI: GPT-5.6 Sol ranks first in the current evidence-weighted comparison. Choose for extended daily chat sessions where the model's tendency to catch and correct its own mistakes improves reliability over earlier OpenAI versions.[1]
What should I watch out for with OpenAI: GPT-5.6 Sol?
Verify your thinking level is actually active, as the UI may default to 5.5 instant with effort indicator hidden, requiring manual selection of 'think harder' to access Sol's capabilities.[2]
What is an alternative to OpenAI: GPT-5.6 Sol?
Google: Gemini 3.6 Flash is the next-ranked option. Select when benchmark-validated instruction following is prioritized, as it scores 75.37% on LiveBench Instruction Following (#8 of 58).[3]

Sources

  1. 1

    “5.6 Sol is a joy to use for "daily chat" as well. Compared to earlier OpenAI models it catches and corrects its mistakes very reliably. It also seems way smarter in tuning its replies to areas I am more less knowledgeable about (i.e. when I ask it a law question, it assumes I know as much as a toddler which is true, but on political topics it more easily throws around terminology) and including analogies. On medium thinking, it's a very good compromise between speed and quality.”

    leokennis · Hacker News · Aug 18, 2026
  2. 2

    “For me on a paid plan, the effort indicator was hidden and the model was 5.5 instant. I had to press the + button to select “think harder” before the dial that allowed Sol medium or high to be selected to show. It made me wonder how many paid subscribers realize they are using the same 5.5 instant model as free users by default. A dark pattern or oversight?”

    aryehof · Hacker News · Aug 7, 2026
  3. 3

    “Scores 75.37% on LiveBench Instruction Following (#8 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026
  4. 4

    “## 背景 Pi Web Desktop 中 openai-codex(ChatGPT 订阅 OAuth)会话频繁出现「输出一半停止 / fetch failed / terminated」,其他模型(aliyun/deepseek)偶发但次数少。 ## 实锤数据(2026-08-01,scripts/session-stops.mjs 修复后全量扫描 185 会话) - transport-error 共 37 条:openai-codex/GPT 19(gpt-5.6-sol×11 + gpt-5.6-terra×8)、aliyun 9、pi-router 8、deepseek 1。 - 最大异常会话:311 消息 / ~2.2M tokens,会话开头连续 6 次 fetch failed/terminated。 - compaction 仅 1 次且在非中断点,**不是主因**;主因是 transport error。 ## 工具修复(commit bfe020d) session-stops.mjs 原判定只在独立 type=error 事件时识别 transport-er…”

    dust617 · GitHub · Aug 1, 2026
  5. 5

    “### Environment - ReadAware: v0.2.10 - OS: macOS - Provider: Custom (OpenAI-compatible) ### Configuration - API format: OpenAI-compatible API format provided by the Sub2API upstream - Smart model: `gpt-5.6-sol` - Fast model: `gpt-5.6-luna` - API key: valid and redacted ### Steps to reproduce 1. Configure a Custom (OpenAI-compatible) provider with the settings above. 2. Save the configuration. 3. Click **Test Connection**. 4. Try a normal chat request. ### Observed behavior Both the connection t…”

    lei1024 · GitHub · Jul 22, 2026
  6. 6

    “## Summary A brand new customer who connects **Google Gemini with an API key** during onboarding is permanently stuck at "Applying your changes" and can never reach chat. The customer plane pushes the provider id `gemini`; admin only accepts `google`, and rejects the push with a 502. This is onboarding-blocking for the default path: `gemini-3.6-flash` is the `is_default` model in admin's catalog, and a fresh wizard run configures exactly one model, which routes through the broken code path. ##…”

    kavin-114 · GitHub · Jul 30, 2026
  7. 7

    “## Problem When proxying Antigravity/Gemini responses to the OpenAI chat-completions format, the `usage` object in a streamed chunk (and potentially non-stream responses) can be emitted **without** `completion_tokens`. Strict OpenAI clients that deserialize `usage` with required fields (e.g. Rust serde in Grok CLI) fail the whole turn: Observed with `gemini-3.6-flash-high` via Antigravity OAuth. Heavy-reasoning responses appear to trigger it: upstream sends a `usageMetadata` chunk containing `p…”

    amikai · GitHub · Jul 21, 2026
  8. 8

    “**Describe the bug** Returns “No response: RAW” whenever using ‘book chat’ feature. First it does show a streaming response to my prompt, but when it finishes I get the error. **To Reproduce** Using Openrouter and Gemini 3.6 Flash, when I use the ‘Book Chat’. **Expected behavior** It to show me the response to the chat I submitted. **e-reader (please complete the following information):** - Device: Musnap Ocean - OS: Android 14 - KOReader **Additional context** Haven’t had any issues with any o…”

    rockettemorton · GitHub · Jul 24, 2026
  9. 9

    “## 背景 現在は単一Gemini呼び出し(検索1回+抽出1回)の結果が全て揃ってから画面に表示される。検索フェーズだけで15〜45秒かかるため、ユーザーが待たされる体感が長い。 ## やりたいこと ユーザーを飽きさせないよう、パースできた情報から部分的にストリーミング出力したい。一気に全部出るより、分かった情報から順次出す体験にしたい。 ## 論点(要設計検討) - 現在の設計は「検索1回→抽出1回」の単一呼び出しに意図的に統合した経緯がある(多段パイプラインからの移行、PR #98)。ストリーミング化はこの設計判断と衝突しうる - Gemini APIのstreamGenerateContentを使うか、部分JSONの逐次パースが必要になる - 既存のSuspense分割(改札欄・出口欄が別Promiseで先出しされる仕組み、RouteGateStat/RouteExitStat)を活用できる可能性がある - レイテンシ改善(gemini-3.6-flash移行、PR #101)である程度は緩和されている点も考慮 ## スコープ 実装前に設計方針を固める必要あり。まずは調査・設計…”

    sinoda1114 · GitHub · Jul 22, 2026
  10. 10

    “Ranks #1 of 146 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.”

    LMArena instruction-following category · Benchmark · Sep 13, 2026
  11. 11

    “Scores 88.67% on LiveBench Reasoning (#16 of 58), an objective, ground-truth-scored evaluation refreshed with new questions.”

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  12. 12

    “### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…”

    NStestUser1954 · GitHub · Jul 22, 2026
  13. 13

    “### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…”

    NStestUser1954 · GitHub · Jul 22, 2026
  14. 14

    “Scores 90% on LiveBench Reasoning (#9 of 58), an objective, ground-truth-scored evaluation refreshed with new questions.”

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  15. 15

    “## 背景 真实新用户纯净态验收:清空全部 kito + opencode 本地数据(备份在案)→ dev `kito-preview` 冷启动 → 多 agent 扮演新用户实测。sidecar `127.0.0.1:64830`,确认 `first launch onboarding pending: true`。 ## 实测结果 ### ✅ 通过项 - **未登录免费模型发消息可用**(原始投诉场景):7/8 免费模型正常返回回复(space-bunny / mimo-v2.6-flash / muse-spark-1.3 / ling-3.0-flash / nemotron-3.5-lightning / muse-spark-1.2 / big-pickle) - Onboarding 自动完成(建 `~/Documents/Default Project` + 草稿会话),全简体中文 UI,品牌 Kito - 数据隔离:kito 目录自重建(data 0700),opencode 目录零再生,注册文件一致,无越界写 - Telegram 登录对话框正常(二维码+校验码+…”

    chenjiangjiang85-jpg · GitHub · Sep 24, 2026
  16. 16

    “Proxy: opencodex 2.51.0, Codex 0.147.0, default provider openai. Model: opencode-go/muse-spark-1.3-contributor (also affects 1.2, both enabled in catalog). Error on every call, old and new threads, all thinking levels (including explicit max): Upstream request failed: [invalid_request_error] unknown parameter Notes: - DeepSeek Flash via the same opencode-go provider works, so auth and proxy routing look fine. - Catalog maps Spark under an opus Claude alias family; suspect the proxy forwards a f…”

    jlicerio · GitHub · Sep 19, 2026
  17. 17

    “### 分支选择 dev-stable ### 模块选择 LLM wiki ### Checklist - [x] 我已经搜索过相关问题,但没有得到预期的帮助。 - [x] 最新版本中该错误尚未修复。 - [x] 请注意,如果您提交的Bug描述缺少相应的环境信息和最小可复现的demo,我们将很难复现和解决该问题,从而降低收到反馈的可能性,甚至该问题将被关闭。 ### 🐞 问题详细描述 ●内置模型来源:后端 Zen 提供的 9 个 free 模型(前端"内置免费模型"列表) 问题汇总 # 前端显示名 后端标识符(已知/待确认) 现象 错误码 错误类型 1 DeepSeek V4 Flash deepseek-v4-flash-free 调用报错 181001 / HTTP 400 BadRequestError: Model is unavailable 2 OX Alpha Free 待确认(推测带-free后缀) 对话无反应 / 无输出 无(无报错) 静默失败(疑似超时或未捕获错误) 3 Muse Spark 1.2 待确认(推测带-free后缀) 调用报错 181001 / H…”

    openjiuwen-collaboration-bot[bot] · GitHub · Aug 24, 2026
  18. 18

    “Scores 85.64% on LiveBench Reasoning (#28 of 58), an objective, ground-truth-scored evaluation refreshed with new questions.”

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  19. 19

    “## Problem High-volume conversation-processing routes were moved to `gpt-5.6-luna`, but their callers still use the pre-GPT-5.6 prompt-cache contract: a global `prompt_cache_key` plus `prompt_cache_retention="24h"`, without GPT-5.6 request-wide cache options or explicit cache breakpoints. **Operational observation from a privacy-safe closed production analysis (2026-07-23 17:00–23:00 UTC):** - the Luna population read **0.84%** of input tokens from cache, versus **20.5%** for the preceding matc…”

    Git-on-my-level · GitHub · Jul 24, 2026
  20. 20

    “## Summary Chatting with the default model `openai:gpt-5.6-luna` fails with a 500. The underlying OpenAI error is: Reproduced against a running instance as `admin@example.com`: `openai:gpt-5.6-luna` is currently the system default returned by `GET /api/llm/models`, so this is the out-of-the-box chat path. ## Root cause Every agent model is built by `init_graph()` in `backend/src/agents/__init__.py`: No reasoning configuration is ever passed, and `ChatOpenAI` defaults to the Chat Completions end…”

    ryaneggz · GitHub · Aug 7, 2026
  21. 21

    “### What version of the Codex App are you using (From “About Codex” dialog)? 1 ### What subscription do you have? 1 ### What platform is your computer? _No response_ ### What issue are you seeing? 我是 ChatGPT Pro 20X 用户,在 Windows 桌面端 Work 持续收到错误: “Selected model is at capacity. Please try a different model.” 原项目中 GPT‑6 Astra 和 GPT‑5.6 Sol 均报错;新建对话后,仅发送“在吗”,GPT‑5.6 Sol 和 GPT‑5.6 Luna 仍然报同样的错误。用量页面显示共享周额度剩余 88%。 我已完成新对话、极短消息和切换模型测试。请结合这些失败请求的后台记录,核查是否存在其他用量限制、模型访问问题或服务异常,并提供具体处理办法。 ### What steps…”

    shr723shr-cloud · GitHub · Sep 17, 2026
  22. 22

    “## Destination A single lean llama.cpp `llama-server` (HIP, gfx1100) serving Qwen3.8-27B (unsloth `UD-Q3_K_XL`) on the RX 7900 GRE — KV-cache pre-sized to fit in VRAM so it physically cannot OOM, exposed OpenAI-compatible, fronted by the Hollama browser UI as the opencode replacement, and launched under a systemd unit with a hard memory cap. "Done" = a usable model I can chat/code with: no swap, no 27GB RAM spike, correct context actually provisioned. ## Notes - Domain: local LLM serving, llama…”

    darylerivers · GitHub · Aug 19, 2026
  23. 23

    “## Why Two workload shapes with opposite requirements are currently served by one always-on GPU: - **Sprinkled interactive traffic** (Discord, chat, vision, classifier, pi turns). Needs a resident model and low TTFT. Served well today by the 4090 running Qwen3.8-27B at 262K context. - **Clumped batch work** (`model-bench` sweeps, bulk classification, embeddings backfills, autonomous queue jobs). Latency-insensitive, preemption-tolerant, and wants a much larger model than the 4090 can hold. The…”

    jomcgi · GitHub · Aug 26, 2026
  24. 24

    “Scores 72.66% on LiveBench Instruction Following (#15 of 58), including paraphrasing, simplifying, story generation, and summarization.”

    LiveBench Instruction Following · Benchmark · Jun 25, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.