问题
自托管 LLM 工作流的开发者在用 n8n 跑 agent、用 Open WebUI 接推理模型时,输出层处处漏气:Ollama Chat Model 节点强制流式,"Log files are hard to read when the model writes in many small chunks",下游 JSON 解析也被打断;AI Agent 节点配流式 Webhook 时,"Any intermediate actions—such as tool invocations, inputs, and outputs—execute silently in the background",前端看不到工具调用过程;Open WebUI 会把历史里的 reasoning/thinking 标签重新注入 prompt,导致模型"imitating/amplifying"这些标签,最终"break frontend reasoning parsing/rendering"。三个问题指向同一层:自托管 LLM 栈的流式输出缺乏统一治理。
目标用户
自托管 LLM 技术栈的开发者和自动化工程师:用 n8n 搭 agent 工作流、通过 Ollama 跑本地模型、用 Open WebUI 接推理模型的个人开发者与小团队。市场是自托管 AI 工具链用户群体,规模可观但付费习惯弱,多为开源工具免费用户。
解决方案
- 流式模式开关:对 Ollama 请求提供流式/非流式切换,默认可配,解决日志和 JSON 解析问题(对应 signal 1)
- 结构化中间事件流:捕获 LangChain 的 handleToolStart/handleToolEnd 等回调,转成带类型的 SSE 事件或 JSON chunk 转发给前端(对应 signal 2)
- reasoning 标签治理:在历史注入 prompt 前剔除 thinking/reasoning 块,默认关闭可配,防止模型模仿标签破坏前端解析(对应 signal 3)
- 统一部署形态:作为本地代理或 n8n 自定义节点交付,不改动用户现有工作流
AI 本身不是产品核心;产品核心是对 LLM 流式输出(文本块、工具调用、推理标签)的解析与重组层,用规则 + 事件处理实现,模型无关。
为什么是现在
三条信号都指向近期才出现的技术变化:推理型模型(thinking/reasoning 标签)开始普及,自托管用户通过 Ollama 等本地方式跑这类模型,导致历史注入、标签污染这类新问题集中爆发(signal 3 有 15 分热度、44 条评论);同时 agent 工作流进入生产化阶段,开发者开始要求前端能看到中间工具调用(signal 2)。新模型行为 + 自托管工作流生产化,共同催生了对输出流治理层的即时需求。
MVP 范围
做: 一个本地部署的 HTTP 代理服务(或 n8n 自定义节点 + Open WebUI 反代插件):Ollama 流式/非流式模式切换开关;把 agent 的 tool call 生命周期转成带类型的结构化事件流;可选剔除历史中的 reasoning/thinking 标签再转发。 不做: 不做 UI 前端、不做模型托管、不做多平台统一 API 网关、不做可观测性大盘——只解决输出流这一层。
风险
平台依赖: 三个痛点都在 n8n/Open WebUI 的 issue 跟踪里,平台方加一个默认开关即可让独立产品失去存在理由,这是最大风险。 付费意愿缺失: 全部信号来自免费开源工具用户,willingness_to_pay 均为 none,可能只能走开源 + 服务模式。 碎片化: 不同 provider/model 的标签和事件格式不一致,维护成本会随模型生态膨胀。 证据单薄: 三条信号相互独立、engagement 普遍偏低,共享痛点是推断而非直接验证。
已有产品
| 产品 | 定位 | 价格 |
|---|---|---|
| n8n | 用户在其社区直接请求流式开关与中间事件转发,平台内置后可覆盖部分需求 | — |
| Open WebUI | 用户已提交 reasoning 标签剔除的功能请求,说明平台在补齐该能力 | — |
| LangChain | 其 callback 机制(handleToolStart/handleToolEnd)是解决中间事件可见性的关键底层依赖 | — |
信号证据
这张卡片依据的原始讨论。摘录保持原文,点“原文”查看上下文。
n8n 社区 · 功能请求9月24日功能请求▲ 10 条评论
“Log files are hard to read when the model writes in many small chunks.”
n8n 的 Ollama Chat Model 节点始终以流式方式调用 Ollama,用户无法关闭流式,导致日志难读、代理缓冲异常和下游 JSON 解析困难。原文
n8n 社区 · 功能请求9月20日功能请求▲ 00 条评论
“Any intermediate actions—such as tool invocations, inputs, and outputs—execute silently in the background.”
When using the AI Agent node with a streaming Webhook, only the final LLM text chunks are currently piped to the HTTP SSE response stream and intermediate tool calls execute silently.原文
GitHub Issues · open-webui/open-webui4月2日功能请求▲ 1544 条评论
“For some providers/models, once these tags appear in history, the model starts **imitating/amplifying** them (multiple/nested/variant thinking tags), which can eventually **break frontend reasoning parsing/rendering**”
Open WebUI reinjects reasoning/thinking tags from chat history into prompts, causing models to imitate the tags and breaking frontend reasoning parsing and rendering.原文