问题
重度使用 AI 编程智能体(如 Codex)的开发者,每个工作步骤都会产生大量工具调用输出,这些输出全部回传给模型,导致上下文不断膨胀、token 成本失控。有用户自述 "After maxing out sub and burning $700/day per person on api",另一条讨论也在主动求 "hacks to shave even more costs"。核心痛点不是模型不好用,而是用得越多、账单越贵。
目标用户
重度使用 AI 编程智能体的独立开发者和工程团队(尤其订阅额度打满、转入 API 按量付费的团队)。市场随智能体编程普及快速扩大,但单点付费意愿尚待验证。
解决方案
- 拦截编程智能体工具调用响应的代理层(本地或云部署);
- 微调小模型(信号中提到基于 qwen)自动识别并剥离工具输出中的冗余内容;
- 保持 prompt 前缀不变以保留 KV cache,避免缓存失效带来的额外成本;
- 节省量统计面板,展示每天省了多少 token 和多少钱;
- 兼容 Codex 优先,后续扩展到其他智能体。AI 的角色是被微调的压缩模型,以及可选的信息摘要决策。
为什么是现在
编程智能体(Codex 等)刚进入重度使用阶段,长对话 + 工具输出导致 token 消耗暴涨,成本问题集中爆发;同时小模型微调成本已足够低,使得用 "fine tuned a compression model" 做实时裁剪在工程和经济上都刚刚变得可行。
MVP 范围
第一版只做一件事:作为本地/云端代理拦截 Codex 的工具调用响应,用微调压缩模型裁剪后返回,保留 KV cache。明确不做:多 agent 框架、不支持 Codex 以外的 agent、不做压缩质量评估面板。
变现方式
按开发者席位订阅,或按节省 token 量抽成。付费意愿为暗示性证据:目标用户已在以 "burning $700/day per person on api" 的强度付费给模型厂商,对确定性省钱工具有天然付费动机,但尚无对该产品本身的直接出价数据。
风险
平台依赖:压缩代理卡在 Codex 等智能体的协议中间,上游协议一变就可能失效; 上游吞并:模型厂商或 agent 厂商可能原生集成上下文压缩,直接消灭第三方空间; 技术风险:裁剪工具输出可能丢掉关键信息导致 agent 决策变差,且要小心不破坏 KV cache(这是省钱的关键); 竞品风险:Astra 已发布同类产品,Esus 从框架层面解决同一问题,窗口期不长。
已有产品
信号证据
这张卡片依据的原始讨论。摘录保持原文,点“原文”查看上下文。
Hacker News · Show HN9月30日讨论▲ 115 条评论隐含付费意愿
“Just lmk ur feedback and hacks to shave even more costs on astra!”
People are looking for hacks and methods to reduce token costs when using coding agents.原文
Hacker News · Show HN9月30日新品发布▲ 115 条评论隐含付费意愿
“After maxing out sub and burning $700/day per person on api, we fine tuned a compression model to trim codex's tool call output to reduce input token + cache.”
Coding agents like Codex send large tool call outputs back to the model, inflating API token costs, so a compression proxy was built to trim them.原文
Hugging Face 论坛 · Show and Tell9月24日新品发布▲ 0
“In my current benchmark, Esus makes more model calls than Claude for the same prompts (75 vs 35), but consumes far fewer tokens: about 55k vs 3.9M.”
AI agent workflows keep accumulating context in one long conversation, consuming huge token counts; a new agent framework (Esus) keeps workflow state outside the model in a deterministic kernel so LLM calls stay small and independent.原文