OneLastIdea

AI 编程代理 harness 定制与评测工具

让 AI 编程代理的高级用户定制执行回路、子代理与路由,并用评测数据迭代自己的 harness 配置。

  • AI + 服务
  • B2B
  • 全球
  • 难度 4/5
  • 启动 $500–5k
  • MVP 约 8 周

问题

Claude Code 等编程代理的 power user 和 agent 开发者面对的是黑盒 harness:行为不合意时,他们现在可以通过官方新增的 Claude Mods 系统 “customizing the execution loop, UI, subagents, routing, and behavior”,但如何改、改了之后是否更好,缺乏可评测的反馈闭环;研究侧(Claw-Eval)也指出可以由 “the same frozen model… as the proposer, reads the complete run records and directly edits the harness” 来优化 harness,说明 harness 的可编辑与可评测正在成为真实需求。注意:三条信号均为厂商发布与论文,属于能力信号而非直接的用户抱怨,痛点强度存疑。

目标用户

AI 编程代理的 power user(如 Claude Code 重度用户)与构建 agent harness 的开发者/研究者。信号未给出市场规模数据;从信号看这是一个随代理普及而扩张、但付费意愿未经验证的开发者细分。

解决方案

构建一个代理 harness 的「定制 + 评测」工具层:

  • 可视化 harness 配置:解析 Claude Code 的 mods / 执行回路 / 子代理 / 路由定义,让用户以图形方式理解和修改;
  • A/B 评测沙盒:同一任务集跑两版 harness 配置,对比成功率、步数、token 成本;
  • 运行记录驱动的优化建议:借鉴 Claw-Eval 思路,让模型读取完整 run records 后提出 harness 编辑建议(AI 角色:既是被测的 solver,也是提出改动的 proposer);
  • 可移植的 mods 管理:把定制从单一厂商生态中抽出,导出/复用到其他代理(Vibe、Codex 等)。

为什么是现在

三条能力信号在短期内密集出现:Claude Mods 把 harness 定制开放为官方能力并被称为 “mutable software” 的早期预览;Mistral 发布 Medium 3.5 驱动的远程编程代理;Claw-Eval 论文证明冻结模型可以自优化 harness。三者共同说明「harness 可编辑」正从研究走向产品可用阶段,第三方工具层存在时间窗口。

MVP 范围

第一版只做:导入 Claude Code 的 mods/配置与运行日志 → 可视化展示执行回路、子代理与路由结构 → 按改动生成 A/B 评测对比(任务成功率、步数、成本)→ 支持把胜出的改动写回配置。不做:多模型/多平台适配、自动改写 harness 代码(self-optimization 只做只读分析)、团队协作与云端托管。

风险

平台依赖(最大风险):Claude Mods 明确是 early preview,接口随时可能变动或收紧;Anthropic 官方推进 Mods 即在吞掉本机会。技术风险:harness 差异大(Claude Code vs Codex vs Vibe),做跨平台通用层的工程量被低估的风险高。竞争风险:厂商自建(Mistral 的 Vibe 远程代理、Anthropic 的 Mods)速度快且有分发优势。研究风险:自优化 harness 的方法尚在论文阶段(Claw-Eval),产品化成熟度低。

已有产品

产品定位价格
Codex另一个主流代理 harness,其插件与配置生态是潜在替代—
Claude Code被定制的目标代理本体,其内置能力演进可能吞掉第三方层—
Claude ModsAnthropic 官方的 Claude Code harness 定制系统,正是本机会最直接的占位者—
VibeMistral 的远程编程代理平台,代表厂商自建 harness 定制方向的竞争—

信号证据

这张卡片依据的原始讨论。摘录保持原文,点“原文”查看上下文。

  1. Mistral AI News10月2日新能力

    “Introducing Mistral Medium 3.5, remote coding agents in Vibe, plus new Work mode in Le Chat for complex tasks.”

    Mistral AI launched remote coding agents in Vibe powered by the new Mistral Medium 3.5, plus a new Work mode in Le Chat for complex tasks.原文

  2. arXiv cs.AI10月1日新能力

    “the same frozen model, on the same version of the harness, first solves tasks as the solver and then, as the proposer, reads the complete run records and directly edits the harness that runs it.”

    Researchers can improve agent harnesses by having the same frozen model act as its own optimizer, editing the harness that runs it across diverse tasks.原文

  3. Latent Space9月29日新能力

    “Claude Mods: customizing the execution loop, UI, subagents, routing, and behavior of Claude Code”

    Claude Code users can now customize the agent's execution loop, UI, subagents, routing and behavior via the new Claude Mods system.原文

这个 Idea 怎么样?

相似的 Idea

  • 43

    开发者的长任务 AI 智能体工作流助手

    帮开发者编排、监控和调度跨工具的 AI 长任务,把零散的 agent 操作变成可复用的工作流。

    2 条信号,2 个来源,最近 13小时前

    开发者工具
    • SaaS
    • B2B
    • 全球
    • 难度 3/5
    • 启动 $500–5k
    • MVP 约 8 周
  • 48

    开发者的统一 Agent 会话 CLI 管理器

    在一个终端里创建、接管、恢复和监控多个 AI 编码 Agent 会话,不用在网页和 CLI 之间切换。

    2 条信号,2 个来源,最近 13小时前

    开发者工具
    • 桌面 App
    • B2B
    • 全球
    • 难度 3/5
    • 启动 < $500
    • MVP 约 4 周
  • 44

    遗留代码库的 AI 渐进式现代化工具

    让维护老旧代码的工程团队用 Agent 分模块读懂、重构、验证遗留系统,先出报告再渐进改造。

    3 条信号,3 个来源,最近 13小时前

    开发者工具
    • 浏览器插件
    • B2B
    • 全球
    • 难度 4/5
    • 启动 $5k–50k
    • MVP 约 12 周

本页内容由 AI 根据公开讨论整理,最后更新于 14分钟前。发现错误或需要下架原文,请看这里。

AI 编程代理 harness 定制与评测工具 · OneLastIdea