<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>AI 模型评测 on Text Matrix</title><link>https://txtmix.com/categories/ai-%E6%A8%A1%E5%9E%8B%E8%AF%84%E6%B5%8B/</link><description>Recent content in AI 模型评测 on Text Matrix</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Tue, 21 Jul 2026 20:06:14 +0800</lastBuildDate><atom:link href="https://txtmix.com/categories/ai-%E6%A8%A1%E5%9E%8B%E8%AF%84%E6%B5%8B/index.xml" rel="self" type="application/rss+xml"/><item><title>To Run or Not to Run——ISSTA 2026 论文深度解读，LLM 程序修复的代码执行成本收益分析</title><link>https://txtmix.com/posts/tech/arxiv-2606-26978-code-execution-cost-effectiveness-llm-program-repair/</link><pubDate>Tue, 30 Jun 2026 20:56:00 +0800</pubDate><guid>https://txtmix.com/posts/tech/arxiv-2606-26978-code-execution-cost-effectiveness-llm-program-repair/</guid><description>&lt;h2 id="译序为什么这篇-issta-2026-论文值得完整读">译序：为什么这篇 ISSTA 2026 论文值得完整读&lt;/h2>
&lt;p>论文标题就一句话：&lt;strong>To Run or Not to Run&lt;/strong>——LLM 编程 Agent 是否应该默认执行代码？&lt;/p>
&lt;p>这是 2026 年所有 AI 编程助手（Claude Code / Codex / OpenCode / Cursor / OpenHands / SWE-agent）都默认开启的能力。你问任何 AI 工程师，TA 都会说&amp;quot;执行测试反馈是 agent 修复 bug 的关键&amp;quot;。&lt;/p></description></item><item><title>Ornith-1.0 深度解读：自改进开源 agentic coding 模型，4 个尺寸 + Self-Improving Framework 突破</title><link>https://txtmix.com/posts/tech/ornith-1-self-improving-agentic-coding-model-deep-dive/</link><pubDate>Tue, 30 Jun 2026 16:24:00 +0800</pubDate><guid>https://txtmix.com/posts/tech/ornith-1-self-improving-agentic-coding-model-deep-dive/</guid><description>&lt;h2 id="译序这份-readme-不只是模型发布说明">译序：这份 README 不只是模型发布说明&lt;/h2>
&lt;p>Ornith-1.0 是 deepreinforce-ai 团队在 2026 年发布的开源 agentic coding 模型系列。README 表面上看起来像&amp;quot;模型发布说明&amp;quot;，但读进去会发现它有三个不寻常的地方：&lt;/p></description></item><item><title>Qwen 3.6 27B 是本地开发的甜蜜点：一份 226 行实测译文 + 工程化拆解</title><link>https://txtmix.com/posts/tech/quesma-qwen-3-6-blog-deep-dive/</link><pubDate>Tue, 30 Jun 2026 16:07:00 +0800</pubDate><guid>https://txtmix.com/posts/tech/quesma-qwen-3-6-blog-deep-dive/</guid><description>&lt;h2 id="译序为什么这篇译文值得完整保留">译序：为什么这篇译文值得完整保留&lt;/h2>
&lt;p>Quesma 创始人 Piotr Migdał 在 2026-06-29 写了一篇关于 Qwen 3.6 27B 的实测长文。标题是 &amp;ldquo;Qwen 3.6 27B is the sweet spot for local development&amp;rdquo;——一句话点题：27B 是当前能在 Macbook 或单卡 RTX 上跑出实用水平的&lt;strong>最优尺寸&lt;/strong>。&lt;/p></description></item></channel></rss>