<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>音视频生成 on Text Matrix</title><link>https://txtmix.com/tags/%E9%9F%B3%E8%A7%86%E9%A2%91%E7%94%9F%E6%88%90/</link><description>Recent content in 音视频生成 on Text Matrix</description><generator>Hugo</generator><language>zh-cn</language><lastBuildDate>Tue, 21 Jul 2026 20:06:14 +0800</lastBuildDate><atom:link href="https://txtmix.com/tags/%E9%9F%B3%E8%A7%86%E9%A2%91%E7%94%9F%E6%88%90/index.xml" rel="self" type="application/rss+xml"/><item><title>LTX-2 音视频联合 DiT 拆解：第一个一体化音视频基础模型</title><link>https://txtmix.com/posts/tech/lightricks-ltx-2-audio-video-foundation-model-guide/</link><pubDate>Thu, 18 Jun 2026 21:03:00 +0800</pubDate><guid>https://txtmix.com/posts/tech/lightricks-ltx-2-audio-video-foundation-model-guide/</guid><description>&lt;h1 id="ltx-2-音视频联合-dit-拆解第一个一体化音视频基础模型">LTX-2 音视频联合 DiT 拆解：第一个一体化音视频基础模型&lt;/h1>
&lt;p>&lt;code>Lightricks/LTX-2&lt;/code> 想做的事情在 README 第一句话里就讲清楚了：&lt;em>&amp;ldquo;the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one model&amp;rdquo;&lt;/em>。这是它和市面上&amp;quot;先出视频、再配音&amp;quot;两阶段方案的最大差异——&lt;strong>音视频联合训练、联合推理&lt;/strong>，而不是两个模型拼出来的伪同步。&lt;/p></description></item></channel></rss>