<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>🏭 NVIDIA on Elon&#39;s AD Insight</title>
    <link>https://auto-driving-blog.pages.dev/tags/-nvidia/</link>
    <description>Recent content in 🏭 NVIDIA on Elon&#39;s AD Insight</description>
    <image>
      <title>Elon&#39;s AD Insight</title>
      <url>https://auto-driving-blog.pages.dev/images/share.png</url>
      <link>https://auto-driving-blog.pages.dev/images/share.png</link>
    </image>
    <generator>Hugo</generator>
    <language>zh-cn</language>
    <lastBuildDate>Tue, 28 Jul 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://auto-driving-blog.pages.dev/tags/-nvidia/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>论文精读｜NVIDIA Cosmos 3：全模态世界基础模型开启物理AI新纪元</title>
      <link>https://auto-driving-blog.pages.dev/posts/paper-reading/cosmos3-%E4%B8%96%E7%95%8C%E5%9F%BA%E7%A1%80%E6%A8%A1%E5%9E%8B%E7%B2%BE%E8%AF%BB/</link>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.pages.dev/posts/paper-reading/cosmos3-%E4%B8%96%E7%95%8C%E5%9F%BA%E7%A1%80%E6%A8%A1%E5%9E%8B%E7%B2%BE%E8%AF%BB/</guid>
      <description>NVIDIA Cosmos 3 用 Mixture-of-Transformers 双塔架构将语言、图像、视频、音频、动作五模态统一进单一世界基础模型。Reasoner 塔负责语义推理，Generator 塔负责高保真生成，两塔通过共享跨模态注意力深度耦合。在 8 项物理 AI 基准上取得开放模型第一，并以 OpenMDW-1.1 宽松许可证开源。</description>
    </item>
    <item>
      <title>论文精读｜VIMA：基于多模态提示的通用机器人操作——多模态大模型驱动机器人</title>
      <link>https://auto-driving-blog.pages.dev/posts/paper-reading/vima-multimodal-prompts-robot-manipulation/</link>
      <pubDate>Tue, 28 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.pages.dev/posts/paper-reading/vima-multimodal-prompts-robot-manipulation/</guid>
      <description>VIMA提出了一种统一的多模态提示接口，将多样化的机器人操作任务转化为序列建模问题。通过Transformer编码器-解码器架构和物体中心表示，VIMA在零样本泛化设置下任务成功率最高提升2.9倍。</description>
    </item>
    <item>
      <title>论文精读｜Vesta：统一的具身推理通用模型</title>
      <link>https://auto-driving-blog.pages.dev/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2606-20905/</link>
      <pubDate>Sun, 19 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://auto-driving-blog.pages.dev/posts/paper-reading/%E8%AE%BA%E6%96%87%E7%B2%BE%E8%AF%BB-2606-20905/</guid>
      <description>NVIDIA 的 Vesta 用单一 Qwen3-VL-8B 基础模型统一定位、空间推理、导航和长程规划四大具身智能能力，取代了传统的拼装专家路线。它通过统一的条件语言生成范式和简洁的多模态记忆挽具，让自注意力在历史帧与当前观察间建立长程依赖。平均超越各类别最优专家 20% 以上，真机任务成功率提升超 35%。</description>
    </item>
  </channel>
</rss>
