跳到正文
  1. Hugging Face Blog74

    Liquid AI 发布面向端侧的多模态决策模型 d1-3B 与 d1-omni-600M

    Liquid AI 发布两款开放权重决策模型 d1-3B 和实验性的 d1-omni-600M,分别支持文本与图像、文本与图像或文本与音频输入。在七个公开数据集上,d1-3B 的平均分为 82.9,d1-omni-600M 为 78.4;d1-3B 在 Jetson AGX Thor 上单个问题延迟为 16 ms,在 Jetson Orin Nano 上为 50 ms。

    推荐理由:文章将模型结构、公开数据集对比和端侧延迟放在一起,便于判断两款决策模型在质量、模态支持与设备部署之间的取舍。

  1. Google DeepMind77

    Google DeepMind 发布 EmbeddingGemma 2 多模态嵌入模型

    Google DeepMind 发布 EmbeddingGemma 2,将文本、代码、图像、视频和音频映射到统一嵌入空间,面向端侧推理。

    推荐理由:EmbeddingGemma 2 将文本、图像、音频和视频统一到共享嵌入空间,并公开端侧内存、上下文和向量压缩数据,便于判断其本地检索适用性。

已经到底了