Speech

  1. Qwen2.5-Omni: Alibaba's End-to-End Multimodal AI Model

    Qwen2.5-Omni is Alibaba's flagship open-source multimodal AI model, developed by the QwenLM team at Alibaba Cloud. As a single end-to-end model...

    AI
  2. MiniCPM-o: Open-Source Multimodal LLM for Vision, Speech, and Text

    Multimodal AI models that can simultaneously process vision, speech, and text represent the cutting edge of artificial intelligence. OpenAI's GPT-4o...

    AI