Vision

  1. MiniCPM-o: Open-Source Multimodal LLM for Vision, Speech, and Text

    Multimodal AI models that can simultaneously process vision, speech, and text represent the cutting edge of artificial intelligence. OpenAI's GPT-4o...

    AI
  2. GEMS: General Multimodal Sensing Framework

    The real world does not present information in a single modality. We experience it through vision, language, audio, and physical sensation...

    AI