LLMs and models

Moshi

Reference entryImage, video and audio models
Reference entryMoshi

Moshi is Kyutai’s open speech-language model for real-time, full-duplex spoken conversation with low-latency audio input, understanding, generation, and output.

Where it sits