Stars
MOSS-Transcribe-Diarize 0.9B is an open-source SOTA end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.
KVAE-Audio: a continuous full-band audio waveform autoencoder
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
Write HTML. Render video. Built for agents.
HY-WU (Part I): An Extensible Functional Neural Memory Framework and An Instantiation in Text-Guided Image Editing
A SOTA Industrial-Grade Voice Activity Detection & Audio Event Detection, supporting 100+ languages, outperforming Silero-VAD, TEN-VAD, FunASR-VAD and WebRTC-VAD
IDE Opener is a macOS utility that lets you launch your favorite IDE (VS Code, Cursor, or Windsurf) directly from the Finder toolbar. Skip the hassle of navigating through directories or using the …
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
Pre-built wheels that erase Flash Attention 3 installation headaches.
Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, a…
HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
CUDA Python: Performance meets Productivity
Text-audio foundation model from Boson AI
Tiny-FSDP, a minimalistic re-implementation of the PyTorch FSDP
[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.
🌐 Make websites accessible for AI agents. Automate tasks online with ease.
[NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL
ACE-Step: A Step Towards Music Generation Foundation Model
[ICCV2025] From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers
Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation





