Skip to content
View RayeRen's full-sized avatar
🎯
Focusing
🎯
Focusing

Organizations

@msra-alumni @MLNLP-World @NATSpeech

Block or report RayeRen

Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

MOSS-Transcribe-Diarize 0.9B is an open-source SOTA end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness.

Python 1,374 71 Updated Jul 24, 2026

KVAE-Audio: a continuous full-band audio waveform autoencoder

Python 102 6 Updated Jul 23, 2026

UniRL is a Framework for Unified Multimodal Model Reinforcement Learning

Python 868 59 Updated Jul 31, 2026

Perfect Green Screen Keys

Python 14,497 883 Updated May 28, 2026

Write HTML. Render video. Built for agents.

TypeScript 39,001 3,680 Updated Aug 1, 2026

HY-WU (Part I): An Extensible Functional Neural Memory Framework and An Instantiation in Text-Guided Image Editing

Python 297 13 Updated Mar 18, 2026

A SOTA Industrial-Grade Voice Activity Detection & Audio Event Detection, supporting 100+ languages, outperforming Silero-VAD, TEN-VAD, FunASR-VAD and WebRTC-VAD

Python 478 32 Updated May 6, 2026

IDE Opener is a macOS utility that lets you launch your favorite IDE (VS Code, Cursor, or Windsurf) directly from the Finder toolbar. Skip the hassle of navigating through directories or using the …

9 Updated Jul 27, 2025

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Python 22,316 2,721 Updated Jul 14, 2026

Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.

Python 8,478 1,367 Updated Jul 8, 2026

Pre-built wheels that erase Flash Attention 3 installation headaches.

Python 112 9 Updated Jul 21, 2026

Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, a…

Rust 6,893 786 Updated Aug 1, 2026

HunyuanImage-2.1: An Efficient Diffusion Model for High-Resolution (2K) Text-to-Image Generation​

Python 674 53 Updated Oct 14, 2025

A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone

Python 26,070 2,041 Updated Jul 23, 2026

VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo

Python 2,117 239 Updated Aug 1, 2026

CUDA Python: Performance meets Productivity

Cython 3,327 315 Updated Aug 1, 2026

Open-source unified multimodal model

Python 6,132 548 Updated May 4, 2026

Text-audio foundation model from Boson AI

Python 8,308 643 Updated Jun 5, 2026

Tiny-FSDP, a minimalistic re-implementation of the PyTorch FSDP

Python 112 9 Updated Aug 20, 2025

[ICLR2025, ICML2025, NeurIPS2025 Spotlight] Quantized Attention achieves speedup of 2-5x compared to FlashAttention, without losing end-to-end metrics across language, image, and video models.

Cuda 3,526 447 Updated Jan 17, 2026

🌐 Make websites accessible for AI agents. Automate tasks online with ease.

Python 107,440 11,818 Updated Jul 31, 2026
Python 219 18 Updated Mar 21, 2023
Python 24 Updated May 28, 2025

[NeurIPS 2025] An official implementation of Flow-GRPO: Training Flow Matching Models via Online RL

Python 2,449 170 Updated May 7, 2026

ACE-Step: A Step Towards Music Generation Foundation Model

Python 4,707 604 Updated Feb 15, 2026

[ICCV2025] From Reusing to Forecasting: Accelerating Diffusion Models with TaylorSeers

Python 409 23 Updated Mar 2, 2026

Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation

Python 4,705 372 Updated Jun 21, 2025

The uncompromising Python code formatter

Python 41,773 2,831 Updated Jul 31, 2026
Python 6,080 472 Updated Jun 15, 2026
Next