-
Hugging Face
- Bern, Switzerland
Highlights
- Pro
Stars
Local AI text-to-speech with voice cloning and voice design, powered by GGML. C++17 port of Qwen3-TTS (QwenLM/Qwen3-TTS). 10 languages, 24 kHz mono output, runs on CPU, CUDA, ROCm, Metal, Vulkan.
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
End-to-end deployment of a fully local speech-to-speech conversation pipeline (STT + LLM + TTS) on Reachy Mini Lite, accelerated on AMD Strix Halo (gfx1151) via ROCm 7.13 / TheRock. Includes instal…
Making a mini version of the BDX droid. https://discord.gg/UtJZsgfQGe
Pure-PyTorch inference for CohereLabs/cohere-transcribe-03-2026 (2B Conformer + Transformer ASR, 14 languages).
🤗 ml-intern: an open-source ML engineer that reads papers, trains models, and ships ML models
Give your agents the power of the Hugging Face ecosystem
Mount Hugging Face Buckets and repos as local filesystems. No download, no copy, no waiting.
Easy fine-tuning for Qwen3-TTS: Fast voice cloning and high-quality multilingual speech synthesis.
The most accurate natural language detection library for Python, suitable for short text and mixed-language text
Per-collection OCR leaderboards using VLM-as-judge
AssetOpsBench - Industry 4.0: A unified benchmark and framework for building, orchestrating, and evaluating domain-specific AI agents for Industry 4.0 asset operations and maintenance, with 460+ sc…
Complete setup instructions for getting ESP32 Trinity working with HUB75 LED matrix panels including a web interface.
Talk with Reachy Mini!
✨🤝✨ Build instant multiplayer webapps, no server required — Magic WebRTC matchmaking over BitTorrent, Nostr, MQTT, IPFS, Supabase, and Firebase
A lightweight, local-first, and free experiment tracking library from Hugging Face 🤗
StreamingVLM: Real-Time Understanding for Infinite Video Streams
The open source codebase powering HuggingChat
📓 computational document system build on uv and markdown
🤗 The largest hub of ready-to-use datasets for AI models with fast, easy-to-use and efficient data manipulation tools
Harnessing 1.4M GPT4V-synthesized Data for A Lite Vision-Language Model





