E.M. Lewis-Jong, founder and CEO of Mozilla Data Collective, has been named to the Foundever AI 100 UK 2026, a list recognising the people shaping how AI is built and used across the UK. EM's focus is on fixing the data issue in AI. Systems trained on narrow data fail the people left out of it, from farmers in Kenya and India to speakers of the hundreds of languages mainstream models barely consider. At Mozilla Data Collective, that has meant building the infrastructure for a different model where data creators set the terms of use, contributors are compensated, and consent is written into the license. Today the platform offers datasets spanning more than 450 languages. The public vote is now open and will help decide the final order of the list. If this work matters to you, we'd be grateful for your vote. Vote here: https://lnkd.in/ewPdZ669
Mozilla Data Collective
Software Development
The Data-Sharing Platform for Human Agency and Fair Value Exchange. Multilingual, Multicultural and Multimodal.
About us
Mozilla Data Collective is the platform for data agency and fair value exchange. We enable communities to build a tech future that's more multilingual, multicultural and multimodal - on their own terms.
- Website
-
https://mozilladatacollective.com/
External link for Mozilla Data Collective
- Industry
- Software Development
- Company size
- 11-50 employees
- Headquarters
- London
- Type
- Privately Held
Locations
-
Primary
Get directions
London, GB
Employees at Mozilla Data Collective
Updates
-
That's a wrap for Mozilla Data Collective at Interspeech 2026 #Interspeech2026 in Sydney! We had a great time learning all about the incredible research the community is doing, showcasing our platform for ethical data and value exchange, and hearing about the #data pain points you experience every day. Special shout out to our Visiting Researcher, Assistant Professor Jay L. Cunningham, Ph.D., and to our Regional Researcher Dr. Katerina Vylomova - so great to see you both in person. We also got to chat with Professor Vukosi Marivate about all things data science, and with Dr Sue Keay about the growing need for sovereign data and infrastructure - what a delight! Special thanks to Professor Steven Bird for a keynote that was absolute 🔥 and for the wonderful chat. De-colonising language technology is close to our heart. We also chatted with Shinji Watanabe about large datasets and #CommonVoice, with Michael McAuliffe about his excellent work on the Montreal Forced Aligner, and with Shaomei Wu who is helping us make gating work better for sensitive datasets. Thank you all so much for your insights! A huge shout out to fellow exhibitors Magic Data and DataBaker Technology for being such great neighbours, and we're looking forward to seeing what we might be able to do together. Likewise, to Oxford Wave Research Ltd, to Sarvam, to Mundo AI and to so many other potential partners - it was a delight to meet you, learn about your business and potential opportunities. Lastly a huge shout out to all the greenshirt volunteers and behind the scenes staff at International Convention Centre Sydney (ICC Sydney) for making Interspeech such a delight 👋
-
-
5 hours of speech collected in a single day, directly from contributors who recorded the data that reflects how the language really sounds. That is how Safi Data built the first sizable corpus of real-world Sheng speech housed on Mozilla Data Collective. Safi Data sends collection requests over WhatsApp and Telegram, speeding up the logistics of gathering important datasets. Contributors record themselves saying whatever they'd normally say, and their compensation is routed to them as soon as the session is processed. What would have traditionally taken 4-8 weeks of studio time, Safi Data compresses into one day. The result is a dataset that reflects how Sheng actually sounds in unscripted tone, code switched, and ready to incorporate into AI models and tools. If you're building or evaluating ASR for East Africa, a one-hour sample from more than 70 speakers is live now on MDC's Compensated Marketplace: https://lnkd.in/g4CygSSx #SpeechAI #ASR #Kenya #Sheng #DataForAI #MozillaDataCollective
-
-
Mozilla Data Collective reposted this
Mozilla Data Collective Raises $5 Million From Mozilla. Mozilla Data Collective has raised $5 million from Mozilla to support the expansion of its AI data-sharing platform. Founded by E.M. Lewis-Jong, the UK-based company enables communities and organisations to make cultural and linguistic datasets available to AI developers under defined access and licensing terms. The company says its platform supports more than 350 organisations sharing over 1,700 datasets across more than 450 languages. The new funding will support multimodal cultural video datasets, larger text corpora across EU, African and South Asian languages, new licensing and subscription options, and additional security and data-improvement capabilities.
-
-
Synthetic speech, also called #TTS or text-to-speech, is human-sounding speech produced by a generative AI model. As creating #TTS models has become technically much easier over the last few years, many legal and ethical issues have arisen. Join our Head of R&D, Kathy Reid MBA, who will Chair the Special Session at #Interspeech2026 on "Safeguarding Synthetic Speech: ethical, legal and technical issues", with accepted papers covering many aspects of this important discussion. When: Tuesday 29th September 2pm to 4pm at Area 14-3
-
-
Languages can be spoken by hundreds of thousands of people, taught in schools, and used in daily conversation, and still have almost nothing recorded or transcribed in a form an AI model could learn from. When that happens, those languages become invisible to the systems shaping how people communicate, search, and find information. Saturday was the European Day of Languages, a useful reminder of how much linguistic diversity exists on a single continent. Europe is home to more than 200 indigenous languages. The EU recognises 24 official languages and protects more than 60 regional and minority languages on top of that. Mozilla Data Collective exists to make sure that knowledge shapes the AI being built today, on terms set by the communities who hold it. In our collection you'll find Breton speech recordings, Welsh corpora like Bangor Siarad and CorCenCC that took years of fieldwork to build, Chuvash text-to-speech, and Catalan, which accounts for some of the largest volumes of recorded speech on the platform. Each one comes with clear licensing and documented consent from the people who contributed to it. See how many European languages are already on the platform, ready to use: https://lnkd.in/ey_Tmu3e #EuropeanDayOfLanguages #LinguisticDiversity #AI #LanguageData #MozillaDataCollective
-
-
We're delighted to be exhibiting in booth E7 at #Interspeech2026, next door to DataBaker Technology, ElevenLabs and StepFun, all this week. We'd love for you to drop by to learn more about the multilingual, multicultural and multimodal datasets we have available on the platform, and learn how your organisation can generate a revenue stream from your dataset. The booth will be staffed by our Head of R&D, Kathy Reid MBA, and she loves nothing more than learning about your speech research and your data needs! Rumour has it she has a stash of great stickers and tote bags to give away, so don't miss out! #MozillaDataCollective #InterSpeech2026
-
-
Mozilla Data Collective reposted this
Mozilla Data Collective secures $5 million in funding from Mozilla to scale its data-sharing platform for a more equitable AI data ecosystem. The investment will support multilingual, multicultural dataset expansion and new licensing options for AI builders. Founded by E.M. Lewis-Jong, the platform curates vetted, consentful datasets spanning over 450 languages and 350 contributing organisations, promoting fair value exchange for data contributors. More at: https://lnkd.in/ezjg5yTB #AI #DataEconomy #Funding
-
-
Better AI language and voice capabilities could help a health worker access guidance in an emergency, a farmer get timely weather advice, or a student use a learning tool in a language they actually understand. For years, we've worked directly with organisations and communities to close this gap. Today, that work becomes part of a shared five-year commitment, alongside organisations across technology, research, philanthropy and civil society, to help ensure people who speak languages currently underrepresented in AI can use AI tools in their own language and voice. Closing that gap starts with the datasets AI is built on. Which datasets get created, whose knowledge is represented and which languages receive investment all shape who these technologies ultimately work for. At Mozilla Data Collective, we've spent the last two years working with organisations and communities around the world to bring high-quality datasets from hundreds of languages into the AI ecosystem, on terms that work for the people behind them. This is bigger than any one organisation, which is why commitments like this matter. We're proud to be part of the work to make AI useful to more people, in the languages they actually speak. If you're building for an underrepresented language, we'd love to hear about it. Share it with us: partnerships@mozilladatacollective.com Read the full joint statement: https://lnkd.in/dkKrR2EW #CollectivePower #AIforGood Gates Foundation
-
-
Mozilla Data Collective reposted this
Mozilla Data Collective raises $5 million from Mozilla to scale a more equitable data ecosystem for AI. "The idea that we have to choose between giving AI builders the data they need and giving people agency and fair value is a false choice. We can do both." – E.M. Lewis-Jong, Founder and CEO of Mozilla Data Collective For full story, see comments #MozillaDataCollective #Mozilla #AI #DataEcosystem #EquitableAI #TechIntelPro
-