OpenSpeaks
OpenSpeaks is a community network of language documenters and archivists. We work with speakers of endangered, Indigenous and other local languages in India, Nepal and Sri Lanka to record their languages on audio and video, archive the recordings where they can be cited, and use them on Wikipedia and other Wikimedia projects. Community documenters are paid for this work, at or above local market rates. Between 2017 and 2026, we documented over 20 South Asian languages.
OpenSpeaks began in 2017 as a multimedia toolkit for language digital activists, written on Wikimedia Commons with a grant from the National Geographic Society. The toolkit is now kept on Wikiversity.
Where to find OpenSpeaks
- OpenSpeaks Archives on Meta-Wiki: languages, fellows, plans and reports.
- Recordings on Wikimedia Commons.
- Tools on Wikimedia Toolforge, with source code on Wikimedia GitLab.
- Archiving the Present, the network, manual and course that grew out of our 2026 workshops.
OpenSpeaks Archives
OpenSpeaks Archives is our public archive of community-led recordings in low-resourced languages. It started as a pilot in July 2024, when five community archivists documented their languages. From July 2025 to June 2026, seven OpenSpeaks Fellows documented 13 languages, uploaded 143 files and improved 350 pages across about 70 Wikimedia projects. Those pages were read 599,429 times in that year. Altogether, OpenSpeaks Archives has improved more than 1,000 Wikipedia and Wikimedia pages in over 127 languages.
| Fellow or coordinator | Languages |
|---|---|
| Kimmi Pal | Marcha (Rongpo) |
| Surendra Singh Pangtey | Johari |
| Arun Gour | Jaunpuri, Jaunsari |
| Jaiprakash Chauhan | Bangani |
| Opino Gomango | Sora, Juray, Juang, Gorum (Parengi) |
| Nenavath Mohan | Lambadi |
| Sanjib Chaudhary | Saptariya Tharu (Nepal) |
| Uday Raj Aaley | Raji (Nepal) |
| Sajehan Buckman | Sri Lankan Malay (Sri Lanka) |
Recordings are kept with a citation at the Endangered Languages Archive, the Language Archive Cologne (Kusunda, Sora) and the Library of Congress, which holds over 71,000 of our audio files. Our plan for July 2026 to June 2027 adds communities in Bangladesh and Mali.
OpenSpeaks Tools
We build a tool only when nothing else does the job for a language archivist. All four are free and open source, run on Wikimedia Toolforge, and three of them, Subtitler, Tome and Bento, have been stable since August 2026.
- Subtitler captions audio and video from Wikimedia Commons or your own files, translates the subtitles and uploads them to Commons. Ranjithsiji is the lead developer. Arpitha Bhandary fixed bugs and added features, and Opino Gomango tested it in the field.
- Tome writes metadata for oral knowledge recordings and generates consent and licensing agreements. Its language list maps ISO 639 codes. Govind Lal T. L. added features, with technical guidance from Jnanaranjan Sahu.
- Bento sorts, measures and compresses media files and records codec, resolution, loudness and SHA-256 checksums. It works offline.
- OpenSpeaks Media Dashboard shows which Wikimedia pages use our Commons files and how many people read those pages each month.
Arpitha Bhandary and Govind Lal T. L. built their changes at the Indic Wikimedia Hackathon, Hyderabad, 25 to 28 June 2026. Thank you, Bharathesha A, and the organisers at IIIT Hyderabad.
Framework and research
The OpenSpeaks Oral Knowledge Framework sets out how to record, describe, license, archive and cite oral knowledge. It follows the FAIR principles (findable, accessible, interoperable, reusable) and the CARE principles (collective benefit, authority to control, responsibility, ethics), so that communities keep control over recordings of their own languages.
- Subhashish Panigrahi, Opino Gomango and Kimmi Pal. OpenSpeaks Archives: Citing Low-Resourced Language Oral History Multimedia. Wiki Workshop 2026.
- Subhashish Panigrahi. Citing Oral History: A Framework for Low-Resourced Language Audio-visual Documentation. Language Documentation and Archiving 2026, Berlin.
- Subhashish Panigrahi. Building Foundational Language Data with A Community-First Approach. Language Documentation and Archiving 2024, Berlin.
- Subhashish Panigrahi. OpenSpeaks before AI: Frameworks for Creating the AI/ML Building Blocks for Low-Resource Languages. Interactions 30(3), 2023.
- Subhashish Panigrahi. OpenSpeaks: Transforming Learning of Tech Innovations in Low-Resource Language Documentations to Open Educational Resources. Linguapax Review 9, 2021.
In 2026 we presented this work at the Wikimedia Futures Lab in Frankfurt, a seminar at TU Darmstadt, WikiConference India in Kochi (with Ramjit Tudu and Sanjib Chaudhary) and IndiaFOSS in Bengaluru (with Sneha Mundari).
Archiving the Present
Archiving the Present is a community network for archiving oral languages, memories and knowledges, organised by OpenSpeaks. It grew out of our Community Language Documentation and Archiving workshops: eight online sessions between 18 May and 24 August 2026, for 101 applicants from 14 countries working in about 115 languages. The network launched in Kochi on 3 September 2026. Nineteen community documenters and archivists came, working in 27 languages and varieties. The sessions became a manual, in print and as an e-book, and a free course on WikiLearn. Sama Digital Foundation donated 10 laptops to participants who asked for a device, through Ansh Arora.
Films
Our documentaries use the OpenSpeaks toolkit. Gyani Maiya (2019) is the memoir of Gyani Maiya Sen Kusunda, co-written with Uday Raj Aaley and Sanjib Chaudhary. Mage Porob (2019) was made with Ho villagers in Odisha and Remosam (2019) with the Bondak people of the Bonda hills. See all our films.
Support and partners
OpenSpeaks has been supported by the National Geographic Society (2017), Mozilla Open Leaders (2017), the Online News Association MJ Bear Fellowship (2017), Creative Commons, Grant for the Web (2021), the Wikimedia Foundation (2024 to 2027) and Samagata Foundation. Kiwix is our fiscal sponsor for Wikimedia Foundation grants.
We work with the Language Archive Cologne, the Endangered Languages Archive, Rising Voices, Odia Wikimedians User Group, Wikimedians of Santali Language User Group, Igbo Wikimedians User Group, Dagbani Wikimedians User Group, Been There Doon That / Humanities Himalaya and Nature Science Initiative. Mandana Seyfeddinipur and Padmini Ray Murray advise us.
Work with us
If you document a language and want to work with us, write to us through the contact page. Try the tools and report problems on Wikimedia GitLab.